Zhenwei Shi 0001

dblp:94/5806-1 · DBLP profile ↗
← Back
193ranked-venue papers
13as first author
136since 2021 · last 2026
0000-0002-4772-3172ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 140 · 4 first-author · 109 since 2021Artificial intelligence and machine learning · 34 · 8 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 1 first-author · 15 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Domain generalization via domain uncertainty shrinkage
Jun-Zheng Chu, Bin Pan, Tianyang Shi, Zhenwei Shi 0001
Pattern Recognit.4
2026 A copula-guided temporal dependency method for multitemporal hyperspectral images unmixing
Ruiying Li, Bin Pan, Qiaoying Qu, Zhenwei Shi 0001
Pattern Recognit.5
2026 Be Bayesian by attachments to catch more uncertainty
Bin Pan, Tianyang Shi, Tao Li 0022, Zhenwei Shi 0001
Pattern Recognit.5
2026 Blur-Resistant Hyperspectral Image Super-Resolution via Dual-Degradation Fusion Model
abstract
The deep unfolding network represents a promising research avenue in fusion-based hyperspectral image super-resolution (HSI-SR). However, most current deep unfolding methodologies are anchored in idealized observation models, which overlook the degradation of the multispectral image (MSI), hindering their SR performance and practical applicability. To address this problem, this paper establishes a novel Dual-Degradation Fusion (D2-Fusion) model, which incorporates both HSI degradation and MSI blurring into the HSI-SR modeling process. Subsequently, we apply the second-order semismooth Newton algorithm to solve the optimization problem in D2-Fusion model. The solution steps are then mapped into an end-to-end trainable network, termed Blur-resistant Hyperspectral image Super-Resolution Network (BHSR-Net). To the best of our knowledge, the proposed network is the first successful attempt to consider MSI blurring artifacts in the HSI-SR task. It offers several distinct advantages: 1) The network structure maintains a strict mathematical correspondence with the optimization algorithm, ensuring each module retains strong physical interpretability; 2) The network exhibits superior SR performance and strong generalization ability on both standard and real-world scenarios across five datasets; 3) The network demonstrates excellent learning efficiency with a compact architecture, and its lightweight variant achieves comparable results with only 38K parameters. The code is available at https://github.com/Dou0405/BHSR-Net.
Mai Xu, Yongxuan Dou, Xin Deng 0002, Zhenwei Shi 0001
IEEE Trans. Image Process.5
2026 SLM-VINS: Advancing Visual-Inertial SLAM via Hierarchical Spatial Line Integration and Multi-Mechanism Marginalization
abstract
Simultaneous Localization and Mapping (SLAM) has emerged as a cornerstone technology in intelligent transportation systems (ITS) and autonomous robots, finding widespread applications in various scenarios and multiple platforms autonomous driving tasks. However, in environments with weak textures and motion blur, achieving efficient and robust visual-inertial SLAM remains challenging. Current research utilizes line features to enhance SLAM performance in such environments, but incurs difficulties in line extraction, structured scene representation, and computational overheads involved in joint optimization. To address these challenges, this paper introduces a novel SLAM framework with both high pose estimation accuracy and high back-end computational efficiency. Firstly, an intelligent spatial line integration method is proposed to effectively reduce data redundancy in joint optimization by leveraging the remarkable stability of long line segments in structured scenes, thereby enhancing the spatial consistency of structural lines. Secondly, to combat low-light and high-speed motion environments, an optical flow tracking accuracy verification method is designed to bolster the system’s tracking performance and robustness in complex scenarios. Finally, to relieve the substantial computational overhead arising from high-dimensional optimization parameters in bundle adjustment (BA), a multi-mechanism marginalization strategy is presented to enhance the accuracy and computational efficiency of BA, while also preventing scale explosion. Comparative evaluations against state-of-the-art algorithms on both the EuRoC MAV and TUM VI benchmark datasets demonstrate that the proposed framework significantly improves localization accuracy and joint optimization efficiency.
Baoqi Huang, Bing Jia, Lifei Hao, Zhenwei Shi 0001
IEEE Trans. Intell. Transp. Syst.6
2026 QoE Evaluation for VR with Vibrotactile Feedback Based on Inter-user Brain Spatial Information
abstract
Subjective measurement remains one of the most widely used approaches for evaluating Quality of Experience (QoE) in tactile virtual environments. However, its reliability is often compromised by factors such as conscious bias, variations in user expressiveness, and contextual influences, which may distort the accuracy of evaluation outcomes. In light of the fact that Electroencephalography (EEG) provides a direct window into neurophysiological correlates of emotion and cognitive states, this article proposes an objective QoE evaluation method for Virtual Reality (VR) with vibrotactile feedback based on Inter-user Brain Spatial Information (IBSI). The proposed IBSI feature extraction method enhances the conventional Common Spatial Pattern (CSP) algorithm through the introduction of a regularization constraint term designed to mitigate overfitting. Moreover, the covariance matrix of each non-target user is weighted according to its Kullback–Leibler divergence from the target user, enhancing cross-user alignment and supporting effective transfer of neural information. To validate the performance of the proposed method, we design a VR shooting interaction experiment involving 64 participants. The study comprises three main phases: preparation, interaction, and subjective feedback. During the preparation phase, participants receive task explanations and familiarize themselves with the VR environment. In the interaction phase, participants complete a standardized shooting task, while EEG data are recorded synchronously. Finally, subjective feedback is collected through questionnaires. QoE assessment is accomplished by classifying IBSI features using classical classifiers, with subjective QoE ratings serving as ground truth. Experimental results show that our method outperforms existing methods in classification accuracy, and the dominant activation pattern in the alpha rhythm is consistent with neural mechanisms associated with motor perception. Meanwhile, the mutual interpretability between subjective and objective data characterizing QoE further validates the rationality of the experimental paradigm.
Yan Zhang 0052, Riting Xia, Zhenwei Shi 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Unified Multi-Agent Trajectory Modeling with Masked Trajectory Diffusion
Songru Yang, Zhenwei Shi 0001, Zhengxia Zou
ICCV2
2025 Open-CD: A Comprehensive Toolbox for Change Detection
abstract
We present Open-CD, a change detection toolbox that contains a rich set of change detection methods as well as related components and modules. The toolbox started from a series of open source general vision task tools, including OpenMMLab Toolkits, PyTorch Image Models (Timm), etc. It gradually evolves into a unified platform that covers many popular change detection methods and contemporary modules. It not only includes training and inference codes, but also provides some useful scripts for data analysis. We believe this toolbox is by far the most comprehensive change detection toolbox. In this report, we introduce the features, supported methods and applications of Open-CD. In addition, we also conduct a benchmarking study on different methods and components. We wish that the toolbox and benchmark could serve the growing research community by providing a flexible toolkit to re-implement existing methods and develop their own new change detectors. Code and models are available at https://github.com/likyoo/open-cd.
Kaiyu Li 0001, Chengxi Han, Yupeng Deng 0001, Keyan Chen 0001, Zhuo Zheng, Hao Chen 0045, Ziyuan Liu 0006, Yuantao Gu, Zhengxia Zou, Zhenwei Shi 0001, Sheng Fang 0001, Deyu Meng, Zhi Wang 0002, Xiangyong Cao
ACM Multimedia11
2025 Heterogeneous Mixture of Experts for Remote Sensing Image Super-Resolution
abstract
Remote sensing image super-resolution (SR) aims to reconstruct high-resolution remote sensing images from low-resolution inputs, thereby addressing limitations imposed by sensors and imaging conditions. However, the inherent characteristics of remote sensing images, including diverse ground object types and complex details, pose significant challenges to achieving high-quality reconstruction. Existing methods typically employ a uniform structure to process various types of ground objects without distinction, making it difficult to adapt to the complex characteristics of remote sensing images. To address this issue, we introduce a Mixture of Experts (MoE) model and design a set of heterogeneous experts. These experts are organized into multiple expert groups, where experts within each group are homogeneous while being heterogeneous across groups. This design ensures that specialized activation parameters can be employed to handle the diverse and intricate details of ground objects effectively. To better accommodate the heterogeneous experts, we propose a multi-level feature aggregation strategy to guide the routing process. Additionally, we develop a dual-routing mechanism to adaptively select the optimal expert for each pixel. Experiments conducted on the UCMerced and AID datasets demonstrate that our proposed method achieves superior SR reconstruction accuracy compared to state-of-the-art methods. The code will be available at https://github.com/Mr-Bamboo/MFG-HMoE.
Bowen Chen 0002, Keyan Chen 0001, Mohan Yang, Zhengxia Zou, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.5
2025 A Second-Order Stationarity-Based Confidence Assessment Method for Temperature Forecast
abstract
Remote sensing observations have the potential to improve the accuracy of temperature forecasts. However, the task of quantifying the confidence in these predictions remains challenging. Existing methods for confidence estimation, such as Bootstrap and Bayesian models, often suffer from computational inefficiencies and may impose modifications on the underlying predictor structures. To address these limitations, this letter introduces a novel and efficient confidence assessment framework for temperature forecasting, termed the second-order stationarity-based confidence assessment (SOS-CA). The proposed method is premised on the assumption that the second-order differences in temperature data adhere to a Gaussian distribution. Leveraging this assumption, SOS-CA employs statistical techniques to evaluate the Gaussianity of these second-order differences. Predictions that exhibit greater second-order stationarity are deemed to possess higher confidence. Moreover, we present a rigorous theoretical proof establishing the asymptotic equivalence of the mathematical transformations underpinning the SOS-CA methodology. To enhance its applicability, SOS-CA is extended to multiple variants to accommodate diverse forecasting scenarios. Extensive experiments using real-world remote sensing data substantiate the effectiveness of the proposed approach, demonstrating that SOS-CA achieves performance on par with or superior to existing methods while significantly reducing computational overhead.
Bin Pan, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.5
2025 Structural Representation-Guided GAN for Remote Sensing Image Cloud Removal
abstract
Optical remote sensing imagery is often compromised by cloud cover, making effective cloud-removal techniques essential for enhancing the usability of such data. We designed a novel structural representation-guided generative adversarial network (GAN) framework for cloud removal, in which structure and gradient branches are integrated into the network, helping the model focus on the structural representations of ground objects during image reconstruction. Different from previous methods that concentrate on recovering pixel information, we emphasize learning the structural information of remote sensing images. We then utilize error feedback to fuse features from the structural auxiliary branch, guiding the image reconstruction process. During the training phase, synthetic cloud images are used to supervise the optimization of the cloud-removal network, while real cloud images are employed in an adversarial training manner for unsupervised learning to improve the generalization ability of the network. Additionally, multitemporal revisit images from remote sensing satellites are employed as auxiliary inputs, aiding the network to remove thick clouds reliably. We evaluated our framework on a dataset derived from SEN12MS-CR, and the proposed method outperformed classical cloud-removal methods in both objective performance and subjective visual quality. Furthermore, compared to other methods, our approach achieved superior cloud-removal results on real images.
Keyan Chen 0001, Liqin Liu, Zhengxia Zou, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.6
2025 Multi-Grained Guided Diffusion for Quantity-Controlled Remote Sensing Object Generation
abstract
Accurate object counts represent essential semantical information in remote sensing imagery, significantly impacting applications like traffic monitoring and urban planning. Despite the recent advances in text-to-image generation in remote sensing, existing methods still face challenges in precisely controlling the quantity of object instances in generated images. To address this challenge, we propose a novel method, Multi-Grained Guidend Diffusion (MGDiff). During training, unlike previous methods that relied solely on latent-space noise constraints, MGDiff imposes constraints at three distinct granularities: latent pixel, global counting and spatial distribution. The multi-grained guidance mechanism matches the quantity prompts with object spatial layouts in the feature space, enabling our model to achieve precise control over object quantities. To benchmark this new task, we present Levir-QCG, a dataset comprising 10,504 remote sensing images across five object categories, annotated with precise object counts and segmentation masks. We conducted extensive experiments to benchmark our method against previous methods on the Levir-QCG dataset. Compared to previous models, the MGDiff achieves an approximately +40% improvement in counting accuracy while maintaining higher visual fidelity and strong zero-shot generalization. To the best of our knowledge, this is the first work to research accurate object quantity control in remote sensing text-to-image generation. The dataset and code will be publicly available at https://github.com/YZPioneer/MGDiff.
Zhiping Yu, Chuyu Zhong, Zhengxia Zou, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.5
2025 A Late-Stage Bitemporal Feature Fusion Network for Semantic Change Detection
abstract
Semantic change detection (SCD) is an important task in geoscience and Earth observation. By producing a semantic change map for each temporal phase, both the land use land cover (LULC) categories and change information can be interpreted. Recently some multitask learning-based SCD methods have been proposed to decompose the task into semantic segmentation (SS) and binary change detection (BCD) subtasks. However, previous works comprise triple branches in an entangled manner, which may not be optimal and hard to adopt foundation models. Besides, lacking explicit refinement of bitemporal features during fusion may cause low accuracy. In this letter, we propose a novel late-stage bitemporal feature fusion network to address the issue. Specifically, we propose local–global attentional aggregation module to strengthen feature fusion, and propose local global context enhancement module to highlight pivotal semantics. Comprehensive experiments are conducted on two public datasets, including SECOND and Landsat-SCD. Quantitative and qualitative results show that our proposed model achieves new state-of-the-art performance on both datasets.
Chenyao Zhou, Haotian Zhang 0010, Zhengxia Zou, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.5
2025 Diffusion Models for Imperceptible and Transferable Adversarial Attack
abstract
Many existing adversarial attacks generate -norm perturbations on image RGB space. Despite some achievements in transferability and attack success rate, the crafted adversarial examples are easily perceived by human eyes. Towards visual imperceptibility, some recent works explore unrestricted attacks without -norm constraints, yet lacking transferability of attacking black-box models. In this work, we propose a novel imperceptible and transferable attack by leveraging both the generative and discriminative power of diffusion models. Specifically, instead of direct manipulation in pixel space, we craft perturbations in the latent space of diffusion models. Combined with well-designed content-preserving structures, we can generate human-insensitive perturbations embedded with semantic clues. For better transferability, we further "deceive" the diffusion model which can be viewed as an implicit recognition surrogate, by distracting its attention away from the target regions. To our knowledge, our proposed method, DiffAttack, is the first that introduces diffusion models into the adversarial attack field. Extensive experiments conducted across diverse model architectures (CNNs, Transformers, and MLPs), datasets (ImageNet, CUB-200, and Standford Cars), and defense mechanisms underscore the superiority of our attack over existing methods such as iterative attacks, GAN-based attacks, and ensemble attacks. Furthermore, we provide a comprehensive discussion on future research avenues in diffusion-based adversarial attacks, aiming to chart a course for this burgeoning field.
Jianqi Chen, Hao Chen 0045, Keyan Chen 0001, Yilan Zhang, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 MetaEarth: A Generative Foundation Model for Global-Scale Remote Sensing Image Generation
abstract
The recent advancement of generative foundational models has ushered in a new era of image generation in the realm of natural images, revolutionizing art design, entertainment, environment simulation, and beyond. Despite producing high-quality samples, existing methods are constrained to generating images of scenes at a limited scale. In this paper, we present MetaEarth - a generative foundation model that breaks the barrier by scaling image generation to a global level, exploring the creation of worldwide, multi-resolution, unbounded, and virtually limitless remote sensing images. In MetaEarth, we propose a resolution-guided self-cascading generative framework, which enables the generating of images at any region with a wide range of geographical resolutions. To achieve unbounded and arbitrary-sized image generation, we design a novel noise sampling strategy for denoising diffusion models by analyzing the generation conditions and initial noise. To train MetaEarth, we construct a large dataset comprising multi-resolution optical remote sensing images with geographical information. Experiments have demonstrated the powerful capabilities of our method in generating global-scale images. Additionally, the MetaEarth serves as a data engine that can provide high-quality and rich training data for downstream tasks. Our model opens up new possibilities for constructing generative world models by simulating Earth's visuals from an innovative overhead perspective.
Zhiping Yu, Liqin Liu, Zhenwei Shi 0001, Zhengxia Zou
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 SeG-SR: Integrating Semantic Knowledge Into Remote Sensing Image Super-Resolution via Vision-Language Model
abstract
High-resolution (HR) remote sensing imagery plays a vital role in a wide range of applications, including urban planning and environmental monitoring. However, due to limitations in sensors and data transmission links, the images acquired in practice often suffer from resolution degradation. Remote Sensing Image Super-Resolution (RSISR) aims to reconstruct HR images from low-resolution (LR) inputs, providing a cost-effective and efficient alternative to direct HR image acquisition. Existing RSISR methods primarily focus on low-level characteristics in pixel space, while neglecting the high-level understanding of remote sensing scenes. This may lead to semantically inconsistent artifacts in the reconstructed results. Motivated by this observation, our work aims to explore the role of high-level semantic knowledge in improving RSISR performance. We propose a Semantic-Guided Super-Resolution framework, SeG-SR, which leverages Vision-Language Models (VLMs) to extract semantic knowledge from input images and uses it to guide the super resolution (SR) process. Specifically, we first design a Semantic Feature Extraction Module (SFEM) that utilizes a pretrained VLM to extract semantic knowledge from remote sensing images. Next, we propose a Semantic Localization Module (SLM), which derives a series of semantic guidance from the extracted semantic knowledge. Finally, we develop a Learnable Modulation Module (LMM) that uses semantic guidance to modulate the features extracted by the SR network, effectively incorporating high-level scene understanding into the SR pipeline. We validate the effectiveness and generalizability of SeG-SR through extensive experiments: SeG-SR achieves state-of-the-art performance on three datasets, and consistently improves performance across various SR architectures. Notably, for the ×4 SR task on the UCMerced dataset, it attained a PSNR of 29.3042 dB and an SSIM of 0.7961. Codes can be found at https://github.com/Mr-Bamboo/SeG-SR.
Bowen Chen 0002, Keyan Chen 0001, Mohan Yang, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 TriDF: Triplane-Accelerated Density Fields for Few-Shot Remote Sensing Novel View Synthesis
abstract
Remote sensing novel view synthesis (NVS) offers significant potential for 3D interpretation of remote sensing scenes, with important applications in urban planning and environmental monitoring. However, remote sensing scenes frequently lack sufficient multi-view images due to acquisition constraints. While existing NVS methods tend to overfit when processing limited input views, advanced few-shot NVS methods are computationally intensive and perform sub-optimally in remote sensing scenes. This paper presents TriDF, an efficient hybrid 3D representation for fast remote sensing NVS from as few as 3 input views. Our approach decouples color and volume density information, modeling them independently to reduce the computational burden on implicit radiance fields and accelerate reconstruction. We explore the potential of the triplane representation in few-shot NVS tasks by mapping high-frequency color information onto this compact structure, and the direct optimization of feature planes significantly speeds up convergence. Volume density is modeled as continuous density fields, incorporating reference features from neighboring views through image-based rendering to compensate for limited input data. Additionally, we introduce depth-guided optimization based on point clouds, which effectively mitigates the overfitting problem in few-shot NVS. Comprehensive experiments across multiple remote sensing scenes demonstrate that our hybrid representation achieves a 30× speed increase compared to NeRF-based methods, while simultaneously improving rendering quality metrics over advanced few-shot methods (7.4% increase in PSNR and 3.4% in SSIM). The code is publicly available at https://github.com/kanehub/TriDF.
Jiaming Kang, Keyan Chen 0001, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Exploring Fine-Grained Image-Text Alignment for Referring Remote Sensing Image Segmentation
abstract
Given a language expression, referring remote sensing image segmentation (RRSIS) aims to identify ground objects and assign pixelwise labels within the imagery. One of the key challenges for this task is to capture discriminative multimodal features via image-text alignment. However, the existing RRSIS methods use one vanilla and coarse alignment, where the language expression is directly extracted to be fused with the visual features. In this article, we argue that a “fine-grained image-text alignment” can improve the extraction of multimodal information. To this point, we propose a new RRSIS method to fully exploit the visual and linguistic representations. Specifically, the original referring expression is regarded as context text, which is further decoupled into the ground object and spatial position texts. The proposed fine-grained image-text alignment module (FIAM) would simultaneously leverage the features of the input image and the corresponding texts, obtaining better discriminative multimodal representation. Meanwhile, to handle the various scales of ground objects in remote sensing, we introduce a text-aware multiscale enhancement module (TMEM) to adaptively perform cross-scale fusion and intersections. We evaluate the effectiveness of the proposed method on two public referring remote sensing datasets including RefSegRS and RRSIS-D, and our method obtains superior performance over several state-of-the-art methods. The code will be publicly available athttps://github.com/Shaosifan/FIANet.
Sen Lei, Xinyu Xiao, Heng-Chao Li 0001, Zhenwei Shi 0001, Qing Zhu 0012
IEEE Trans. Geosci. Remote. Sens.5
2025 MarsSeg: Mars Surface Semantic Segmentation With Multilevel Extractor and Connector
abstract
The segmentation and interpretation of the Martian surface play a pivotal role in Mars exploration, providing essential data for the trajectory planning and obstacle avoidance of rovers. However, the complex topography, self-similar surface features, and the lack of extensive annotated data pose significant challenges to the high-precision semantic segmentation of the Martian surface. To address these challenges, we propose a novel encoder-decoder-based Mars segmentation network, termed MarsSeg. To facilitate a high-level semantic understanding across the multi-level feature maps, we introduce a feature enhancement module, which incorporates Multi-scale Feature Pyramid (MFP) and Strip Attention Pyramid Pooling Module (SAPPM). The MFP is specifically designed for shallow feature enhancement, thereby enabling the expression of local details and small objects. Conversely, the SAPPM is employed for deep feature enhancement, facilitating the extraction of high-level semantic category-related information. To effectively fuse features from different levels, we propose a feature fusion module, which contains Mars Polarized Self Attention (Mars-PSA) and Pixel Attention Head (PA-Head). Mars-PSA enables the fusion of multi-level information while directing the model’s attention to salient features. The PA-Head focuses on detailed information at the pixel level. Experimental results derived from the MarsSeg and AI4Mars datasets prove that the proposed MarsSeg outperforms other state-of-the-art methods in segmentation performance, validating the efficacy of each proposed component.
Keyan Chen 0001, Gengju Tian, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 RSBEV-Mamba: 3-D BEV Sequence Modeling for Multiview Remote Sensing Scene Segmentation
abstract
Multiview collaborative perception has been demonstrated to be highly effective in extracting 3-D information from remote sensing scenes by remote sensing bird’s-eye-view (RSBEV). However, inherent depth uncertainty in purely visual methods limits view fusion accuracy, and high computational complexity makes it challenging to model long sequences efficiently. To address these issues, we reformulate the BEV segmentation problem as a 3-D sequence modeling task and propose RSBEV-Mamba, a novel framework comprising a 3-D BEV module, a 3-D VMamba module, and a dense BEV contrastive learning module. The 3-D BEV module projects multiview 2-D image features into 3-D world coordinates, thus establishing a foundation for accurate spatial representation. The 3-D VMamba module, based on state-space models (SSMs), optimizes the processing of densely projected features with linear computational complexity in global 3-D spatial modeling. It incorporates a 3-D selective scanning strategy (SS3D) block with 16 scanning strategies, transforming previously ignored projections at different heights into valid 3-D sequences and enriching the contextual depth and precision of BEV encoding. By employing a contrastive learning strategy with the CLIP model, we align BEV and ground truth (GT) features within the same dimensional framework, ensuring spatial integrity after side-view projection. Our approach achieves a 4% improvement mIoU, thus reaching a score of 0.7368 on LEVIR-MDS and surpassing previous state-of-the-art methods. This establishes the 3-D VMamba module as a general model for 3-D perception tasks and sets a new benchmark in remote sensing technology.
Baihong Lin, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Topographic Informed Kolmogorov-Arnold Neural Interpolator for Downscaling and Correcting Meteorological Fields From In Situ Observations
abstract
Obtaining accurate weather forecasts at station locations is a critical challenge due to systematic biases arising from the mismatch between multi-scale, continuous atmospheric characteristic and their discrete, gridded representations. Previous works have primarily focused on modeling gridded meteorological data, inherently neglecting the off-grid, continuous nature of atmospheric states and leaving such biases unresolved. To address this, we propose theKolmogorov–Arnold Neural Interpolator(KANI), a novel framework that redefines meteorological field representation as continuous neural functions derived from discretized grids. Grounded in the Kolmogorov–Arnold theorem, KANI captures the inherent continuity of atmospheric states and leverages sparse in-situ observations to correct these biases systematically. Furthermore, KANI introduces an innovativezero-shotdownscaling capability, guided by high-resolution topographic textures without requiring high-resolution meteorological fields for supervision. Experimental results across three sub-regions of the continental United States indicate that KANI achieves an accuracy improvement of 40.28% for temperature and 67.41% for wind speed, highlighting its significant improvement over traditional interpolation methods. This enables continuous neural representation of meteorological variables through neural networks, transcending the limitations of conventional grid-based representations.
Hao Chen 0045, Lei Bai 0001, Wenyuan Li 0002, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 Infrared Small Target Detection Based on Prior Guided Dense Nested Network
abstract
Infrared small target detection (IRSTD) has been widely applied and developed in military and civilian fields, playing a vital role. Despite the extensive research foundation of traditional manual feature-based methods, they are still constrained by the inherent problem of infrared small targets lacking prior features. In recent years, the advancement of deep learning methods has enriched the research landscape in this field, yet they are still constrained by the imbalance of positive and negative samples between the target and the background. To address these issues, we propose a novel prior guided dense nested network (PGDN-Net), which ingeniously integrates traditional manual features with a deep learning network model. First, three prior features are extracted, including the high-order Riesz transform feature, the compactness and heterogeneity feature (CH), and the corner feature of the structure tensor (ST). Then, these features are input into a dense nested network for guidance, supported by a two-orientation attention aggregation module and a channel and spatial attention module. Different features play their respective guiding roles in different depths of the network. Through multiple attention mechanisms and feature fusion operations on the interested target area, the extraction and preservation of target features can be improved, while easily removing irrelevant backgrounds. Experiments on public datasets demonstrate the effectiveness and progressiveness of our PGDN-Net. Compared with other state-of-the-art methods, it achieves better performance in background suppression, target enhancement, probability of detection, and false alarm rate. In addition, the PGDN-Net model can effectively maintain and restore the original shape of the target while performing robust detection, which is beneficial for subsequent fine-grained recognition tasks.
Chang Liu 0090, Xuedong Song, Dianyu Yu, Linwei Qiu, Fengying Xie, Yue Zi, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.7
2025 Unsupervised Domain Adaptation for VHR Urban Scene Segmentation via Prompted Foundation Model-Based Hybrid Training Joint-Optimized Network
abstract
Unsupervised Domain Adaptation for Remote Sensing Semantic Segmentation (UDA-RSSeg) is to adapt a model trained on the source domain data to the target domain samples, thereby minimizing the need for annotated data across diverse remote sensing scenes. In urban planning and monitoring, the task of UDA-RSSeg on Very-High-Resolution (VHR) images has garnered significant research interest. While recent deep learning techniques have demonstrated huge success in tackling the UDA-RSSeg task for VHR urban scenes, a persistent challenge in addressing the domain shift issue remains. Specifically, there are two primary problems: (1) severe inconsistencies in feature representation across diverse domains, characterized by notably differing data distributions, and (2) the domain gap problem due to the representation bias of the source domain patterns when translating features to predictive logits. To solve these problems, we propose a prompted foundation model based hybrid training joint-optimized network (PFM-JONet) for UDA-RSSeg on VHR urban scene. Our approach integrates the notable “Segment Anything Model” (SAM) as prompted foundation model to leverage its robust generalized representation capabilities, thereby alleviating feature inconsistencies. Based on the feature extracted by SAM-Encoder, we introduce a mapping decoder designed to convert SAM-Encoder features into predictive logits. Additionally, a prompted segmentor is employed to generate class-agnostic maps, which guide the mapping decoder’s feature representations. To efficiently optimize the entire network in an end-to-end manner, we design a hybrid training scheme that integrates feature-level and logits-level adversarial training strategies alongside a self-training mechanism. This scheme enhances the model from diverse, compatible perspectives. To evaluate the performance of our proposed PFM-JONet, we conduct extensive experiments on urban scene benchmark datasets, including ISPRS (Potsdam/Vaihingen) and CITY-OSM (Paris/Chicago). On ISPRS dataset, PFM-JONet surpasses previous SOTA methods by 1.60% in mean IoU value across four adaptation tasks. For CITY-OSM’s adaptation task, it outperforms SOTA by 4.84% in mean IoU value. These results demonstrate the effectiveness of our method. Furthermore, visualization and analysis reinforce the method’s interpretability. The code of this paper is available at https://github.com/CV-ShuchangLyu/PFM-JONet.
Shuchang Lyu, Qi Zhao 0037, Yaxuan Sun, Yiwei He, Guangbiao Wang, Jinchang Ren, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.8
2025 M3-CR: Multiscale Multibranch Mamba for SAR-Assisted Optical Image Thick Cloud Removal
abstract
SAR-assisted thick cloud removal from optical remote sensing images has long been a challenging task. Current mainstream methods face challenges in achieving an effective global receptive field, fully utilizing multi-scale features, and deeply integrating features from both modalities. To overcome these limitations, we propose the Multi-scale Multi-branch Mamba model(M3-CR) for SAR-assisted thick cloud removal. Specifically, we integrate the Mamba model into the task of SAR-assisted cloud removal, effectively modeling global dependencies within the images. Concurrently, a Multi-scale Multi-branch structure is introduced to extract and integrate multi-scale information, and in combination with a convolutional branch to fully exploit the global and local geographic proximities inherent in remote sensing images. Furthermore, we present a novel feature fusion module leveraging the Modal-Traversing 2D Selective Scan(MTSS2D) to enable deep interaction and integration of features from optical and SAR images. The experimental results on two benchmark databases show that the M3-CR achieves superior performance compared to state-of-theart cloud removal approaches, while requiring fewer parameters and reduced FLOPs. The code for M3-CR will be made publicly available at https://github.com/LinpengPan/M3CR.
Linpeng Pan, Xuedong Song, Fengying Xie, Xiaozhe Zhang, Haolin Ji, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 Efficient Semantic Splatting for Remote Sensing Multiview Segmentation
abstract
Remote sensing multi-view image segmentation is essential for achieving accurate and consistent stereoscopic perception of target scenes. This task involves processing RGB images from multiple viewpoints to generate high-accuracy, view-consistent semantic segmentation across all views. Traditional training-based methods struggle with maintaining cross-view consistency, while optimization-driven approaches using implicit neural networks improve view consistency but suffer from slow parameter optimization and inference. To overcome these limitations, we propose a novel Gaussian Splatting-based semantic segmentation framework. Our method efficiently projects the color attributes and semantic features of 3D Gaussians onto the image plane, enabling the simultaneous generation of both RGB images and segmentation outputs. By leveraging explicit spatial structures and a splatting rendering strategy, our approach significantly enhances optimization efficiency and rendering speed. Additionally, we incorporate SAM2 to generate pseudo-labels for boundary regions, addressing the lack of supervision in sparsely labeled views (e.g., 3%). To further enforce cross-view consistency and feature coherence of 3D Gaussians, we introduce a two-level aggregation loss that operates at both the 2D feature map and 3D spatial levels. Extensive experiments across nine datasets demonstrate the superiority of our method, achieving competitive segmentation quality with limited supervisory views. Notably, our approach reduces rendering (inference) times by 90%, while improving the average mIoU by up to 3.5%.
Zipeng Qi, Hao Chen 0045, Haotian Zhang 0010, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Radiation-Tolerant Unsupervised Deep Image Stitching for Remote Sensing
Linwei Qiu, Fengying Xie, Chang Liu 0090, Xiaoling Che, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 High-Precision Wave-Parameter Perception via Spatiotemporal Coupling of Sea-Clutter Imagery and Ship-Motion Responses
abstract
Accurate detection of near-wave field parameters is crucial for safe navigation and efficient offshore operations. However, estimating precise wave parameter remains challenging due to strong nonlinear wave dynamics and inherent measurement uncertainties. Most of existing deep learning methods relying on single modality data face limitations in insufficient accuracy. Motivated by advances in multimodal fusion algorithms, we propose a novel maritime multimodal fusion inversion model, MR-FuNet. The proposed model integrates a Convolutional Neural Network (CNN) and Bidirectional Long Short-Term Memory (BiLSTM) in parallel, enabling effective fusion of spatial information from X-band radar sea clutter images and temporal patterns from ship motion data at the feature level. A multi-dimensional attention strategy was employed to enable the model to dynamically calibrate and integrate heterogeneous information across modalities and feature domains which can significantly improve the accuracy of inversion for significant wave height and characteristic wave period. To address the lack of comprehensive and high-quality public datasets in this research area, this study constructs and releases a large-scale and multimodal datasetRadar Images and Ship Motion Dataset(RSD). RSD is a large-scale multimodal dataset generated via numerical simulation and covers 99 representative sea states. It provides a valuable benchmark for future research. Extensive experiments validate the effectiveness of the proposed model, demonstrating substantial performance gains across various metrics. Compared to traditional Artificial Neural Networks (ANN), the proposed multimodal fusion model achieves notable reductions in RMSE by 61.6% and 62.8% for the characteristic period and significant wave height inversion tasks, respectively. Dataset is available at https://github.com/felixfelixXu/Radar-Images-and-Ship-Motion-Dataset.
Guangbiao Wang, Zihang Xu, Limin Huang, Shuchang Lyu, Jingjun Li, Linzhou Tang, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.8
2025 CDMamba: Incorporating Local Clues Into Mamba for Remote Sensing Image Binary Change Detection
abstract
Recently, the Mamba architecture based on state-space models has demonstrated remarkable performance in a series of natural language processing tasks and has been rapidly applied to remote sensing change detection (CD) tasks. However, most methods enhance the global receptive field by directly modifying the scanning mode of Mamba, neglecting the crucial role that local information plays in dense prediction tasks (e.g., binary CD). In this article, we propose a model called CDMamba, which effectively combines global and local features for handling binary CD tasks. Specifically, the scaled residual ConvMamba (SRCM) block is proposed to utilize the ability of Mamba to extract global features and convolution to enhance the local details, to alleviate the issue that current Mamba-based methods lack detailed clues and are difficult to achieve fine detection in dense prediction tasks. Furthermore, considering the characteristics of bi-temporal feature interaction required for CD, the adaptive global–local guided fusion (AGLGF) block is proposed to dynamically facilitate the bi-temporal interaction guided by other temporal global/local features. Our intuition is that more discriminative change features can be acquired with the guidance of other temporal features. Extensive experiments on five datasets demonstrate that our proposed CDMamba is comparable to the current methods (such as the F1/intersection over union (IoU) scores are improved by 2.10%/3.00%, 2.44%/2.91%, on LEVIR+CD and CLCD, respectively). Our code is open-sourced athttps://github.com/zmoka-zht/CDMamba.
Haotian Zhang 0010, Keyan Chen 0001, Hao Chen 0045, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 FoBa: A Foreground-Background Co-Guided Method and New Benchmark for Remote Sensing Semantic Change Detection
abstract
Despite the remarkable progress achieved in remote sensing semantic change detection (SCD), two major challenges remain. At the data level, existing SCD datasets suffer from limited change categories, insufficient change types, and a lack of fine-grained class definitions, making them inadequate to fully support practical applications. At the methodological level, most current approaches underutilize change information, typically treating it as a post-processing step to enhance spatial consistency, which constrains further improvements in model performance. To address these issues, we construct a new benchmark for remote sensing SCD, LevirSCD. Focused on the Beijing area, the dataset covers 16 change categories and 210 specific change types, with more fine-grained class definitions (e.g., roads are divided into unpaved and paved roads). Furthermore, we propose a foreground-background co-guided SCD (FoBa) method, which leverages foregrounds that focus on regions of interest and backgrounds enriched with contextual information to guide the model collaboratively, thereby alleviating semantic ambiguity while enhancing its ability to detect subtle changes. Considering the requirements of bi-temporal interaction and spatial consistency in SCD, we introduce a gated interaction fusion (GIF) module along with a simple consistency loss to further enhance the model’s detection performance. Extensive experiments on three datasets (SECOND, JL1, and the proposed LevirSCD) demonstrate that FoBa achieves competitive results compared to current SOTA methods, with improvements of 1.48%, 3.61%, and 2.81% in the SeK metric, respectively. Our code and dataset are available at https://github.com/zmoka-zht/FoBa.
Haotian Zhang 0010, Keyan Chen 0001, Hao Chen 0045, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 FIE-Net: Foreground Instance Enhancement Network for Domain Adaptation Object Detection in Remote Sensing Imagery
abstract
Domain adaptation methods can mitigate performance degradation in remote sensing image object detection that arises from inter-domain differences. However, current approaches often overlook the focused attention on foreground features, and the differences between foreground and background in remote sensing images implicitly diminishes the characteristics of the foreground. To address this challenge, we propose a foreground instance enhancement network (FIE-Net) to balance the differences between foreground and background in remote sensing images, while enhancing the alignment and application of foreground features. The FIE-Net cooperates the foregroundfocused multi-granularity feature alignment (FMA) module with the label filtering and application (LFA) module, progressively focusing on the salient foreground features during the processes of feature alignment and label application. Specifically, FMA directs feature alignment towards the foreground focus during the multi-granularity feature alignment process, through foregroundfocus perception attention and instance-centered emphasis approach. LFA balances the foreground and background difference information contained in labels, through the hybrid threshold label filtering method and the progressive label switching strategy. The experimental results indicate superior performance and generalization capabilities of our proposed FIE-Net in multiple remote sensing adaptation scenarios. Code is released at https://github.com/Lab-PANbin/.
Jun Zhang 0050, Xupeng Zhang, Bin Pan, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 DiffPR-Net: Few-Shot Remote Sensing Scene Classification Based on Generative Diffusion and Prototype Rectified Model
abstract
Few-shot remote sensing scene classification (FSRSSC) aims to identify unseen scene classes from limited labeled samples, facing the challenge of accurately modeling data distribution and preserving image details in complex backgrounds with high intraclass variance and interclass similarity. To address this challenge, we propose a novel Diffusion Prototype Rectified Network (DiffPR-Net), which is comprised of three core modules: diffusion augmentation (DA), dual attention fusion module (DAFM) and prototype rectified module (PRM). The DA is constructed to generate high-quality remote sensing images with the objective of augmenting the training dataset. Besides, the DAFM facilitates the model to focus discriminative regions by transmitting highly fused image detail features from higher to lower layers. What’s more, the PRM addresses prototype deviation by adaptively assigning temporary labels to unlabeled data based on prediction confidence, thereby correcting the initial prototypes. Experiments indicate that our proposed method is highly promising, achieving competitive or state-of-the-art classification performance while addressing the scarcity of remotely sensed data and enhancing focus on discriminative regions.
Jiaxin Han, Bin Pan, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Physical Adversarial Camouflage Generation in Optical Remote Sensing Images
abstract
Physical adversarial examples in optical remote sensing have garnered significant attention in recent years due to their practicality and high adversarial threat potential. However, existing methods focus on position-fixed adversarial patches, neglecting tailored considerations for the domain-specific texture patterns and mobility required by aerial platforms. To address the issues above, we proposed a novel method of physical adversarial camouflage generation for the first time in optical remote sensing, which paints adversarial camouflage with specialized textures onto the targets to escape detection from DNN-based models. In pursuit of achieving a synthesis of visual harmony and adversarial attack potency, we propose a "latent variable-based" adversarial camouflage generation approach, in which we introduce a texture generator controlled by a group of latent variables to generate camouflage patterns with adversarial properties. By employing this idea, we can constrain the searching domain for adversarial examples to the domain characterized by camouflage exhibiting textures with high visual harmony, and easily focus on finding the most threatening ones during the optimization. We chose airplanes as the object of interest and object detection as the typical reconnaissance method in experiments. Our method achieved high attack success rates (ASRs) against a majority of existing detection models. Comparison with existing pixel-level optimization methods confirmed that the integration of a dedicated generator helps solve the trade-off dilemma between visual harmony and adversarial potency. Real-world experiments involving targets painted by our developed adversarial camouflage confirmed the adversarial attack potency and practicality, with a more than 50% increase on average in the ASRs compared to the conventional camouflage.
Zhenbang Peng, Jianqi Chen, Zhenwei Shi 0001, Zhengxia Zou
IEEE Trans. Inf. Forensics Secur.3
2025 HUNTNet: Homomorphic Unified Nexus Topology for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) is challenging for both human and computer vision, as targets often blend into the background by sharing similar color, texture, or shape. While many feature enhancement techniques exist, single-view methods tend to overemphasize certain Recognizing that camouflaged objects exhibit different concealment strategies under varying observational perspectives, we propose HUNTNet, a network that establishes a dynamic detection mechanism to decouple target features from RGB images and perform topological decamouflage across multiple homomorphic feature spaces through a unified feature focusing architecture. We adopt PVTv2 as the backbone to extract multi-perspective spatial features. Detail representation is enhanced via a feature module that integrates Dual-Channel Recursive (DCR), Wavelet-Gabor Transform (WGT), and Anisotropic Gradient Responding (AGR), which together improve boundary discrimination and edge contour detection. To further boost performance, the Simplicial Feature Integration (SFI) module recursively fuses multi-layer features, enabling high-resolution focus on target regions. Experiments show that HUNTNet surpasses state-of-the-art methods in both accuracy and generalization, offering a robust solution for COD and improving segmentation in complex scenes. Our code is available at https://github.com/HaolinJi817/HUNTNet.
Haolin Ji, Fengying Xie, Linpeng Pan, Yushan Zheng, Zhenwei Shi 0001
IEEE Trans. Image Process.5
2025 Transformer for Multitemporal Hyperspectral Image Unmixing
abstract
Multitemporal hyperspectral image unmixing (MTHU) holds significant importance in monitoring and analyzing the dynamic changes of surface. However, compared to single-temporal unmixing, the multitemporal approach demands comprehensive consideration of information across different phases, rendering it a greater challenge. To address this challenge, we propose the Multitemporal Hyperspectral Image Unmixing Transformer (MUFormer), an end-to-end unsupervised deep learning model. To effectively perform multitemporal hyperspectral image unmixing, we introduce two key modules: the Global Awareness Module (GAM) and the Change Enhancement Module (CEM). The GAM computes self-attention across all phases, facilitating global weight allocation. On the other hand, the CEM dynamically learns local temporal changes by capturing differences between adjacent feature maps. The integration of these modules enables the effective capture of multitemporal semantic information related to endmember and abundance changes, significantly improving the performance of multitemporal hyperspectral image unmixing. We conducted experiments on one real dataset and two synthetic datasets, demonstrating that our model significantly enhances the effect of multitemporal hyperspectral image unmixing.
Qiankun Dong, Xueshuo Xie, Tao Li 0022, Zhenwei Shi 0001
IEEE Trans. Image Process.6
2025 Jointly Optimizing the Energy and Time for Multi-UAV 3-D Coverage of Terrestrial Regions
abstract
Multi-rotor unmanned aerial vehicles (UAVs) have been widely employed in various sensing tasks, e.g., environmental monitoring and disaster rescuing, many of which often require full coverage of terrestrial regions by UAVs. Efforts have been devoted to minimizing one of two objectives, i.e., energy consumptions and time costs of UAVs fulfilling such tasks, whereas it is still challenging to jointly optimize both objectives due to their complicated interdependent relationship. Therefore, this paper deals with the tasks of sensing terrestrial regions with multiple UAVs, and focuses on the three-dimensional (3-D) coverage problem by formulating a multi-objective optimization problem of jointly minimizing both objectives. Specifically, in order to optimize energy consumption effectively, an advanced closed-form energy consumption model for multi-rotor UAVs is developed based on a rigorous theoretical analysis by introducing the influences of torque and acceleration, which are often ignored by existing heuristic models. Moreover, considering the NP-hardness of the problem, an innovative swarm intelligence optimization framework is established by leveraging a multitasking learning pattern to exploit cross-task knowledge transfer and adopting an improved multi-objective salp swarm algorithm. Therein, two novel operators, i.e., a variable characteristic-guided hybrid solution initialization operator and a large-scale search-space-oriented multi-mechanism solution update operator, are designed to handle continuous, discrete and even high-dimensional variables involved. Real-world experiments validate the proposed energy model due to the reduction of power consumption estimation error by up to 59% compared to baselines, and besides, extensive simulations demonstrate that the proposed algorithm significantly outperforms the benchmarks in terms of both energy consumptions and time costs.
Baoqi Huang, Bing Jia, Lifei Hao, Zhenwei Shi 0001
IEEE Trans. Mob. Comput.5
2025 Zero-Shot Image Harmonization With Generative Model Prior
abstract
We propose a zero-shot approach to image harmonization, aiming to overcome the reliance on large amounts of synthetic composite images in existing methods. These methods, while showing promising results, involve significant training expenses and often struggle with generalization to unseen images. To this end, we introduce a fully modularized framework inspired by human behavior. Leveraging the reasoning capabilities of recent foundation models in language and vision, our approach comprises three main stages. Initially, we employ a pretrained vision-language model (VLM) to generate descriptions for the composite image. Subsequently, these descriptions guide the foreground harmonization direction of a text-to-image generative model (T2I). We refine text embeddings for enhanced representation of imaging conditions and employ self-attention and edge maps for structure preservation. Following each harmonization iteration, an evaluator determines whether to conclude or modify the harmonization direction. The resulting framework, mirroring human behavior, achieves harmonious results without the need for extensive training. We present compelling visual results across diverse scenes and objects, along with quantitative comparisons validating the effectiveness of our approach.
Jianqi Chen, Yilan Zhang, Zhengxia Zou, Keyan Chen 0001, Zhenwei Shi 0001
IEEE Trans. Multim.5
2025 CR-Famba: A Frequency-Domain Assisted Mamba for Thin Cloud Removal in Optical Remote Sensing Imagery
abstract
Optical remote sensing images are inevitably affected by cloud cover. To remove clouds from optical remote sensing images, a series of deep learning-based thin cloud removal methods have been developed. However, these methods have not explored the long-range modeling ability of state space models in optical remote sensing image thin cloud removal. In this paper, we propose a frequency-domain assisted Mamba for thin cloud removal, which is called CR-Famba. In CR-Famba, to better extract global and local features of images, we design a frequency-domain assisted state space layer (FDA-SSL). The FDA-SSL consists of two core components: residual state space block (RSSB) and frequency domain detail enhancement block (FDDEB). The RSSB utilizes the visual state space module (VSSM) to extract long-range dependencies of images from a spatial perspective while adding convolutional layers to overcome local pixel forgetting. Due to the rich detailed information of remote sensing images, we present FDDEB equipped with discrete wavelet transform (DWT) to supplement the extracted local information from the frequency domain perspective. We conduct experiments on different types of cloud-containing datasets, and the results show that our method can recover images with clearer texture details compared to other methods.
Jiao Liu 0003, Bin Pan, Zhenwei Shi 0001
IEEE Trans. Multim.3
2024 Time Travelling Pixels: Bitemporal Features Integration with Foundation Model for Remote Sensing Image Change Detection
abstract
Change detection, a prominent research area in remote sensing, is pivotal in observing and analyzing surface transformations. Despite significant advancements achieved through deep learning-based methods, executing high-precision change detection in spatiotemporally complex remote sensing scenarios still presents a substantial challenge. The recent emergence of foundation models, with their powerful universality and generalization capabilities, offers potential solutions. However, bridging the gap of data and tasks remains a significant obstacle. In this paper, we introduce Time Travelling Pixels (TTP), a novel approach that integrates the latent knowledge of the SAM foundation model into change detection. TTP can effectively address the domain shift in general knowledge transfer and the challenge of expressing homogeneous and heterogeneous characteristics of multi-temporal images. The state-of-the-art results obtained on the LEVIR-CD underscore the efficacy of the TTP. The code has been made publicly available at https://github.com/KyanChen/TTP.
Keyan Chen 0001, Chengyang Liu, Wenyuan Li 0002, Hao Chen 0045, Haotian Zhang 0010, Zhengxia Zou, Zhenwei Shi 0001
IGARSS8
2024 Learning to Detect Cloud and Snow in Remote Sensing Images from Noisy Labels
abstract
Detecting clouds and snow in remote sensing images is an essential preprocessing task for remote sensing imagery. Previous works draw inspiration from semantic segmentation models in computer vision, with most research focusing on improving model architectures to enhance detection performance. However, unlike natural images, the complexity of scenes and the diversity of cloud types in remote sensing images result in many inaccurate labels in cloud and snow detection datasets, introducing unnecessary noises into the training and testing processes. By constructing a new dataset and proposing a novel training strategy with the curriculum learning paradigm, we guide the model in reducing overfitting to noisy labels. Additionally, we design a more appropriate model performance evaluation method, that alleviates the performance assessment bias caused by noisy labels. By conducting experiments on models with UNet and Segformer, we have validated the effectiveness of our proposed method. This paper is the first to consider the impact of label noise on the detection of clouds and snow in remote sensing images.
Hao Chen 0045, Wenyuan Li 0002, Keyan Chen 0001, Zipeng Qi, Zhengxia Zou, Zhenwei Shi 0001
IGARSS8
2024 Pixel-Level Change Detection Pseudo-Label Learning For Remote Sensing Change Captioning
abstract
The existing Remote Sensing Image Change Captioning (RSICC) methods perform well in simple scenes but exhibit poorer performance in complex scenes. This limitation is primarily attributed to the model’s constrained visual ability to distinguish and locate changes. Acknowledging the inherent correlation between change detection (CD) and RSICC tasks, we believe pixel-level CD is significant for describing the differences between images through language. Regrettably, the current RSICC dataset lacks readily available pixel-level CD labels. To address this deficiency, we leverage a model trained on existing CD datasets to derive CD pseudo-labels. We propose an innovative network with an auxiliary CD branch, supervised by pseudo-labels. Furthermore, a semantic fusion augment (SFA) module is proposed to fuse the feature information extracted by the CD branch, thereby facilitating the nuanced description of changes. Experiments demonstrate that our method achieves state-of-the-art performance and validate that learning pixel-level CD pseudo-labels significantly contributes to change captioning.
Keyan Chen 0001, Zipeng Qi, Haotian Zhang 0010, Zhengxia Zou, Zhenwei Shi 0001
IGARSS7
2024 Multi-View Remote Sensing Image Segmentation with Sam Priors
abstract
Multi-view segmentation in Remote Sensing (RS) seeks to segment images from diverse perspectives within a scene. Recent methods leverage 3D information extracted from an Implicit Neural Field (INF), bolstering result consistency across multiple views while using limited accounts of labels (even within 3-5 labels) to streamline labor. Nonetheless, achieving superior performance within the constraints of limited-view labels remains challenging due to inadequate scene-wide supervision and insufficient semantic features within the INF. To address these. we propose to inject the prior of the visual foundation model-Segment Anything(SAM), to the INF to obtain better results under the limited number of training data. Specifically, we contrast SAM features between testing and training views to derive pseudo labels for each testing view, augmenting scene-wide labeling information. Subsequently, we introduce SAM features via a transformer into the INF of the scene, supplementing the semantic information. The experimental results demonstrate that our method outperforms the mainstream method, confirming the efficacy of SAM as a supplement to the INF for this task.
Zipeng Qi, Hao Chen 0045, Yongchang Wu, Zhengxia Zou, Zhenwei Shi 0001
IGARSS7
2024 Residual Group Enhanced GAN for Remote Sensing Image Cloud Shadow Removal
abstract
Cloud shadow is an unignorable factor affecting the quality of remote sensing images, but there are few researches specifically focusing on cloud shadow removal and its impact on downstream remote sensing tasks. In this paper, we propose a residual group enhanced generative adversarial network (RGE-GAN) for cloud shadow removal. We design an encoder-decoder with residual group enhancement (RGE) module to remove cloud shadows from remote sensing images. RGE module can effectively enhance the deep features extracted by encoder. We further introduce a discriminator network and employ adversarial training strategy to constrain the generator to reconstruct high-quality cloud shadow removed images conforming to the distribution of remote sensing images. The joint experiments of cloud shadow removal and building extraction on real remote sensing dataset show that our cloud shadow removal method can effectively enhance the quality of remote sensing images and improve the performance of downstream remote sensing processing tasks.
Keyan Chen 0001, Zhengxia Zou, Zhenwei Shi 0001
IGARSS5
2024 RSMamba: Remote Sensing Image Classification With State Space Model
abstract
Remote sensing image classification forms the foundation of various understanding tasks, serving a crucial function in remote sensing image interpretation. The recent advancements of Convolutional Neural Networks (CNNs) and Transformers have markedly enhanced classification accuracy. Nonetheless, remote sensing scene classification remains a significant challenge, especially given the complexity and diversity of remote sensing scenarios and the variability of spatiotemporal resolutions. The capacity for whole-image understanding can provide more precise semantic cues for scene discrimination. In this paper, we introduce RSMamba, a novel architecture for remote sensing image classification. RSMamba is based on the State Space Model (SSM) and incorporates an efficient, hardware-aware design known as the Mamba. It integrates the advantages of both a global receptive field and linear modeling complexity. To overcome the limitation of the vanilla Mamba, which can only model causal sequences and is not adaptable to two-dimensional image data, we propose a dynamic multi-path activation mechanism to augment Mamba’s capacity to model non-causal data. Notably, RSMamba maintains the inherent modeling mechanism of the vanilla Mamba, yet exhibits superior performance across multiple remote sensing image classification datasets,e.g., F1 scores of 95.25, 92.63, and 95.18 on the UC Merced, AID, and RESISC45 classification datasets respectively, exceeding those of concurrent Vim and VMamba. This indicates that RSMamba holds significant potential to function as the backbone of future visual foundation models. The code is available at https://github.com/KyanChen/RSMamba.
Keyan Chen 0001, Bowen Chen 0002, Wenyuan Li 0002, Zhengxia Zou, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.6
2024 RSCaMa: Remote Sensing Image Change Captioning With State Space Model
abstract
Remote Sensing Image Change Captioning (RSICC) aims to describe surface changes between multi-temporal remote sensing images in language, including the changed object categories, locations, and dynamics of changing objects (e.g., added or disappeared). This poses challenges to spatial and temporal modeling of bi-temporal features. Despite previous methods progressing in the spatial change perception, there are still weaknesses in joint spatial-temporal modeling. To address this, in this paper, we propose a novel RSCaMa model, which achieves efficient joint spatial-temporal modeling through multiple CaMa layers, enabling iterative refinement of bi-temporal features. To achieve efficient spatial modeling, we introduce the recently popular Mamba (a state space model) with a global receptive field and linear complexity into the RSICC task and propose the Spatial Difference-aware SSM (SD-SSM), overcoming limitations of previous CNN- and Transformer-based methods in the receptive field and computational complexity. SD-SSM enhances the model’s ability to capture spatial changes sharply. In terms of efficient temporal modeling, considering the potential correlation between the temporal scanning characteristics of Mamba and the temporality of the RSICC, we propose the Temporal-Traversing SSM (TT-SSM), which scans bi-temporal features in a temporal cross-wise manner, enhancing the model’s temporal understanding and information interaction. Experiments validate the effectiveness of the efficient joint spatial-temporal modeling and demonstrate the outstanding performance of RSCaMa and the potential of the Mamba in the RSICC task. Additionally, we systematically compare three different language decoders, including Mamba, GPT-style decoder, and Transformer decoder, providing valuable insights for future RSICC research. The code will be available at https://github.com/Chen-Yang-Liu/RSCaMa.
Keyan Chen 0001, Bowen Chen 0002, Haotian Zhang 0010, Zhengxia Zou, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.6
2024 Road Graph Extraction via Transformer and Topological Representation
abstract
Road graph extraction from remote sensing images aims at extracting topological maps composed of road vertices and edges, which has broad prospects in urban planning, traffic management and other applications. However, existing methods are easily affected by complex remote sensing scenes, and also have shortcomings such as poor continuity and slow processing speed. In this paper, we propose a novel end-to-end road extraction method named “Road2Graph”, which encodes road graphs into topological representations for prediction1. We proposed a transformer-based model to encode the deep convolutional features, and then fuse them with the output of the feature extractor to make the network pay more attention to the global multiscale road topology context. We also design an efficient topological representation that encodes attributes such as road segmentation, midpoint map, vertex map, and connection relationships with few parameters and low redundancy. The obtained topological representation can be decoded to obtain the road extraction result in graph format. We conduct experiments on two public datasets - CityScale dataset and SpaceNet dataset. The results show that our method achieves the state-of-art and improves both accuracy (TOPO-F1 +1.55% on CityScale dataset and +2.23% on SpaceNet dataset) and continuity (APLS +7.03% on CityScale dataset and +3.05% on SpaceNet dataset) compared to the other methods.
Yifan Zao, Zhengxia Zou, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.3
2024 Learning at a Glance: Towards Interpretable Data-Limited Continual Semantic Segmentation via Semantic-Invariance Modelling
abstract
Continual semantic segmentation (CSS) based on incremental learning (IL) is a great endeavour in developing human-like segmentation models. However, current CSS approaches encounter challenges in the trade-off between preserving old knowledge and learning new ones, where they still need large-scale annotated data for incremental training and lack interpretability. In this paper, we present Learning at a Glance (LAG), an efficient, robust, human-like and interpretable approach for CSS. Specifically, LAG is a simple and model-agnostic architecture, yet it achieves competitive CSS efficiency with limited incremental data. Inspired by human-like recognition patterns, we propose a semantic-invariance modelling approach via semantic features decoupling that simultaneously reconciles solid knowledge inheritance and new-term learning. Concretely, the proposed decoupling manner includes two ways, i.e., channel-wise decoupling and spatial-level neuron-relevant semantic consistency. Our approach preserves semantic-invariant knowledge as solid prototypes to alleviate catastrophic forgetting, while also constraining sample-specific contents through an asymmetric contrastive learning method to enhance model robustness during IL steps. Experimental results in multiple datasets validate the effectiveness of the proposed method. Furthermore, we introduce a novel CSS protocol that better reflects realistic data-limited CSS settings, and LAG achieves superior performance under multiple data-limited conditions.
Bo Yuan 0009, Danpei Zhao, Zhenwei Shi 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Joint Variational Inference Network for domain generalization
Jun-Zheng Chu, Bin Pan, Tianyang Shi, Zhenwei Shi 0001, Tao Li 0022
Pattern Recognit.5
2024 Dense Pixel-to-Pixel Harmonization via Continuous Image Representation
abstract
High-resolution (HR) image harmonization is of great significance in real-world applications such as image synthesis and image editing. However, due to the high memory costs, existing dense pixel-to-pixel harmonization methods are mainly focusing on processing low-resolution (LR) images. Some recent works resort to combining with color-to-color transformations but are either limited to certain resolutions or heavily depend on hand-crafted image filters. In this work, we explore leveraging the implicit neural representation (INR) and propose a novel image Harmonization method based on Implicit neural Networks (HINet), which to the best of our knowledge, is the first dense pixel-to-pixel method applicable to HR images without any hand-crafted filter design. Inspired by the Retinex theory, we decouple the MLPs into two parts to respectively capture the content and environment of composite images. A Low-Resolution Image Prior (LRIP) network is designed to alleviate the Boundary Inconsistency problem, and we also propose new designs for the training and inference process. Extensive experiments have demonstrated the effectiveness of our method compared with state-of-the-art methods. Furthermore, some interesting and practical applications of the proposed method are explored. Our code is available at https://github.com/WindVChen/INR-Harmonization.
Jianqi Chen, Yilan Zhang, Zhengxia Zou, Keyan Chen 0001, Zhenwei Shi 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation Based on Visual Foundation Model
abstract
Leveraging the extensive training data from SA-1B, the Segment Anything Model (SAM) demonstrates remarkable generalization and zero-shot capabilities. However, as a category-agnostic instance segmentation method, SAM heavily relies on prior manual guidance, including points, boxes, and coarse-grained masks. Furthermore, its performance in remote sensing image segmentation tasks remains largely unexplored and unproven. In this paper, we aim to develop an automated instance segmentation approach for remote sensing images, based on the foundational SAM model and incorporating semantic category information. Drawing inspiration from prompt learning, we propose a method to learn the generation of appropriate prompts for SAM. This enables SAM to produce semantically discernible segmentation results for remote sensing images, a concept we have termed RSPrompter. We also propose several ongoing derivatives for instance segmentation tasks, drawing on recent advancements within the SAM community, and compare their performance with RSPrompter. Extensive experimental results, derived from the WHU building, NWPU VHR-10, and SSDD datasets, validate the effectiveness of our proposed method. The code for our method is publicly available at https://kychen.me/RSPrompter.
Keyan Chen 0001, Hao Chen 0045, Haotian Zhang 0010, Wenyuan Li 0002, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.7
2024 Spectral-Cascaded Diffusion Model for Remote Sensing Image Spectral Super-Resolution
abstract
Hyperspectral remote sensing images (HSIs) have unique advantages in urban planning, precision agriculture, and ecology monitoring since they provide rich spectral information. However, hyperspectral imaging usually suffers from low spatial resolution and high cost, which limits the wide application of hyperspectral data. Spectral super-resolution provides a promising solution to acquire hyperspectral images with high spatial resolution and low cost, taking RGB images as input. Existing spectral super-resolution methods utilize neural networks following a single-shot framework, i.e., final results are obtained by one-stage spectral super-resolution, which struggles to capture and model the complex relationships between spectral bands. In this article, we propose a spectral-cascaded diffusion model (SCDM), a coarse-to-fine spectral super-resolution method based on the diffusion model. The diffusion model fits the real data distribution through stepwise denoising, which is naturally suitable for modeling rich spectral information. We cascade the diffusion model in the spectral dimension to gradually refine the spectral trends and enrich spectral information of the pixels. The cascade solves the highly ill-posed problem of spectral super-resolution step-by-step, mitigating the inaccuracies of previous single-shot approaches. To better utilize the potential of the diffusion model for spectral super-resolution, we design image condition mixture guidance (ICMG) to enhance the guidance of image conditions and progressive dynamic truncation (PDT) to limit cumulative errors in the sampling process. Experimental results demonstrate that our method achieves state-of-the-art performance in spectral super-resolution. Codes can be found athttps://github.com/Mr-Bamboo/SCDM.
Bowen Chen 0002, Liqin Liu, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Digital-to-Physical Visual Consistency Optimization for Adversarial Patch Generation in Remote Sensing Scenes
abstract
In contrast to digital image adversarial attacks, adversarial patch attacks involve physical operations that project crafted perturbations into real-world scenarios. During the digital-to-physical transition, adversarial patches inevitably undergo information distortion. Existing approaches focus on data augmentation and printer color gamut regularization to improve the generalization of adversarial patches to the physical world. However, these efforts overlook a critical issue within the adversarial patch crafting pipeline—namely, the significant disparity between the appearance of adversarial patches during the digital optimization phase and their manifestation in the physical world. This unexplored concern, termed “Digital-to-Physical Visual Inconsistency", introduces inconsistent objectives between the digital and physical realms, potentially skewing optimization directions for adversarial patches. To tackle this challenge, we propose a novel harmonization-based adversarial patch attack. Our approach involves the design of a self-supervised harmonization method, seamlessly integrated into the adversarial patch generation pipeline. This integration aligns the appearance of adversarial patches overlaid on digital images with the imaging environment of the background, ensuring a consistent optimization direction with the primary physical attack goal. We validate our method through extensive testing on the aerial object detection task. To enhance the controllability of environmental factors for method evaluation, we construct a dataset of 3D simulated scenarios using a graphics rendering engine. Extensive experiments on these scenarios demonstrate the efficacy of our approach. Our code and dataset are publicly accessible at https://github.com/WindVChen/VCO-AP.
Jianqi Chen, Yilan Zhang, Keyan Chen 0001, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Robust Haze and Thin Cloud Removal via Conditional Variational Autoencoders
abstract
Existing methods for remote sensing image dehazing and thin cloud removal treat this image restoration task as a clear pixel estimation problem, yielding a single prediction result through a deterministic pipeline. However, image restoration is a highly ill-posed problem, as the sharp pixel value corresponding to the input cannot be uniquely determined solely from the degraded image. In this paper, we present a novel algorithm for haze and thin cloud removal using Conditional Variational Autoencoders (CVAE) to generate multiple realistic restored images for each input. By sampling from the latent space to capture the pixel diversity, the proposed method mitigates the limitations arising from inaccuracies in a single estimation. In this uncertainty pipeline, we can generate a more accurate restored image based on these multiple predictions. Furthermore, we have developed a Dynamic Fusion Network (DFN) for combining multiple plausible outcomes to obtain a more accurate result. DFN dynamically predicts the kernels used for restored result generation conditioned on inputs, improving haze and thin cloud thanks to its adaptive nature. Quantitative and qualitative experiments demonstrate that the proposed method outperforms existing state-of-the-art techniques by a significant margin on dehazing and thin cloud removal benchmarks.
Haidong Ding, Fengying Xie, Linwei Qiu, Xiaozhe Zhang, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 A Reversible Generative Network for Hyperspectral Unmixing With Spectral Variability
abstract
Spectral variability is one of the challenges for hyperspectral unmixing. Recently, deep generative models are developed to describe the spectral variability, which have attracted increasing attention. However, generative unmixing methods may suffer the problems of mode collapse and image blur, which tend to generate uncontrollable endmember distribution. To address this issue, in this paper, we propose a Reversible Generative Network (Rev-Net) for hyperspectral imagery unmixing, which targets at the spectral variability challenge. Our motivation is that if the endmember distribution can be described by an explicit mathematical expression and the expression is reversible, then the generation process will be more stable. To achieve this purpose, Rev-Net mainly includes two contributions: a flow-based endmember learning module, and a theoretical proof for the reversibility of the endmember generation process. In the endmember learning module, we develop a new flow-based structure with a series of reversible transformation, so as to obtain an explicit mathematical expression for the endmember distribution. Moreover, to guarantee the existence of the explicit expression, we have theoretically proven the reversibility of the endmember learning module. Through the flow-based endmember learning module and the correspond theoretical analysis, the proposed Rev-Net can make the endmember generation process more stable and thus avoiding the problems of mode collapse and image blur. In addition, we also construct an abundance guidance module to further assist in the generation process of endmember by image reconstruction. Experimental results on real hyperspectral datasets and synthetic datasets indicate that Rev-Net has certain competitiveness. TThe codes are available at https://github.com/Lab-PANbin/Rev-Net.
Yuyou Gao, Bin Pan, Xinyu Song 0004, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 PETDet: Proposal Enhancement for Two-Stage Fine-Grained Object Detection
abstract
Fine-grained object detection (FGOD) extends object detection with the capability of fine-grained recognition. In recent two-stage FGOD methods, the region proposal serves as a crucial link between detection and fine-grained recognition. However, current methods overlook that some proposal-related procedures inherited from general detection are not equally suitable for FGOD, limiting the multitask learning from generation, representation, to utilization. In this article, we present a proposal enhancement for two-stage FGOD (PETDet) to better handle the subtasks in two-stage FGOD methods. First, an anchor-free quality-oriented proposal network (QOPN) is proposed with dynamic label assignment and attention-based decomposition to generate high-quality-oriented proposals. In addition, we present a bilinear channel fusion network (BCFN) to extract independent and discriminative features of the proposals. Furthermore, we designed a novel adaptive recognition loss (ARL) that offers guidance for the region-based convolutional neural networks (R-CNNs) head to focus on high-quality proposals. Extensive experiments validate the effectiveness of PETDet. Quantitative analysis reveals that PETDet with ResNet50 reaches state-of-the-art performance on various FGOD datasets, including FAIR1M-v1.0 (42.96 AP), FAIR1M-v2.0 (48.81 AP), MAR20 (85.91 AP), and ShipRSImageNet (74.90 AP). The proposed method also achieves superior compatibility between accuracy and inference speed. Our code and models will be released athttps://github.com/canoe-Z/PETDet.
Danpei Zhao, Bo Yuan 0009, Yue Gao 0008, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 RSBEV: Multiview Collaborative Segmentation of 3-D Remote Sensing Scenes With Bird's-Eye-View Representation
abstract
Perception of 3-D remote sensing scenes plays a crucial role in accurately recognizing and locating ground objects, as it enables a deeper understanding of complex environments by capturing scene geometry, object relationships, and occlusion patterns. Inspired by the powerful multisensor fusion capabilities in autonomous driving, we explore a new task in this article: given a set of multiview images of a 3-D remote sensing scene, we aim to obtain bird’s-eye-view (BEV) scene information under the common view area in the world coordinate system. In this work, we focus on the task of semantic segmentation to demonstrate the feasibility of our approach and introduce a BEV modeling technique tailored for remote sensing scenes, which facilitates the projection of 3-D scene details from multiple perspective views onto a BEV. We then utilize a dual-encoder structure based on the vision transformer (VIT) architecture to extract relevant spatial information using self-attention mechanisms. Within the decoder, we employ a feature pyramid network (FPN) to integrate BEV patch encoding with spatial feature residuals, enabling fine-grained segmentation results at the original input resolution. Furthermore, we curated the LEVIR-MDS multidrone segmentation dataset, comprising scenes from ten community-level areas across three continents, totaling 243k images and their corresponding annotated BEV semantic maps, amounting to approximately 500 GB. This dataset serves as a robust benchmark to assess the effectiveness and generalization capability of our proposed method. To our knowledge, this is the first semantic segmentation dataset designed specifically for collaborative multidrone applications. We further show that our method achieves a 12% improvement in mean IoU (mIoU), reaching 69.73%, compared to a pure convolutional network model.
Baihong Lin, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Deriving Accurate Surface Meteorological States at Arbitrary Locations via Observation-Guided Continuous Neural Field Modeling
abstract
Accurately retrieving surface meteorological states at arbitrary locations is of great application significance in weather forecasting and climate modeling. Since meteorological variables are typically provided as coarse-resolution gridded fields, common methods that obtain the states at a specific location directly through spatial interpolation can lead to significant accuracy deviations compared to actual observations. Traditional downscaling, the process of obtaining fixed-scale high-resolution meteorological fields from low-resolution inputs, has been proposed as a way to indirectly improve the accuracy of retrieving states at arbitrary locations by providing more detailed subgrid-scale information. However, for arbitrary locations at the station scale, their states are influenced by subgrid information, resulting in systematic biases between the downscaled results after interpolation and the actual observations at specific station locations. To address this issue, in this article, we propose a new task called station-scale downscaling, which aims to directly derive accurate meteorological states at any given station location from a coarse-resolution meteorological field. To achieve this, we propose a new downscaling model based on hypernetwork architecture, namely, HyperDS, which efficiently integrates the multiscale observational information to guide the continuous neural field modeling of the meteorological variables, enabling accurate sampling of the states at any target location. Through extensive experiments, our proposed method outperforms other specially designed baseline models on multiple surface variables. Notably, the mean squared error (mse) for wind speed and surface pressure improved by 67% and 19.5% compared with other methods, respectively.
Hao Chen 0045, Lei Bai 0001, Wenyuan Li 0002, Keyan Chen 0001, Wanli Ouyang, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.9
2024 MambaDS: Near-Surface Meteorological Field Downscaling With Topography Constrained Selective State-Space Modeling
abstract
In an era of frequent extreme weather and global warming, obtaining precise, fine-grained near-surface weather forecasts is increasingly essential for human activities. Downscaling (DS), a crucial task in meteorological forecasting and remote sensing, enables the reconstruction of high-resolution meteorological states for target regions from global-scale forecast results. Previous downscaling methods, inspired by convolutional neural network (CNN) and Transformer-based super-resolution (SR) models, lacked tailored designs for meteorology and encountered structural limitations. Notably, they failed to efficiently integrate topography, a crucial prior to the downscaling process. In this article, we address these limitations by pioneering the selective state-space model (SSM) into the meteorological field downscaling and propose a novel model called MambaDS. This model retains the advantages of Mamba in long-range dependency modeling and linear computational complexity while enhancing the learning ability of multivariate correlation. In addition, by designing an efficient topography constraint layer, this prior information can be used more efficiently than ever before. Through extensive experiments in both China mainland and the continental United States (CONUS), we validated that our proposed MambaDS achieves state-of-the-art (SOTA) results in three different types of meteorological field downscaling settings.
Hao Chen 0045, Lei Bai 0001, Wenyuan Li 0002, Wanli Ouyang, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.7
2024 Change-Agent: Toward Interactive Comprehensive Remote Sensing Change Interpretation and Analysis
abstract
Monitoring changes in the Earth’s surface is crucial for understanding natural processes and human impacts, necessitating precise and comprehensive interpretation methodologies. Remote sensing (RS) satellite imagery offers a unique perspective for monitoring these changes, leading to the emergence of RS image change interpretation (RSICI) as a significant research focus. Current RSICI technology encompasses change detection and change captioning, each with its limitations in providing comprehensive interpretation. To address this, we propose an interactive Change-Agent, which can follow user instructions to achieve comprehensive change interpretation and insightful analysis, such as change detection and change captioning, change object counting, and change cause analysis. The Change-Agent integrates a multilevel change interpretation (MCI) model as the eyes and a large language model (LLM) as the brain. The MCI model contains two branches of pixel-level change detection and semantic-level change captioning, in which the BI-temporal iterative interaction (BI3) layer is proposed to enhance the model’s discriminative feature representation capabilities. To support the training of the MCI model, we build the LEVIR-MCI dataset with a large number of change masks and captions of changes. Experiments demonstrate the state-of-the-art (SOTA) performance of the MCI model in achieving both change detection and change description simultaneously and highlight the promising application value of our Change-Agent in facilitating comprehensive interpretation of surface changes, which opens up a new avenue for intelligent RS applications. To facilitate future research, we will make our dataset and codebase publicly available athttps://github.com/Chen-Yang-Liu/Change-Agent.
Keyan Chen 0001, Haotian Zhang 0010, Zipeng Qi, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Cascaded Memory Network for Optical Remote Sensing Imagery Cloud Removal
abstract
Cloud removal is an inevitable task in optical remote sensing images, which aims at restoring high-quality images from cloud-contaminated images. In recent years, deep learning-based image cloud removal methods utilize convolution neural network to obtain clean images. However, due to the limitations of the convolution operator, these methods cannot effectively leverage the local and global information of the image. In this paper, we propose a Cascaded Memory Network (CMNet) for optical remote sensing imagery cloud removal. The CMNet recycles previously captured information to form a memory mechanism, which is composed of two cascaded components: local information memory module (LIMM) and global information auxiliary module (GIAM). The LIMM aims to obtain local spatial details information of the image via two sub-networks, and the GIAM tries to further restore the detail of the image from global perspective. In the LIMM, two sub-networks are constructed to capture the details from coarse to fine, each of which includes a continuous memory descriptor that describes local details of the image and a hierarchical memory correlation descriptor that adaptively integrates relevant features. In the GIAM, we design a swin cloud remove transformer layer and explore an adaptive normalization to cope with unevenly distributed thin clouds, and further provide theoretical proof for the existence of the required solution. Experimental results indicate that our method can remove clouds while maintaining the detailed information of the image. https://github.com/Lab-PANbin/.
Jiao Liu 0003, Bin Pan, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Infrared Small Target Detection Based on Monogenic Signal Decomposition
abstract
Robust detection of infrared small target under complex background is of great significance for infrared search and tracking applications. However, the inherent problem of limited prior features for infrared small target has always made its detection task a challenging research topic. In order to solve the problem, we propose a novel infrared small target detection method based on monogenic signal decomposition and feature expansion, which can effectively enrich and extract the potential features of the target. First, a series of local information of the original image is obtained through the monogenic signal constructed by Riesz transform. Then, various features of the small target are extracted from different local signals, including the direction feature, edge feature, and local saliency feature. Finally, the fusion of target features is completed through signal reconstruction, thereby achieving target detection. This method not only pays attention to the local salient characteristic of the target, but also supplements the consideration of other characteristics of the target, providing a new idea for small target detection. The experimental results on real infrared images show that the proposed method framework is reasonable and effective, and possesses better detection performance and good generalization compared to other state-of-the-art methods.
Chang Liu 0090, Fengying Xie, Linwei Qiu, Haolin Ji, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Remote Sensing Image Rectangling With Iterative Warping Kernel Self-Correction Transformer
abstract
Stitched remote sensing images often exhibit irregular boundaries, which can be frustrating for general users and detrimental to downstream tasks such as object detection and segmentation. However, this issue has received insufficient attention and remains unexplored within the remote sensing domain. In this study, we investigate mesh-based rectangling techniques for remote sensing images, aiming to produce rectangular outputs while preserving the original field-of-view (FoV) and avoiding the introduction of unreliable content. Observing that prior rectangling algorithms tend to generate unsatisfactory boundaries or discernible distortions, that is, under-rectangling or over-rectangling, we propose the concept of a warping kernel associated with mesh deformations to account for these phenomena. Consequently, we introduce the iterative warping kernel self-correction transformer (IWKFormer), designed to enhance warping kernel estimation and generate superior rectangular outcomes. It primarily comprises two components: a mesh feature extractor built upon the partial swin transformer block (PSTB) and a corrector module using the swin transformer block (STB). These modules collaborate to derive warping kernels implicitly. The extractor extracts latent features pertinent to mesh deformation, whereas the corrector iteratively refines the warping kernel estimation to improve the ultimate prediction. Furthermore, to bolster further research, we have constructed an aerial imagery stitching rectangling dataset (AIRD), featuring a wide array of stitching scenes. Extensive experimentation on the AIRD demonstrates that our method yields visually appealing and naturally rectangled images, achieving state-of-the-art performance. The code and data will be available athttps://github.com/yyywxk/IWKFormer.
Linwei Qiu, Fengying Xie, Chang Liu 0090, Xuedong Song, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 HiReNet: Hierarchical-Relation Network for Few-Shot Remote Sensing Image Scene Classification
abstract
Few-shot scene classification aims to develop models that can quickly adapt to new scenes with only a few labeled samples that are not present in training sets. In recent years, convolutional neural networks (CNNs) have made significant advancements in few-shot remote sensing image scene classification tasks. However, most existing approaches focus solely on utilizing high-level embeddings of remote sensing images to learn similarity relations, while neglecting intrinsic hierarchical representations that could be crucial in distinguishing scenes with substantial interclass similarities. To address this limitation, we propose a novel few-shot scene classification method for remote sensing images called hierarchical-relation network (HiReNet). This approach leverages the hierarchical features of a query sample and its corresponding support sample to learn discriminative representations. HiReNet consists of an embedding network and a relation network. The embedding network employs a Siamese architecture to extract representations, while the relation network utilizes these representations for classification. Within the relation network, we introduce a hierarchical relation learning (HRL) structure to capture the hierarchical relations among query and support samples. Additionally, to extract stronger features, we introduce a feature aggregation module that concatenates multilevel features and employs channel attention to re- weight these features. Experimental results demonstrate the superior performance of our HiReNet compared to several state-of-the-art few-shot scene classification methods.
Sen Lei, Yingbo Zhou 0001, Jialin Cheng, Guohao Liang, Zhengxia Zou, Heng-Chao Li 0001, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.8
2024 SWIN-TOD: Smooth Wasserstein Distance and Instance-Level Neighboring Enhancement for Remote Sensing Tiny Object Detection
abstract
The advancement of deep neural network has propelled the widespread application of remote sensing target detection. However, compared to natural scenes, remote sensing targets possess inherent characteristics such as weak features and small scale, leading to a significant performance gap in traditional detection methods. To address these challenges, we undertake a systematic analysis of existing approaches, focusing on two key aspects: inadequate extraction of discriminative features and inappropriate regression measurement metrics. To tackle the first issue, an instance-level neighboring enhancement network (INEN) is proposed, enhancing the network’s feature extraction capability through inter-object feature aggregation. To address the second issue, a novel metric, smooth Wasserstein loss (SWL), is devised. Building upon these principles, a new tiny object detection (TOD) network for remote sensing images is developed. Extensive experiments on AI-TOD v1/v2 and DOTA v2 remote sensing tiny target detection datasets demonstrate that our approach achieves state-of-the-art (SOTA) performance. Codes are available athttps://github.com/sevenwgb/SWIN-TOD.
Guangbiao Wang, Hongbo Zhao 0001, Shuchang Lyu, Qing Chang 0003, Wenquan Feng, Qi Zhao 0037, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.8
2024 MiSSNet: Memory-Inspired Semantic Segmentation Augmentation Network for Class-Incremental Learning in Remote Sensing Images
abstract
With remote sensing images constantly being collected rapidly, class-incremental semantic segmentation task has attracted increasing attention. However, the semantic distribution shift problem of the background class in remote sensing images, which is a case of catastrophic forgetting, continues to limit available class-incremental semantic segmentation algorithms. To address this challenge, we present a new Memory-inspired Semantic Segmentation augmentation network (MiSSNet) for class-incremental learning in remote sensing images. The MiSSNet mainly includes two modules: Local Semantic Distillation (LSD) module and Class-Specific Regularization (CSR) module. LSD is a distillation structure that employs the local semantic features in retained memory to maintain correlation between pixels throughout the training process of incremental learning. It constructs a series of pixel-level correlation matrices and implicitly adjusts the semantic distribution shift problem of the background class. CSR is a class-wise regularization term that utilizes the class-specific portion of the preserved memory to help the model keep repeating the learning of the old categories. It alleviates the background classes shift problem by generating countless pixel level instances of old classes. LSD and CSR work together to tackle the semantic distribution shift problem of background class from semantic information and class information aspects, respectively. Specially, MiSSNet only needs additional single inference process for memory extraction and storage, and the whole algorithm does not add any new training parameters. Experimental results on three semantic segmentation datasets indicate the advantage of the proposed method.
Jiajun Xie, Bin Pan, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Topology-Guided Road Graph Extraction From Remote Sensing Images
abstract
Road maps are widely used in traffic management, vehicle navigation, urban planning, and other fields. However, automatically extracting road graphs from remote sensing images is very challenging due to the interference of vegetation and buildings and the problem of data imbalance. In this article, we propose a novel end-to-end road graph extraction method for remote sensing images named “TopoRoad,” which learns vectorized representations of road maps guided by topological graphs of the roads. Our method decouples road graph extraction into the predictions of a vertex/degree map (DM), an orientation map, and a segmentation map, which are then jointly decoded to obtain the vertices and edges of the final road graphs. In the vertex branch (VB), we predict the probability map of the vertices for the subsequent vertex extraction. At the same time, to learn local topological connections, the number of connections for each vertex is predicted. For the orientation branch (OB), the connectivity between vertices is obtained by learning to predict the local extension direction of all road pixels. We conduct experiments on two public road extraction datasets—the CityScale dataset and the SpaceNet dataset. The result suggests that our method can produce accurate and continuous road graph extraction results with vectorized representations. Our method achieves state-of-the-art (SOTA) results compared to the other methods in terms of both local and global topological accuracy.
Yifan Zao, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Generating Imperceptible and Cross-Resolution Remote Sensing Adversarial Examples Based on Implicit Neural Representations
abstract
Deep neural networks (DNNs) have been widely applied in remote sensing, and the research on its adversarial attack algorithm is the key to evaluating its robustness. Current adversarial attack methods primarily prioritize maximizing the attack success rate, disregarding the imperceptibility of the generated adversarial noise to human visual perception. Moreover, research on adversarial sample transferability has mostly focused on cross-model and cross-dataset scenarios, overlooking the investigation of adversarial attacks across different resolutions, while the rarely studied cross-resolution adversarial attacks are critical for remote sensing with different resolutions. In this article, we propose a novel method for generating imperceptible adversarial samples for cross-resolution remote sensing images based on implicit neural representations (INRs). By mapping the discrete images to a continuous neural functional space, we explicitly guarantee the visual quality of adversarial samples and decouple the model input from the image resolution. To enhance the visual fidelity of the generated adversarial samples, a multiscale discriminative learning scheme is proposed for the optimization process. For cross-resolution adversarial attacks, we align with images of different resolutions and generate cross-resolution adversarial perturbation by benefiting from the natural properties of the continuous resolution of INRs. To validate the effectiveness of our method, we compare it with the existing adversarial attacking methods using four evaluation metrics. Experiments show that our method achieves the best results in terms of attack success rate, imperceptibility, and cross-resolution attack transferability. Our code will be made publicly available.
Jianqi Chen, Liqin Liu, Keyan Chen 0001, Zhenwei Shi 0001, Zhengxia Zou
IEEE Trans. Geosci. Remote. Sens.5
2024 Physical Adversarial Attacks Against Aerial Object Detection With Feature-Aligned Expandable Textures
abstract
Physical adversarial attacks in aerial object detection have gained significant attention. Existing adversarial patches exhibit subpar visual effects and encounter limitations when transitioning from digital to physical spaces, restricting applicability in real-world scenarios. To address these challenges, we propose an adversarial texture generation method based on background texture design. This method selectively covers the background environment without interfering with the target surface. We also explore the translational invariance of fully convolutional networks to decouple adversarial textures from shapes, allowing adversarial textures to be arbitrarily expanded during use. The areas where adversarial textures are placed are designated as the “detection failure zone,” rendering the detector ineffective regardless of the aircraft’s position within this zone. This significantly enhances the practicality of the adversarial texture. To improve its concealment, we align the features of the adversarial textures with those of the original image using a pretrained VGG network, ensuring a consistent style and color tone with the background environment. Additionally, we employ a discriminator to further control the visual effects of the adversarial samples, ensuring effective concealment. Furthermore, we simulate the real environment in digital space using operations like affine transformations and Gaussian blur to transfer adversarial textures seamlessly from digital to physical space. This allows for the integration of adversarial textures into real environments without compromising their effectiveness. Experimental results demonstrate the effectiveness and consistent styling of the proposed adversarial texture in real-world environments, showing robustness against environmental changes, weather conditions, and viewing angles.
Jianqi Chen, Zhenbang Peng, Yi Dang, Zhenwei Shi 0001, Zhengxia Zou
IEEE Trans. Geosci. Remote. Sens.5
2024 BiFA: Remote Sensing Image Change Detection With Bitemporal Feature Alignment
abstract
Despite the success of deep learning-based change detection methods, their existing insufficiency in temporal (channel, spatial) and multi-scale alignment have rendered them insufficient capability in mitigating external factors (illumination changes and perspective differences, etc.) arising from different imaging conditions during change detection. In this paper, a Bi-temporal Feature Alignment (BiFA) model is proposed to produce a precise change detection map in a lightweight manner by reducing the impact of irrelevant factors. Specifically, for the temporal alignment, the Bi-temporal Interaction (BI) module is proposed to realize the alignment of the bi-temporal image channel level. Our intuition is introducing the bi-temporal interaction in the feature extraction stage may benefit suppressing the interference, such as illumination changes. Simultaneously, the Alignment module based on Differential Flow Field (ADFF) is proposed to explicitly estimate the offset of the bi-temporal image and realize their spatial level alignment to mitigate the inadequate registration resulting from different perspectives. Furthermore, for the multi-scale alignment, we introduce the Implicit Neural alignment Decoder (IND) to produce more refined prediction maps achieving precise alignment of multi-scale features by learning continuous image representations in coordinate space. Our BiFA outperforms other state-of-the-art methods on six datasets (such as the F1/IoU scores are improved by 2.70%/3.91%, 2.01%/2.94% on LEVIR+-CD and SYSU-CD, respectively) and displays greater robustness in cross-resolutions change detection. Our code is available at https://github.com/zmoka-zht/BiFA.
Haotian Zhang 0010, Hao Chen 0045, Chenyao Zhou, Keyan Chen 0001, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.7
2024 Proxy and Cross-Stripes Integration Transformer for Remote Sensing Image Dehazing
abstract
Existing Transformer-based dehazing methods for remote sensing (RS) images, to avoid quadratic computation complexity with respect to the feature map size, either perform self-attention mechanisms within local windows or capture long-range dependencies in the channel dimension rather than spatial. Each of these methods has its drawbacks. To address these limitations, we propose the Proxy and Cross-Stripes Integration Transformer (PCSformer) for RS image dehazing. PCSformer introduces two innovative Transformer blocks, i.e., sliding cross-stripes Transformer block and local proxy-based global Transformer block. The former allows us to directly model long-range dependencies and capture rich contextual information for large-scale objects in RS images. The latter seeks valuable information for thick haze regions within the whole feature map, generating more consistent and realistic scene details for such regions. Both achieve a large receptive field with cost-effective computational complexity within a single Transformer block. Furthermore, we introduce a shallow deep model with a small receptive field to conduct local refinement, which can mitigate artifacts associated with a large receptive field. Finally, to facilitate the better application of dehazing models to downstream visual tasks, we contribute two large-scale datasets for RS image dehazing. Experiments indicate that the dehazing models trained on our datasets can better assist downstream visual tasks under hazy atmospheric conditions compared to the dehazing models trained on existing datasets. Quantitative and qualitative experiments demonstrate that the proposed PCSformer significantly outperforms existing state-of-the-art techniques on dehazing benchmarks, particularly excelling in the restoration of thick haze scenes. The code and datasets are available athttps://github.com/SmileShaun/PCSformer.
Xiaozhe Zhang, Fengying Xie, Haidong Ding, Shaocheng Yan, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Three-Dimensional Frequency-Domain Transform Network for Cross-Scene Hyperspectral Image Classification
abstract
Reducing interdomain discrepancies effectively enhances the performance of hyperspectral cross-scene classification tasks. However, hyperspectral single-source domain (SD) generalization methods based on mining visual representation information are significantly influenced by interdomain discrepancies. Recent research has demonstrated that frequency-domain information exhibits robust stability. Therefore, this article proposes a three-dimensional frequency domain transform network (TFTnet) for achieving hyperspectral single-SD cross-scene classification tasks. To leverage the advantageous 3-D characteristics of hyperspectral images (HSIs), all frequency domain transforms are implemented within a 3-D framework. The model consists of a generator and a discriminator. The generator incorporates a frequency domain enhancement (FDE) module and a multisource information fusion (MIF) module; the discriminator incorporates a set of weight-sharing adaptive frequency domain transform (AFT) modules. The FDE module generates the extended domain (ED) with a certain domain shift by doing linear interpolation in the amplitude interval of a single SD itself. The MIF module integrates multisource information through interdomain attention, ensuring a balanced approach between the SD and ED, thus generating the effective balance domain (BD). The AFT module empowers the discriminator to selectively acquire HSI frequency domain features, facilitating synergistic collaboration of spatial-spectral features and frequency domain features for enhanced image comprehension. Extensive experiments on three public hyperspectral datasets show the superiority of the method compared with state-of-the-art techniques.
Jun Zhang 0050, Zhenwei Shi 0001, Bin Pan
IEEE Trans. Geosci. Remote. Sens.4
2024 Semantic-CC: Boosting Remote Sensing Image Change Captioning via Foundational Knowledge and Semantic Guidance
abstract
Remote sensing image change captioning (RSICC) aims to articulate the changes in objects of interest within bitemporal remote sensing images using natural language. Given the limitations of current RSICC methods in expressing general features across multitemporal and spatial scenarios, and their deficiency in providing granular, robust, and precise change descriptions, we introduce a novel change captioning (CC) method based on the foundational knowledge and semantic guidance, which we term Semantic-CC. Semantic-CC alleviates the dependency of high-generalization algorithms on extensive annotations by harnessing the latent knowledge of foundation models, and it generates more comprehensive and accurate change descriptions guided by pixel-level semantics from change detection (CD). Specifically, we propose a bitemporal SAM-based encoder for dual-image feature extraction; a multitask semantic aggregation neck for facilitating information interaction between heterogeneous tasks; a straightforward multiscale CD decoder to provide pixel-level semantic guidance; and a change caption decoder based on the large language model (LLM) to generate change description sentences. Moreover, to ensure the stability of the joint training of CD and CC, we propose a three-stage training strategy that supervises different tasks at various stages. We validate the proposed method on the LEVIR-CC and LEVIR-CD datasets. The experimental results corroborate the complementarity of CD and CC, demonstrating that Semantic-CC can generate more accurate change descriptions and achieve optimal performance across both tasks.
Yongshuo Zhu, Keyan Chen 0001, Fugen Zhou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Zero-Shot Text-to-Parameter Translation for Game Character Auto-Creation
abstract
Recent popular Role-Playing Games (RPGs) saw the great success of character auto-creation systems. The bone-drivenface model controlled by continuous parameters (like the position of bones) and discrete parameters (like the hairstyles) makes it possible for users to personalize and customize in-game characters. Previous in-game character auto-creation systems are mostly image-driven, where facial parameters are optimized so that the rendered character looks similar to the reference face photo. This paper proposes a novel text-to-parameter translation method (T2P) to achieve zero-shot text-driven game character auto-creation. With our method, users can create a vivid in-game character with arbitrary text description without using any reference photo or editing hundreds of parameters manually. In our method, taking the power of large-scale pre-trained multi-modal CLIP and neural rendering, T2P searches both continuous facial parameters and discrete facial parameters in a unified framework. Due to the discontinuous parameter representation, previous methods have difficulty in effectively learning discrete facial parameters. T2p, to our best knowledge, is the first method that can handle the optimization of both discrete and continuous parameters. Experimental results show that T2P can generate high-quality and vivid game characters with given text prompts. T2P outperforms other SOTA text-to-3D generation methods on both objective evaluations and subjective evaluations.
Rui Zhao 0019, Wei Li 0224, Zhipeng Hu, Lincheng Li, Zhengxia Zou, Zhenwei Shi 0001, Changjie Fan
CVPR6
2023 Progressive Scale-Aware Network for Remote Sensing Image Change Captioning
abstract
Remote sensing (RS) images contain numerous objects of different scales, which poses significant challenges for the RS image change captioning (RSICC) task to identify visual changes of interest in complex scenes and describe them via language. However, current methods still have some weaknesses in sufficiently extracting and utilizing multi-scale information. In this paper, we propose a progressive scale-aware network (PSNet) to address the problem. PSNet is a pure Transformer-based model. To sufficiently extract multi-scale visual features, multiple progressive difference perception (PDP) layers are stacked to progressively exploit the differencing features of bitemporal features. To sufficiently utilize the extracted multi-scale features for captioning, we propose a scale-aware reinforcement (SR) module and combine it with the Transformer decoding layer to progressively utilize the features from different PDP layers. Experiments show that the PDP layer and SR module are effective and our PSNet outperforms previous methods.
Zipeng Qi, Zhengxia Zou, Zhenwei Shi 0001
IGARSS5
2023 Resolution-Agnostic Remote Sensing Scene Classification With Implicit Neural Representations
abstract
Remote sensing scene classification is an important yet challenging task. In recent years, the excellent feature representation ability of convolutional neural networks (CNNs) has led to substantial improvements in scene classification accuracy. However, handling resolution variations of remote sensing images is still challenging because CNNs are not inherently capable of modeling multiresolution input images. In this letter, we propose a novel scene classification method with scale and resolution adaptation ability by leveraging the recent advances in implicit neural representations (INRs). Unlike previous CNN-based methods that make predictions based on rasterized image inputs, the proposed method converts the images as continuous functions with INRs optimization and then performs classification within the function space. When the image is represented as a function, the image resolution can be decoupled from the pixel values so that the resolution does not have much impact on the classification performance. Our method also shows great potential for multiresolution remote sensing scene classification. Using only a simple multilayer perceptron (MLP) classifier in the proposed function space, our method achieves classification accuracy comparable to deep CNNs but exhibits better adaptability to image scale and resolution changes.
Keyan Chen 0001, Wenyuan Li 0002, Jianqi Chen, Zhengxia Zou, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.5
2023 Inherit With Distillation and Evolve With Contrast: Exploring Class Incremental Semantic Segmentation Without Exemplar Memory
abstract
As a front-burner problem in incremental learning, class incremental semantic segmentation (CISS) is plagued by catastrophic forgetting and semantic drift. Although recent methods have utilized knowledge distillation to transfer knowledge from the old model, they are still unable to avoid pixel confusion, which results in severe misclassification after incremental steps due to the lack of annotations for past and future classes. Meanwhile data-replay-based approaches suffer from storage burdens and privacy concerns. In this paper, we propose to address CISS without exemplar memory and resolve catastrophic forgetting as well as semantic drift synchronously. We present Inherit with Distillation and Evolve with Contrast (IDEC), which consists of a Dense Knowledge Distillation on all Aspects (DADA) manner and an Asymmetric Region-wise Contrastive Learning (ARCL) module. Driven by the devised dynamic class-specific pseudo-labelling strategy, DADA distils intermediate-layer features and output-logits collaboratively with more emphasis on semantic-invariant knowledge inheritance. ARCL implements region-wise contrastive learning in the latent space to resolve semantic drift among known classes, current classes, and unknown classes. We demonstrate the effectiveness of our method on multiple CISS tasks by state-of-the-art performance, including Pascal VOC 2012, ADE20K and ISPRS datasets. Our method also shows superior anti-forgetting ability, particularly in multi-step CISS tasks.
Danpei Zhao, Bo Yuan 0009, Zhenwei Shi 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Object Detection in 20 Years: A Survey
abstract
Object detection, as of one the most fundamental and challenging problems in computer vision, has received great attention in recent years. Over the past two decades, we have seen a rapid technological evolution of object detection and its profound impact on the entire computer vision field. If we consider today’s object detection technique as a revolution driven by deep learning, then, back in the 1990s, we would see the ingenious thinking and long-term perspective design of early computer vision. This article extensively reviews this fast-moving research field in the light of technical evolution, spanning over a quarter-century’s time (from the 1990s to 2022). A number of topics have been covered in this article, including the milestone detectors in history, detection datasets, metrics, fundamental building blocks of the detection system, speedup techniques, and recent state-of-the-art detection methods.
Zhengxia Zou, Keyan Chen 0001, Zhenwei Shi 0001, Yuhong Guo, Jieping Ye
Proc. IEEE3
2023 Continuous Remote Sensing Image Super-Resolution Based on Context Interaction in Implicit Function Space
abstract
Despite its fruitful applications in remote sensing, image super-resolution is troublesome to train and deploy as it handles different resolution magnifications with separate models. Accordingly, we propose a highly-applicable super-resolution framework called FunSR, which settles different magnifications with a unified model by exploiting context interaction within implicit function space. FunSR composes a functional representor, a functional interactor, and a functional parser. Specifically, the representor transforms the low-resolution image from Euclidean space to multi-scale pixel-wise function maps; the interactor enables pixel-wise function expression with global dependencies; and the parser, which is parameterized by the interactor’s output, converts the discrete coordinates with additional attributes to RGB values. Extensive experimental results demonstrate that FunSR reports state-of-the-art performance on both fixed-magnification and continuous-magnification settings, meanwhile, it provides many friendly applications thanks to its unified nature. Our code is available at https://github.com/KyanChen/FunSR.
Keyan Chen 0001, Wenyuan Li 0002, Sen Lei, Jianqi Chen, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 Continuous Cross-Resolution Remote Sensing Image Change Detection
abstract
Most contemporary supervised Remote Sensing (RS) image Change Detection (CD) approaches are customized for equal-resolution bitemporal images. Real-world applications raise the need for cross-resolution change detection, aka, CD based on bitemporal images with different spatial resolutions. Given training samples of a fixed bitemporal resolution difference (ratio) between the high-resolution (HR) image and the low-resolution (LR) one, current cross-resolution methods may fit a certain ratio but lack adaptation to other resolution differences. Toward continuous cross-resolution CD, we propose scale-invariant learning to enforce the model consistently predicting HR results given synthesized samples of varying resolution differences. Concretely, we synthesize blurred versions of the HR image by random downsampled reconstructions to reduce the gap between HR and LR images. We introduce coordinate-based representations to decode per-pixel predictions by feeding the coordinate query and corresponding multi-level embedding features into an MLP that implicitly learns the shape of land cover changes, therefore benefiting recognizing blurred objects in the LR image. Moreover, considering that spatial resolution mainly affects the local textures, we apply local-window self-attention to align bitemporal features during the early stages of the encoder. Extensive experiments on two synthesized and one real-world different-resolution CD datasets verify the effectiveness of the proposed method. Our method significantly outperforms several vanilla CD methods and two cross-resolution CD methods on the three datasets both in in-distribution and out-of-distribution settings. The empirical results suggest that our method could yield relatively consistent HR change predictions regardless of varying bitemporal resolution ratios. Our code will be public.
Hao Chen 0045, Haotian Zhang 0010, Keyan Chen 0001, Chenyao Zhou, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 UnDAT: Double-Aware Transformer for Hyperspectral Unmixing
abstract
Deep-learning-based methods have attracted increasing attention on hyperspectral unmixing, where the transformer models have shown promising performance. However, recently proposed deep-learning-based hyperspectral unmixing methods usually tend to directly apply visual models, while ignoring the characteristics of hyperspectral imagery. In this article, we propose a novel double-aware transformer for hyperspectral Unmixing (UnDAT), which aims at simultaneously exploiting the region homogeneity and spectral correlation of hyperspectral imagery. One of the major assumptions of UnDAT is that hyperspectral remote-sensing images involve many homogeneous regions. Pixels inside a homogeneous region usually present similar spectral features, and the edge pixels are just the reverse. Another observation is that the pixel spectra are continuous and correlated. Based on the above assumption and observation, we construct the UnDAT by developing two modules: Score-based homogeneous-aware (SHA) module and the spectral group-aware (SGA) module. In the SHA module, a feature map rearrangement (FMR) approach is proposed to split the shallow feature maps from a linear encoder into an ordered homogeneous map (HomoMap) and an edge map and develop a homogenous region-aware strategy for deep feature representation. In the SGA module, the dependency among neighboring bands is described by dividing the hyperspectral image into multiple spectral groups and calculating the spectral similarity among bands within each group. Experiments on both real and synthetic datasets indicate the effectiveness of our model. We will publish the code of our approach if the article has the honor to be accepted.
Yuexin Duan, Tao Li 0022, Bin Pan, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 LiCa: Label-Indicate-Conditional-Alignment Domain Generalization for Pixel-Wise Hyperspectral Imagery Classification
abstract
One of the major difficulties for hyperspectral imagery (HSI) classification is the hyperspectral-heterospectra, which refers to the same material presenting different spectra. Although joint spatial-spectral classification methods can relieve this problem, they may lead to falsely high accuracy because the test samples may be involved during the training process. How to address the hyperspectral-heterospectra problem remains a great challenge for pixel-wise hyperspectral imagery classification methods. Domain generalization is a promising technique that may contribute to the heterospectra problem, where the different spectra of the same material can be considered as several domains. In this paper, inspired by the theory of domain generalization, we provide a formulaic expression for hyperspectral-heterospectra. To be specific, we consider the spectra of one material as a conditional distribution and propose a domain-generalization-based method for pixel-wise HSI classification. The key of our proposed method is a new Label-indicate-Conditional-alignment (LiCa) block that focuses on aligning the spectral conditional distributions of different domains. In the LiCa block, we define two loss functions, cross-domain conditional alignment, and cross-domain entropy, to describe the heterogeneity of HSI. Moreover, we have provided the theoretical foundation for the newly-proposed loss functions, by analyzing the upper bound of classification error in any target domains. Experiments on several public data sets indicate that the LiCa block has achieved better generalization performance when compared with other pixel-wise classification methods.
Bin Pan, Tao Li 0022, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Toward Convergence: A Gradient-Based Multiobjective Method With Greedy Hash for Hyperspectral Unmixing
abstract
Multiobjective optimization aims at addressing the conflicting objectives, which has been introduced to improve the performance of sparse hyperspectral unmixing. Recently proposed multiobjective unmixing methods usually employ evolutionary algorithms to improve the unmixing accuracy. However, evolutionary algorithms may suffer the challenge of convergence, in which case the reasonability of the solutions is hard to guarantee. To solve the problem of convergence, in this paper, we present a new gradient-based multiobjective unmixing method, which explores the optimization direction in a theoretically reliable manner. Furthermore, considering the mathematical model of hyperspectral sparse unmixing where sparsity error objective of selected endmembers is discrete, we develop a greedy hash based coding approach which is able to well describe the discrete constraints imposed on endmembers. The major components of the proposed method are a search approach and an update approach. In the search approach, we construct the pareto descent direction via a gradient-based strategy, which contributes to converging to an optimal continuous solution by searching along this direction. In the update approach, we update discrete binary endmember via hash coding under the guidance of greedy principle, which allows our method to handle the problem of discrete objective. The major contribution of the proposed method is designing a new framework that can get the optimal discrete endmembers in a convergent way. Moreover, we provide the theoretical analysis and proof for the convergence. Synthetic and real-world experiments have indicated the advantages of our algorithm when compared with evolutionary multiobjective unmixing methods.
Ruiying Li, Bin Pan, Tao Li 0022, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 CoI2A: Collaborative Inter-domain and Intra-domain Alignments for Multisource Domain Adaptation
abstract
In the remote sensing information interpretation tasks, compared with collecting lots of high-quality image labels for the target domain, a large amount of labeled remote sensing data from multiple source domains are generally available without any extra cost. In this paper, our work focuses on how to exploit the rich knowledge obtained from multiple source domains to guide the interpretation of the target scene, and we propose a novel framework called Collaborative Inter-domain and Intra-domain Alignments for multi-source domain adaptation, namely CoI2A, in which inter-domain and intra-domain alignments are well collaborated to reduce the distribution divergence across sources and target. To reduce the discrepancy across sources, the inter-source alignment is proposed to map multiple sources into a unified representation space. In addition, the cross-domain attention is introduced to enforce the intra-class compactness of the target. Inter-domain alignment aligns each source with target domain separately with the help of cross-domain attention. As for the intra-domain alignment, the multi-head attentive representations of the target obtained by cross-domain attention are correlated into a unified one. The experimental results obtained from different scene classification tasks demonstrate the superiority of our model.
Zhenfeng Zhu, Shenghui Wang 0003, Zhenwei Shi 0001, Yao Zhao 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Diverse Hyperspectral Remote Sensing Image Synthesis With Diffusion Models
abstract
Hyperspectral image synthesis overcomes the limitations of imaging sensors and enables low-cost acquisition of hyperspectral images with high spatial resolution. Using RGB as a conditional input for hyperspectral generation is promising and valuable, as it can leverage abundant existing multispectral/RGB images without the intervention of hyperspectral sensors. However, most existing generation methods follow one-to-one mapping frameworks and ignore generation diversity. In addition, the current evaluation metrics of hyperspectral generation are based on the similarity with the reference image, which cannot reflect the diversity of the generated spectra. In this paper, we propose a novel method for diverse hyperspectral remote sensing image generation based on the diffusion model. The diffusion model uses a denoising model to gradually remove noise from the normal distribution and generates the hyperspectral data step-by-step with the conditional RGB image as input. To address the high-dimensional noise prediction problem caused by a large number of bands in the hyperspectral image, we introduce a conditional VQGAN that maps the high-dimension hyperspectral data into a low-dimension latent space and conduct the diffusion process in the latent space. The latent-diffusion process makes the diffusion process faster and more stable. The conditional VQGAN decodes hyperspectral images from the latent code generated by diffusion, with the conditional RGB image as input, which restricts the diversity to a specific object distribution. We also design two new metrics to evaluate the generation spectral diversity. Experiments on the IEEEgrss_dfc_2018dataset demonstrate that our method can synthesize highly diverse hyperspectral data. In addition, the rationality of the proposed metrics is also verified.
Liqin Liu, Bowen Chen 0002, Hao Chen 0045, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 A Decoupling Paradigm With Prompt Learning for Remote Sensing Image Change Captioning
abstract
Remote sensing image change captioning (RSICC) is a novel task that aims to describe the differences between bi-temporal images by natural language. Previous methods ignore a significant specificity of the task: the difficulty of RSICC is different for unchanged and changed image pairs. They process the unchanged and changed image pairs in a coupled way, which usually causes confusion for change captioning. In this paper, we decouple the task into two issues to ease it: whether and what changes have occurred. An image-level classifier performs binary classification to address the first issue. A feature-level encoder contributes to extracting discriminative features to help the caption generation module address the second issue. Besides, for caption generation, we utilize prompt learning to introduce pre-trained large language models (LLMs) into the RSICC task. A multi-prompt learning strategy is proposed to generate a set of unified prompts and a class-specific prompt conditioned on the image-level classifier’s results. The strategy can prompt a pre-trained LLM to know whether changes exist and generate captions. Finally, the multiple prompts and the visual features of the feature-level encoder are fed into a frozen LLM for language generation. Compared with previous methods, our method can leverage the powerful abilities of the pre-trained LLM in language to generate plausible captions, which is free of training. Extensive experiments show that our method is effective and achieves state-of-the-art performance. Besides, an additional experiment demonstrates that our decoupling paradigm is more promising than the previous coupled paradigm for the RSICC task. We will make our codebase publicly available to facilitate future research at https://github.com/Chen-Yang-Liu/PromptCC.
Rui Zhao 0019, Jianqi Chen, Zipeng Qi, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Hyperspectral Remote Sensing Image Synthesis Based on Implicit Neural Spectral Mixing Models
abstract
Hyperspectral image (HSI) synthesis, as an emerging research topic, is of great value in overcoming sensor limitations and achieving low-cost acquisition of high-resolution remote sensing HSIs. However, the linear spectral mixing model used in recent studies oversimplifies the real-world hyperspectral imaging process, making it difficult to effectively model the imaging noise and multiple reflections of the object spectrum. As a prerequisite for hyperspectral data synthesis, accurate modeling of nonlinear spectral mixtures has long been a challenge. Considering the above difficulties, we propose a novel method for modeling nonlinear spectral mixtures based on implicit neural representations (INRs) in this article. The proposed method learns from INR and adaptively implements different mixture models for each pixel according to their spectral signature and surrounding environment. Based on the above neural mixing model, we also propose a new method for HSI synthesis. Given an RGB image as input, our method can generate an accurate and physically meaningful HSI. As a set of by-products, our method can also generate subpixel-level spectral abundance as well as the solar atmosphere signature. The whole framework is trained end-to-end in a self-supervised manner. We constructed a new dataset for HSI synthesis based on a wide range of Airborne Visible Infrared Imaging Spectrometer (AVIRIS) data. Our method achieves a mean peak signal-to-noise ratio (MPSNR) of 52.36 dB and outperforms other state-of-the-art hyperspectral synthesis methods. Finally, our method shows great benefits to downstream data-driven applications. With the HSIs and abundance directly generated from low-cost RGB images, the proposed method improves the accuracy of HSI classification tasks by a large margin, particularly for those with limited training samples.
Liqin Liu, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 An Imbalanced Discriminant Alignment Approach for Domain Adaptive SAR Ship Detection
abstract
Synthetic aperture radar (SAR) imaging has round-the-clock data acquisition capability regardless of light and climate constraints, so it has been widely used for ship detection. However, SAR images usually suffer lower imaging quality, which may result in indistinct contours and non-negligible noise. Therefore, the manual labeling for SAR images is expensive, leading to a lack of training data in the task of ship detection. In this paper, we propose a route by utilizing domain adaptive methods to transfer information from labeled visible images (source domain) to unlabeled SAR images (target domain) for ship detection. To address the distribution mismatch between domains, we develop a novel imbalanced discriminant alignment (IDA) approach to improve the discriminant ability of the network and prevent negative migration. The core of the IDA approach is applying a new loss function called imbalanced prediction consistency (IPC) loss to describe the domain classifier consistency, and we further provide theoretical analysis for the effectiveness of the IPC loss. IDA ensures consistency at the image level and instance level, and focuses on the consistency of the source domain to enhance the feature extraction capability of the adversarial network. The theoretical discussion has proven that a necessary and sufficient condition for convergence of the IPC loss is that the two discriminant probabilities converge to 0 at the discriminant distance we define. Experimental results have indicated the advantage of IDA when compared with other domain adaptation SAR ship detection methods.
Bin Pan, Zhehao Xu, Tianyang Shi, Tao Li 0022, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Implicit Ray Transformers for Multiview Remote Sensing Image Segmentation
abstract
The mainstream CNN-based remote sensing (RS) image semantic segmentation approaches typically rely on massively labeled training data. Such a paradigm struggles with the problem of RS multi-view scene segmentation with limited labeled views due to the lack of consideration of 3D information within the scene. In this paper, we propose “Implicit Ray-Transformer (IRT)” based on Implicit Neural Representation (INR) for RS scene semantic segmentation with sparse labels (5% of the images being labeled). We explore a new way of introducing the multi-view 3D structure priors to the task for accurate and view-consistent semantic segmentation. The proposed method includes a two-stage learning process. In the first stage, we optimize a neural field to encode the color and 3D structure of the remote sensing scene based on multi-view images. In the second stage, we design a Ray Transformer to leverage the relations between the neural field 3D features and 2D texture features for learning better semantic representations. Different from previous methods that only consider 3D priors or 2D features, we incorporate additional 2D texture information and 3D priors by broadcasting CNN features to different point features along the sampled ray. To verify the effectiveness of the proposed method, we construct a challenging dataset containing six synthetic sub-datasets collected from the Carla platform and three real sub-datasets from Google Maps. Experiments show that the proposed method outperforms the CNN-based methods and the state-of-the-art INR-based segmentation methods in quantitative and qualitative metrics. The ablation study shows that under a limited number of fully annotated images, the combination of the 3D structure priors and 2D texture can significantly improve the performance and effectively complete missing semantic information in novel views. Experiments also demonstrate the proposed method could yield geometry-consistent segmentation results against illumination changes and viewpoint changes. Our data and code will be public.
Zipeng Qi, Hao Chen 0045, Zhenwei Shi 0001, Zhengxia Zou
IEEE Trans. Geosci. Remote. Sens.4
2023 Unsupervised Multimodal Remote Sensing Image Registration via Domain Adaptation
abstract
Registration of multi-modal remote sensing images with geometric distortions is one of the fundamental applications, but it remains difficult since multi-modal remote sensing images have significant differences in both radiometric and geometric features. One of the challenges is the disregarding of modality-specific information, which hinders the model from focusing on the content information of structure and texture due to differences in radiometric features. In this paper, an unsupervised Content-focused Hierarchical Alignment Network (CHA-Net) is proposed, which is constructed based on the theory of domain adaptation. The kernel idea of CHA-Net is to weaken the style differences among different modal images and achieve non-rigid multi-modal remote sensing image registration. CHA-Net is a hierarchical refinement model, where different scales of features are aligned respectively by utilizing the field calibration module and gradually generating the registration field. To be specific, CHA-Net consists of two structures: the Siamese Feature Decoupling (SFD) structure and the Hierarchical Refinement Alignment (HRA) structure. The SFD aims at reducing the style differences caused by cross-modal differences and developing a shared-weight Siamese network to map images to content feature space. The HRA enhances the ability of the network by capturing global distortions based on the Transformer model. Experiments on public datasets indicate that compared with other methods, CHA-Net performs better when geometric and radiometric distortions appear.
Lukui Shi, Ruiyun Zhao, Bin Pan, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Remote Sensing Image Synthesis via Semantic Embedding Generative Adversarial Networks
abstract
Generating photo-realistic remote sensing images conditioned on semantic masks has many practical applications like image editing, detecting deep fake geography, and data augmentation. Although previous methods achieved high-quality synthesis results for natural images like faces and everyday objects, they still underperform in remote sensing scenarios in terms of both visual fidelity and diversity. The high data imbalance and high semantic similarity of remote sensing object categories make the semantic synthesis of remote sensing images more challenging than natural images. To tackle these challenges, we propose a novel method named Conducted Semantic EmBedding GAN (CSEBGAN) for semantic-controllable remote sensing image synthesis. The proposed method decouples different semantic classes into independent Semantic Embeddings, which explores the regularities between classes to improve visual fidelity and naturally supports semantic-level. We further introduce a novel tripartite cooperation adversarial training scheme that involves a conductor network to provide fine-grained semantic feedback for the generator. We also show that the proposed semantic image synthesis method can be utilized as an effective data augmentation approach on improving the performance of the downstream remote sensing image segmentation tasks. Extensive experiments show the superiority of our method compared with the state-of-the-art image synthesis methods.
Chendan Wang, Bowen Chen 0002, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 A Hierarchical Decoder Architecture for Multilevel Fine-Grained Disaster Detection
abstract
As a cutting-edge challenge in the field of disaster evaluation, the detection of disasters in remote sensing images is crucial. However, most existing approaches to disaster detection simply solve the problem as a naive multi-class change detection, lacking accurate damage-level classification. In this paper, we propose a new approach to disaster detection called multi-level disaster detection (MLDD) that focuses on fine-grained damage-level classification. Our proposed approach tackles MLDD through hierarchical-correlation modeling and presents a universal disaster detection architecture. Specifically, we summarize two existing applicative methods, one-step training and pre-training, which are compatible with our proposed architecture. In addition, we propose two novel hierarchical approaches, namely the multi-task (MT) based and graph-encoding (GE) based approaches. The MT approach resolves MLDD through layer-wise learning in a progressive manner, building explicit multi-stage and implicit joint models to probe into the coarse-to-fine correlation for damage-level evaluation. The GE approach enhances hierarchical relationships by encoding multifold messaging directions and probabilities using a graph neural network. Furthermore, all four hierarchical paradigms can be embedded in our hierarchical MLDD architecture, which outperforms state-of-the-art methods on the xBD dataset, particularly in fine-grained damage-level classification. Overall, our proposed approach represents a significant improvement over existing disaster detection methods and has the potential to advance the field of disaster evaluation.
Chenxu Wang 0017, Danpei Zhao, Xinhu Qi, Zhuoran Liu 0006, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 A Bayesian Meta-Learning-Based Method for Few-Shot Hyperspectral Image Classification
abstract
Few-shot learning provides a new way to solve the problem of insufficient training samples in hyperspectral classification. It can implement reliable classification under several training samples by learning meta-knowledge from similar tasks. However, most existing works perform frequency statistics, which may suffer from the prevalent uncertainty in point estimates (PEs) with limited training samples. To overcome this problem, we reconsider the hyperspectral image few-shot classification (HSI-FSC) task as a hierarchical probabilistic inference from a Bayesian view and provide a careful process of meta-learning probabilistic inference. We introduce a prototype vector for each class as latent variables and adopt distribution estimates (DEs) for them to obtain their posterior distribution. The posterior of the prototype vectors is maximized by updating the parameters in the model via the prior distribution of HSI and labeled samples. The features of the query samples are matched with prototype vectors drawn from the posterior; thus, a posterior predictive distribution over the labels of query samples can be inferred via an amortized Bayesian variational inference approach. Experimental results on four datasets demonstrate the effectiveness of our method. Especially given only three to five labeled samples, the method achieves noticeable upgrades of overall accuracy (OA) against competitive methods.
Jing Zhang 0127, Liqin Liu, Rui Zhao 0019, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Classification Matters More: Global Instance Contrast for Fine-Grained SAR Aircraft Detection
abstract
Since significant intraclass differences and inconspicuous interclass variations, fine-grained aircraft detection in synthetic aperture radar (SAR) images is challenging. Also, the inherent lack of detailed features and severe noise interference in SAR images make it difficult to learn class-specific feature representations. Current detection approaches focus more on localization accuracy and ignore classification performance, which is more critical in fine-grained detection. To address the above challenges, we present GICNet: global instance contrast (GIC) for fine-grained SAR aircraft detection a global instance-level contrast module is proposed to improve interclass divergences and intraclass compactness. With a specially constructed global instance set, GICNet can contrast a large number of different aircraft targets while keeping a small batch size. Furthermore, we design a novel quality-aware focal loss (QAFL) to facilitate the accurate classification of well-localized aircraft targets. Meanwhile, to maintain localization performance, we develop a new edge-aware bounding-box refinement (EABR) module to refine predicted coarse bounding boxes. Experimental results show that our GICNet outperforms current advanced detectors and achieves a new state-of-the-art performance on the GaoFen-3 SAR aircraft detection dataset. In particular, GICNet also has advantages in reducing misclassification and recognizing well-located targets.
Danpei Zhao, Yue Gao 0008, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Unmixing Guided Unsupervised Network for RGB Spectral Super-Resolution
abstract
Spectral super-resolution has attracted research attention recently, which aims to generate hyperspectral images from RGB images. However, most of the existing spectral super-resolution algorithms work in a supervised manner, requiring pairwise data for training, which is difficult to obtain. In this paper, we propose an Unmixing Guided Unsupervised Network (UnGUN), which does not require pairwise imagery to achieve unsupervised spectral super-resolution. In addition, UnGUN utilizes arbitrary other hyperspectral imagery as the guidance image to guide the reconstruction of spectral information. The UnGUN mainly includes three branches: two unmixing branches and a reconstruction branch. Hyperspectral unmixing branch and RGB unmixing branch decompose the guidance and RGB images into corresponding endmembers and abundances respectively, from which the spectral and spatial priors are extracted. Meanwhile, the reconstruction branch integrates the above spectral-spatial priors to generate a coarse hyperspectral image and then refined it. Besides, we design a discriminator to ensure that the distribution of generated image is close to the guidance hyperspectral imagery, so that the reconstructed image follows the characteristics of a real hyperspectral image. The major contribution is that we develop an unsupervised framework based on spectral unmixing, which realizes spectral super-resolution without paired hyperspectral-RGB images. Experiments demonstrate the superiority of UnGUN when compared with some SOTA methods.
Qiaoying Qu, Bin Pan, Tao Li 0022, Zhenwei Shi 0001
IEEE Trans. Image Process.5
2022 Semantic Decoupled Representation Learning for Remote Sensing Image Change Detection
abstract
Self-supervised learning (SSL) has recently been introduced to remote sensing (RS) to learn in-domain transferable representations. Here, we propose a semantic decoupled representation learning for RS image change detection (CD). Typically, the object of interest (e.g., building) is relatively small compared to the vast background. Different from existing methods expressing an image into one representation vector that may be dominated by irrelevant land-covers, we disentangle representations of different semantic regions by leveraging the semantic mask. We additionally force the model to distinguish different semantic representations, which benefits the recognition of objects of interest in the downstream CD task. We construct a dataset of bitemporal images with semantic masks in an effort-less manner for pre-training. Experiments on two CD datasets show our model outperforms ImageNet, indomain supervised pre-training, and several recent SSL methods.
Hao Chen 0045, Yifan Zao, Liqin Liu, Zhenwei Shi 0001
IGARSS5
2022 Hyperspectral Image Generation From Rgb Images With Semantic and Spatial Distribution Consistency
abstract
Generating hyperspectral images (HSI) from RGB imagery can obtain HSI with both high spatial and spectral resolution, which overcomes the limitations of imaging hardware conditions. Many HSI generation methods target learning a 3-n mapping from RGB to HSI, lacking concern of the spectral categories and spatial distribution. In this paper, we propose an HSI generation method preserving the band structure similarity and semantic information. We design an MLP based classifier and trained it on many spectra of known semantic categories. Then we use it to map the spectra to semantic space and constrain the distance between the embedding of generated spectra and that of the real ones. Meanwhile, a structure similarity loss is added to constrain the spatial information. Experiment results verified the superiority of the proposed method.
Liqin Liu, Zhenwei Shi 0001, Yifan Zao, Hao Chen 0045
IGARSS2
2022 Enhance Essential Features for Road Extraction from Remote Sensing Images
abstract
In deep learning based road extraction from remote sensing images, the network often learns some features that are not essential to road discrimination, such as trees, buildings, etc. In fact, there is no causal relationship between these features and road discrimination, which will lead to error and omission in final results. In this paper, we propose a novel road extraction network to enhance essential features, including local and global line features and geometric features along the road direction. Multi-scale Line Enhancement Module utilize hough transform to enhance line featues of different scales. Neighboring road prediction branch make the network pre-dict the distance and direction of each pixel to the neighboring road, which helps the network to focus on geometric features along the road direction. Experimental results on the deepglobe dataset show that the network is able to obtain bet-ter road extraction results by enhancing essential features that have a causal relationship with the task. Codes are available at https://github.com/zaoyifan/EssentialFeatures.
Yifan Zao, Hao Chen 0045, Liqin Liu, Zhenwei Shi 0001
IGARSS4
2022 Remote-Sensing Image Captioning Based on Multilayer Aggregated Transformer
abstract
Remote-sensing image (RSI) captioning aims to automatically generate sentences describing the content of RSIs. The multiscale information of RSIs contains attributes and complex relationships of objects of different sizes. However, current methods still have some weaknesses in efficiently utilizing multiscale information to generate accurate and detailed sentences. In this letter, we propose a new model based on the “encoder–decoder” framework to address the problem. In the encoder, we fuse the features of different layers in ResNet-50 to extract multiscale information. In the decoder, we propose multilayer aggregated transformer (MLAT) to utilize the extracted information to generate sentences sufficiently. Specially, as the transformer encoding layer goes deeper, the extracted features will be more similar. To sufficiently utilize the features from different transformer encoding layers, compress redundant information, and extract important information, long short-term memory (LSTM) in MLAT aggregates the features to obtain better feature representations. The self-attention mechanism and the aggregation strategy enable MLAT to utilize the features sufficiently. The experimental results show that MLAT as the decoder can help the model address the multiscale problem, significantly improve the model performance on sentence accuracy and diversity, and show that our proposed method performs better than other current methods. Our code is available athttps://github.com/Chen-Yang-Liu/MLAT.
Rui Zhao 0019, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.3
2022 Remote-Sensing Image Segmentation Based on Implicit 3-D Scene Representation
abstract
Remote sensing image segmentation, as a challenging but fundamental task, has drawn increasing attention in the remote sensing field. Recent advances in deep learning have greatly boosted research on this task. However, the existing deep learning-based segmentation methods heavily rely on a large amount of pixel-wise labeled training data, and the labeling process is time-consuming and labor-intensive. In this paper, we focus on the scenario that leverages the 3D structure of multi-view images and a limited number of annotations to generate accurate novel view segmentation. Under this scenario, we propose a novel method for remote sensing image segmentation based on implicit 3D scene representation, which generates arbitrary-view segmentation output from limited segmentation annotations. The proposed method employs a two-stage training strategy. In the first stage, we optimize the implicit neural representations of a 3D scene and encode their multi-view images into a neural radiance field. In the second stage, we transform the scene color attribute into semantic labels and propose a ray-convolution network to aggregate local 3D consistency cues across different locations. We also design a color-radiance network to help our method generalize to unseen views. Experiments on both synthetic and real-world data suggest that our method significantly outperforms deep convolutional networks (CNN)-based methods and other view synthesis-based methods. We also show that the proposed method can be applied as a novel data augmentation approach that benefits CNN-based segmentation methods.
Zipeng Qi, Zhengxia Zou, Hao Chen 0045, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 An All-Scale Feature Fusion Network With Boundary Point Prediction for Cloud Detection
abstract
Cloud detection is a significant pre-processing for remote sensing images. In recent years, many methods based on deep learning are proposed to detect clouds and multi-scale feature fusion is often used in these methods. However, most existing methods fuse features through concatenation and element-wise summation, which are simple and can be improved in spatial information recovery. Therefore, we explore the way of fusing features to recover the missing spatial information more sufficiently. Besides, we also observe that some cloud detection results are not accurate enough near the boundary of clouds. In view of the above observations, in this letter, we propose a cloud detection network, ABNet, which includes All-scale feature Fusion modules and a Boundary point Prediction module. The All-scale feature Fusion module can optimize the features and recover spatial information by integrating features of all scales. And the Boundary point Prediction module further remedies cloud boundary information by classifying the cloud boundary points separately. Experimental results demonstrate that our method improves the accuracy of cloud detection compared with other methods.
Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Tropical Cyclone Forecast Using Multitask Deep Learning Framework
abstract
A tropical cyclone is a robust weather system that affects human daily life. Accurate and rapid tropical cyclone forecast can guide human disaster prevention and mitigation work against tropical cyclones. The mainstream tropical cyclone forecasting method is numerical forecasting, which requires abundant prior knowledge and luxurious calculation. Nowadays, machine learning methods have received increasing attention for which they can overcome these disadvantages. However, existing machine learning methods usually ignored some potential factors since they mainly concentrated on one aspect of the tropical cyclone forecast. This letter proposes a multitask machine learning framework to forecast tropical cyclone path and intensity, which possesses two modules: one is the prediction module and the other is the estimate module. We use an improved generative adversarial network as the prediction module to predict the tropical cyclone spatial data at a certain moment in the future. Then, we use two different deep neural networks as the estimation module to extract the position and intensity from the generated prediction data. The method we propose is a general and relatively accurate tropical cyclone forecast method. We reach a 24-h path forecast error of 116 km and a 24-h intensity forecast error of 13.06 kt.
Yuqiao Wu, Xiaoyi Geng, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Scene Aggregation Network for Cloud Detection on Remote Sensing Imagery
abstract
There has been a breakthrough in cloud detection by using convolutional neural networks (CNNs) during these years. However, there are still weaknesses among current cloud detection algorithms because only cloud mask information is used. As clouds represent differently in different scenes, the scene information may give hints to improve cloud detection performance. Therefore, different from the previous cloud detection literature, in this letter, we propose an end-to-end new deep learning network named scene aggregation network (SAN), which aggregates the scene information in the framework. Specifically, basic features are first extracted by utilizing all levels of network features. Then, the aggregated features used to produce the final cloud masks are created by fusing the basic features and the specially introduced scene information. Experimental results have demonstrated that with scene information aggregated, our proposed method can be robust on images with different scenes. Additionally, as SAN outperforms other state-of-the-art methods, our proposed method suits for cloud detection and can achieve improvement on this task.
Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Richer U-Net: Learning More Details for Road Detection in Remote Sensing Images
abstract
Road detection in remote sensing images has been an important research topic in the past few decades. However, with complex backgrounds and occlusion of vehicles and trees, it is difficult for most road detection methods to obtain complete and accurate results. There will be a large number of error and omission detections in such complex scenes due to the poor utilization of detailed information. Therefore, in this article, we propose a novel road detection method called Richer U-Net, which alleviates this problem by designing two detail enhancement strategies. First, considering that convolution operation will cause the loss of detailed information in the feature map, an enhanced detail recovery structure (EDRS) is introduced to make full use of those lost information. It combines the output of each convolutional layer at the same level for the detail recovery of decoding network, leading to more accurate segmentation results. Second, an edge-focused loss function is proposed to guide the network to pay more attention to the road edge area. By adding an enhancement factor, the pixels closer to edge will contribute more loss. The corresponding experiments are conducted on two public datasets, and it can be shown that our method effectively improves final detection results.
Yifan Zao, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Text-to-Remote-Sensing-Image Generation With Structured Generative Adversarial Networks
abstract
Synthesizing high-resolution remote sensing images based on the given text descriptions has great potential in expanding the image data set to release the power of deep learning in the remote sensing image processing field. However, there has been no efficient research carried out on this formidable task yet. Given a remote sensing image, the structural rationality of ground objects is critical to judge it whether real or fake, e.g., real bridges are always straight, while a sinuous one can be easily judged as fake. Inspired by this, we propose a multistage structured generative adversarial network (StrucGAN) to synthesize remote sensing images in a structured way given the text descriptions. StrucGAN utilizes structural information extracted by an unsupervised segmentation module to enable the discriminators to distinguish the image in a structured way. The generators of StrucGAN are, thus, forced to synthesize structural reasonable image contents, which could enhance the image authenticity. The multistage framework enables the StrucGAN to generate remote sensing images with increasing resolution stage by stage. The quantitative and qualitative experiments’ results show that the proposed StrucGAN achieves better performance compared with the baseline, and it could synthesize high resolution, realistic, structural reasonable remote sensing images that are semantically consistent with the given text descriptions.
Rui Zhao 0019, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Semantic Segmentation of Remote Sensing Image Based on Regional Self-Attention Mechanism
abstract
In remote sensing images (RSIs), accurate semantic segmentation faces more challenges because of small targets, unbalanced categories, and complex scenes. Restricted by local receptive field of convolution layers, the traditional semantic segmentation models cannot use global information of RSIs. According to the characteristics of RSIs, we propose an RSANet based on regional self-attention mechanism. Our model is no longer limited by the locality of convolution, but transfers the information flow in the whole image. It can mine out the relationship between pixels in the surrounding areas, which is more logical for understanding images content. Moreover, compared with the traditional self-attention mechanism, RSANet can effectively reduce the noise of feature maps and the interference of redundant features. Our model can get better semantic segmentation results than other current models on the DroneDeploy data set and the Chreos semantic segmentation data set. The experiments show that our RSANet achieves 2% higher mean intersection over union (mIoU) than the baseline model, especially in terms of fineness, edge integrity, and classification accuracy.
Danpei Zhao, Chenxu Wang 0017, Yue Gao 0008, Zhenwei Shi 0001, Fengying Xie
IEEE Geosci. Remote. Sens. Lett.4
2022 UGCNet: An Unsupervised Semantic Segmentation Network Embedded With Geometry Consistency for Remote-Sensing Images
abstract
In remote-sensing image (RSI) semantic segmentation, the dependence on large-scale and pixel-level annotated data has been a critical factor restricting its development. In this letter, we propose an unsupervised semantic segmentation network embedded with geometry consistency (UGCNet) for RSIs, which imports the adversarial-generative learning strategy into a semantic segmentation network. The proposed UGCNet can be trained on a source-domain dataset and achieve accurate segmentation results on a different target-domain dataset. Furthermore, for refining the remote-sensing target geometric representation such as densely distributed buildings, we propose a geometry-consistency (GC) constraint that can be embedded in both image-domain adaptation process and semantic segmentation network. Therefore, our model could achieve cross-domain semantic segmentation with target geometric property preservation. The experimental results on Massachusetts and Inria buildings datasets prove that the proposed unsupervised UGCNet could achieve a very comparable segmentation accuracy with the fully supervised model, which validates the effectiveness of the proposed method.
Danpei Zhao, Bo Yuan 0009, Yue Gao 0008, Xinhu Qi, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.5
2022 Neural Rendering for Game Character Auto-Creation
abstract
Many role-playing games feature character creation systems where players are allowed to edit the facial appearance of their in-game characters. This paper proposes a novel method to automatically create game characters based on a single face photo. We frame this "artistic creation" process under a self-supervised learning paradigm by leveraging the differentiable neural rendering. Considering the rendering process of a typical game engine is not differentiable, an "imitator" network is introduced to imitate the behavior of the engine so that the in-game characters can be smoothly optimized by gradient descent in an end-to-end fashion. Different from previous monocular 3D face reconstruction which focuses on generating 3D mesh-grid and ignores user interaction, our method produces fine-grained facial parameters with a clear physical significance where users can optionally fine-tune their auto-created characters by manually adjusting those parameters. Experiments on multiple large-scale face datasets show that our method can generate highly robust and vivid game characters. Our method has been applied to two games and has now provided over 10 million times of online services.
Tianyang Shi, Zhengxia Zou, Zhenwei Shi 0001, Yi Yuan 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Contrastive Learning for Fine-Grained Ship Classification in Remote Sensing Images
abstract
Fine-grained image classification can be considered as a discriminative learning process where images of different subclasses are separated from each other while the same subclass images are clustered. Most existing methods perform synchronous discriminative learning in their approaches. Although achieving promising results in fine-grained visual classification (FGVC) in natural images, these methods may fail in fine-grained ship classification (FGSC) problem in remote sensing (RS) images due to the highly “imbalanced fineness" and “imbalanced appearances" of ships among subclasses. To tackle the issue, we propose an asynchronous contrastive learning-based method for effective FGSC. The proposed method, which we refer to as “Push-and-Pull Network (P2Net)", includes a “push-out stage” and a “pull-in stage”, where the first stage forces all the instances to be de-correlated and then the second one groups them into each subclass. A dual-branch network is designed to separate/de-correlate the images with each other, while an Integration Module is designed to aggregate the de-correlated images into their corresponding subclass together with a Proxy-based Module designed for acceleration. In this way, the correlation between subclasses can be decoupled, which in turn makes the final classification much easier. Our method can be trained end-to-end and requires no additional annotations other than category information. Extensive experiments are conducted on two large-scale FGSC datasets (FGSC-23 and FGSCR-42). Our method outperforms other state-of-the-art approaches. Ablation experiments also suggest the effectiveness of our design. Our code is available at https://github.com/WindVChen/Push-and-Pull-Network.
Jianqi Chen, Keyan Chen 0001, Hao Chen 0045, Wenyuan Li 0002, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 A Degraded Reconstruction Enhancement-Based Method for Tiny Ship Detection in Remote Sensing Images With a New Large-Scale Dataset
abstract
The rapid detection of ships within the wide sea area is essential for intelligence acquisition. Most modern deep learning-based ship detection methods focus on locating ships in high-resolution (HR) remote sensing (RS) images. Seldom efforts have been made on ship detection in medium-resolution (MR) RS images. An MR image covers a much wider area than an HR one of the same size, thus facilitating quick ship detection. To this end, we propose a tiny ship detection method namely, Degraded Reconstruction Enhancement Network (DRENet), for MR RS images. Different from previous methods that mainly focus on feature fusion strategies to improve the expression ability of the detector, we design an additional network branch, i.e., degraded reconstruction enhancer, to learn to regress an object-aware blurred version of the input image in the training phase. Our intuition is that the proposed reconstruction branch may guide the backbone to focus more on tiny ship targets instead of the vast background. Moreover, we incorporate a CRoss-stage Multi-head Attention module in the detector to further improve the feature discrimination by leveraging the self-attention mechanism. To fill the gap of lacking a large-scale MR ship detection dataset, we introduce Levir-Ship, which contains 3876 GF-1/GF-6 multi-spectral images and over 3K tiny ship instances. Experiments on Levir-Ship validate the effectiveness and efficiency of the proposed method. Our method achieves 82.4 AP with 85 FPS, which outperforms many state-of-the-art ship detection methods. Our code and dataset will be made public.
Jianqi Chen, Keyan Chen 0001, Hao Chen 0045, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Semantic-Aware Dense Representation Learning for Remote Sensing Image Change Detection
abstract
Supervised deep learning models depend on massive labeled data. Unfortunately, it is time-consuming and labor-intensive to collect and annotate bitemporal samples containing desired changes. Transfer learning from pretrained models is effective to alleviate label insufficiency in remote sensing (RS) change detection (CD). We explore the use of semantic information during pretraining. Different from traditional supervised pretraining that learns the mapping from image to label, we incorporate semantic supervision into the self-supervised learning (SSL) framework. Typically, multiple objects of interest (e.g., buildings) are distributed in various locations in an uncurated RS image. Instead of manipulating image-level representations via global pooling, we introduce point-level supervision on per-pixel embeddings to learn spatially sensitive features, thus benefiting downstream dense CD. To achieve this, we obtain multiple points via class-balanced sampling on the overlapped area between views using the semantic mask. We learn an embedding space where background and foreground points are pushed apart, and spatially aligned points across views are pulled together. Our intuition is the resulting semantically discriminative representations invariant to irrelevant changes (illumination and unconcerned land covers) may help change recognition. We collect large-scale image-mask pairs freely available in the RS community for pretraining. Extensive experiments on three CD datasets verify the effectiveness of our method. Ours significantly outperforms ImageNet pretraining, in-domain supervision, and several SSL methods. Empirical results indicate our pretraining improves the generalization and data efficiency of the CD model. Notably, we achieve competitive results using 20% training data than baseline (random initialization) using 100% data. Our code is available athttps://github.com/justchenhao/SaDL_CD.
Hao Chen 0045, Wenyuan Li 0002, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Adversarial Instance Augmentation for Building Change Detection in Remote Sensing Images
abstract
Training deep learning-based change detection (CD) models heavily relies on large labeled data sets. However, it is time-consuming and labor-intensive to collect large-scale bitemporal images that contain building change, due to both its rarity and sparsity. Contemporary methods to tackle the data insufficiency mainly focus on transformation-based global image augmentation and cost-sensitive algorithms. In this article, we propose a novel data-level solution, namely, Instance-level change Augmentation (IAug), to generate bitemporal images that contain changes involving plenty and diverse buildings by leveraging generative adversarial training. The key of IAug is to blend synthesized building instances onto appropriate positions of one of the bitemporal images. To achieve this, a building generator is employed to produce realistic building images that are consistent with the given layouts. Diverse styles are later transferred onto the generated images. We further propose context-aware blending for a realistic composite of the building and the background. We augment the existing CD data sets and also design a simple yet effective CD model—CD network (CDNet). Our method (CDNet + IAug) has achieved state-of-the-art results in two building CD data sets (LEVIR-CD and WHU-CD). Interestingly, we achieve comparable results with only 20% of the training data as the current state-of-the-art methods using 100% data. Extensive experiments have validated the effectiveness of the proposed IAug. Our augmented data set has a lower risk of class imbalance than the original one. Conventional learning on the synthesized data set outperforms several popular cost-sensitive algorithms on the original data set. Our code and data are available athttps://github.com/justchenhao/IAug_CDNet.
Hao Chen 0045, Wenyuan Li 0002, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Remote Sensing Image Change Detection With Transformers
abstract
Modern change detection (CD) has achieved remarkable success by the powerful discriminative ability of deep convolutions. However, high-resolution remote sensing CD remains challenging due to the complexity of objects in the scene. Objects with the same semantic concept may show distinct spectral characteristics at different times and spatial locations. Most recent CD pipelines using pure convolutions are still struggling to relate long-range concepts in space-time. Nonlocal self-attention approaches show promising performance via modeling dense relationships among pixels, yet are computationally inefficient. Here, we propose a bitemporal image transformer (BIT) to efficiently and effectively model contexts within the spatial-temporal domain. Our intuition is that the high-level concepts of the change of interest can be represented by a few visual words, that is, semantic tokens. To achieve this, we express the bitemporal image into a few tokens and use a transformer encoder to model contexts in the compact token-based space-time. The learned context-rich tokens are then fed back to the pixel-space for refining the original features via a transformer decoder. We incorporate BIT in a deep feature differencing-based CD framework. Extensive experiments on three CD datasets demonstrate the effectiveness and efficiency of the proposed method. Notably, our BIT-based model significantly outperforms the purely convolutional baseline using only three times lower computational costs and model parameters. Based on a naive backbone (ResNet18) without sophisticated structures (e.g., feature pyramid network (FPN) and UNet), our model surpasses several state-of-the-art CD methods, including better than four recent attention-based methods in terms of efficiency and accuracy. Our code is available athttps://github.com/justchenhao/BIT_CD.
Hao Chen 0045, Zipeng Qi, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Hybrid-Scale Self-Similarity Exploitation for Remote Sensing Image Super-Resolution
abstract
Recently, deep convolutional neural networks (CNNs) have made great progress in remote sensing image super-resolution (SR). The CNN-based methods can learn powerful feature representation from plenty of low- and high-resolution counterparts. For remote sensing images, there are many similar ground targets recurred inside the image itself, both within the same scale and across different scales. In this article, we argue that this internal recurrence can be used for learning stronger feature representation, and we propose a new hybrid-scale self-similarity exploitation network (HSENet) for remote sensing image SR. Specifically, we introduce a single-scale self-similarity exploitation module (SSEM) to compute the feature correlation within the same scale image. Moreover, we design a cross-scale connection structure (CCS) to capture the recurrences across different scales. By combining SSEM and CCS, we further develop a hybrid-scale self-similarity exploitation module (HSEM) to construct the final HSENet, which simultaneously exploits single- and cross-scale similarities. Experimental results demonstrate that HSENet can obtain superior performance over several state-of-the-art methods. Besides, the effectiveness of our method is also verified by the assistance to the remote sensing scene classification task.
Sen Lei, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Transformer-Based Multistage Enhancement for Remote Sensing Image Super-Resolution
abstract
Convolutional neural networks have made a great breakthrough in recent remote sensing image super-resolution (SR) tasks. Most of these methods adopt upsampling layers at the end of the models to perform enlargement, which ignores feature extraction in the high-dimension space, and thus, limits SR performance. To address this problem, we propose a new SR framework for remote sensing images to enhance the high-dimensional feature representation after the upsampling layers. We name the proposed method as a transformer-based enhancement network (TransENet), where transformers are introduced to exploit features at different levels. The core of the TransENet is a transformer-based multistage enhancement structure, which can be combined with traditional SR frameworks to fuse multiscale high-/low-dimension features. Specifically, in this structure, the encoders aim to embed the multilevel features in the feature extraction part and the decoders are used to fuse these encoded embeddings. Experimental results demonstrate that our proposed TransENet can improve super-resolved results and obtain superior performance over several state-of-the-art methods.
Sen Lei, Zhenwei Shi 0001, Wenjing Mo
IEEE Trans. Geosci. Remote. Sens.2
2022 Geographical Knowledge-Driven Representation Learning for Remote Sensing Images
abstract
The proliferation of remote sensing satellites has resulted in a massive amount of remote sensing images. However, due to human and material resource constraints, the vast majority of remote sensing images remain unlabeled. As a result, it cannot be applied to currently available deep learning methods. To fully utilize the remaining unlabeled images, we propose a Geographical Knowledge-driven Representation learning method for remote sensing images (GeoKR), improving network performance and reduce the demand for annotated data. The global land cover products and geographical location associated with each remote sensing image are regarded as geographical knowledge to provide supervision for representation learning and network pre-training. An efficient pre-training framework is proposed to eliminate the supervision noises caused by imaging times and resolutions difference between remote sensing images and geographical knowledge. A large scale pre-training dataset Levir-KR is proposed to support network pre-training. It contains 1,431,950 remote sensing images from Gaofen series satellites with various resolutions. Experimental results demonstrate that our proposed method outperforms ImageNet pre-training and self-supervised representation learning methods and significantly reduces the burden of data annotation on downstream tasks such as scene classification, semantic segmentation, object detection, and cloud / snow detection. It demonstrates that our proposed method can be used as a novel paradigm for pre-training neural networks. Codes will be available on https://github.com/flyakon/Geographical-Knowledge-driven-Representaion-Learning.
Wenyuan Li 0002, Keyan Chen 0001, Hao Chen 0045, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Geographical Supervision Correction for Remote Sensing Representation Learning
abstract
Global land cover (GLC) products can be utilized to provide geographical supervision for remote sensing representation learning, which has significantly improved downstream tasks’ performance and decreased the demand of manual annotations. However, the time differences between remote sensing images and GLC products may introduce deviations in geographical supervision. In this paper, we propose a Geographical supervision Correction method (GeCo) for remote sensing representation learning. Deviated geographical supervision generated by GLC products can be corrected adaptively using the correction matrix during network pre-training and joint optimization process is designed to simultaneously update the correction matrix and network parameters. Additionally, we identify prior knowledge on geographical supervision to guide representation learning and restrict the correction process. The prior knowledge named “minor changes” implies that the geographical supervision may not change significantly, whereas the prior knowledge named “spatial aggregation” implies that land covers are aggregated in their spatial distribution. According to the prior knowledge, corresponding regularization terms are proposed to prevent abrupt changes in geographical supervision correction process and excessive smoothing of network outputs, thereby ensuring the adaptive correction process’s correctness. Experimental results demonstrate that our proposed method outperforms random initialization, ImageNet pre-training, and other representation learning methods on a variety of downstream tasks. In particular, when compared to the method that learns representations directly from deviated geographical supervision, it is proved that our method can eliminate the influence of deviations and further improve the effect of representation learning.
Wenyuan Li 0002, Keyan Chen 0001, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Physics-Informed Hyperspectral Remote Sensing Image Synthesis With Deep Conditional Generative Adversarial Networks
abstract
High-resolution hyperspectral remote sensing images are of great significance to agricultural, urban, and military applications. However, collecting and labeling hyperspectral images are time-consuming, expensive, and usually heavily rely on domain knowledge. In this article, we propose a new method for generating high-resolution hyperspectral images and subpixel ground-truth annotations from RGB images. Given a single high-resolution RGB image as its conditional input, unlike previous methods that directly predict spectral reflectance and ignores the physics behind it, we consider both imaging mechanism and spectral mixing, introduce a deep generative network that first recovers the spectral abundance for each pixel, and then generate the final spectral data cube with the standard USGS spectral library. In this way, our method not only synthesizes high-quality spectral data existing in the real world but also generates subpixel-level spectral abundance with well-defined spectral reflectance characteristics. We also introduce a spatial discriminative network and a spectral discriminative network to improve the fidelity of the synthetic output from both spatial and spectral perspectives. The whole framework can be trained end-to-end in an adversarial training paradigm. We refer to our method as “Physics-informed Deep Adversarial Spectral Synthesis (PDASS).” On the IEEEgrss_dfc_2018dataset, our method achieves an MPSNR of 47.56 on spectral reconstruction accuracy and outperforms other state-of-the-art methods. As latent variables, the generated spectral abundance and the atmospheric absorption coefficients of sunlight also suggest the effectiveness of our method.
Liqin Liu, Wenyuan Li 0002, Zhenwei Shi 0001, Zhengxia Zou
IEEE Trans. Geosci. Remote. Sens.3
2022 Remote Sensing Image Change Captioning With Dual-Branch Transformers: A New Method and a Large Scale Dataset
abstract
Analyzing land cover changes with multi-temporal remote sensing (RS) images is crucial for environmental protection and land planning. In this paper, we explore Remote Sensing Image Change Captioning (RSICC), a new task aiming at generating human-like language descriptions for the land cover changes in multi-temporal RS images. We propose a novel Transformer-based RSICC model (RSICCformer). It consists of three main components: 1) a CNN-based feature extractor to generate high-level features of RS image pairs, 2) a dual-branch Transformer encoder to improve the feature discrimination capacity for the changes, and 3) a caption decoder to generate sentences describing the differences. The dual-branch Transformer encoder consists of a hierarchy of processing stages to capture and recognize multiple changes of interest. Concretely, we use the bi-temporal feature differences as keys to enhance image features (queries) from each temporal image in the dual-branch Transformer encoder. To explore the RSICC task, we build a large-scale dataset named LEVIR-CC, which contains 10077 pairs of bi-temporal RS images and 50385 sentences describing the differences between images. We benchmark existing state-of-the-art synthetic image change captioning methods on the LEVIR-CC dataset, and our RSICCformer outperforms previous methods with a significant margin (+4.98% on BLEU-4 and +9.86% on CIDEr-D). The attention visualization results also suggest that our model can focus on changes of interest and ignore irrelevant changes.
Rui Zhao 0019, Hao Chen 0045, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Large-Factor Super-Resolution of Remote Sensing Images With Spectra-Guided Generative Adversarial Networks
Yapeng Meng, Wenyuan Li 0002, Sen Lei, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Structure-Color Preserving Network for Hyperspectral Image Super-Resolution
abstract
Fusion-based hyperspectral super-resolution (HSR) algorithms usually utilize a low-resolution hyperspectral image (LR-HSI) and a high-resolution multispectral image (MSI) to generate a high-resolution hyperspectral image (HR-HSI), which have attracted increasing attention in recent years. However, how to deal with the abundant spectral information of hyperspectral images and complex structure characteristics of MSIs has always been the focus and difficulty of fusion-based HSR. In this article, we propose a new structure–color preserving network (SCPNet) for HSR, which is developed under the basis of the joint attention mechanism. The SCPNet mainly includes three modules: structure-preserving module (SPM), color-preserving module (CPM), and cross-fusion module. The SPM is constructed based on the spatial attention, which aims to capture and enhance the significant structure information from the high-resolution MSI. Meanwhile, the CPM is constructed based on the channel attention, where the spectral characteristics in the LR-HSI are preserved during the reconstruction process. Finally, we propose a cross attention-based cross-fusion strategy to integrate the features from the two branches and reconstruct the final HR-HSI. The major contribution of SCPNet is that the structure and color information is described and preserved via the joint attention mechanism. Experimental results indicate that the proposed SCPNet has presented advantages on three benchmark datasets when compared with some state-of-the-art HSR methods.
Bin Pan, Qiaoying Qu, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 CANet: Centerness-Aware Network for Object Detection in Remote Sensing Images
abstract
Recently, feature pyramid has been widely exploited in remote sensing detectors, which greatly alleviates the problem arising from scale variation across objects in remote sensing images. However, these object detectors with feature pyramid give insufficient consideration that objects in remote sensing images usually maintain symmetrical shape. To address this issue, we propose an anchor-free-based detector called Centerness-Aware Network (CANet), which could capture the symmetrical shape of objects in remote sensing images. The kernel structure of CANet is a new Centerness-Aware Model (CAM) that contains three components: Multiscale Centerness Descriptor (MSCD), Centerness Detection Head (CDH), and Feature Selective Module (FSM). Considering that symmetrical objects will maintain a rigid appearance around their center region, three components are integrated into the feature pyramid to extract and utilize the features around the center region. More precisely, the MSCD is embedded into the feature pyramid and highlights the center of current objects through the attention mechanism. Guided by the MSCD, the CDH could accurately capture the center of objects by per-pixel prediction. Furthermore, the FSM is connected to the CDH, which guides the CDH to adaptively select the optimal feature level from the pyramidal features. The selected feature level could describe the best semantic information around the center region, which helps the network progressively fit the symmetrical shape of remote sensing objects. Besides, we also design the hybrid loss function to effectively train CAM in the end-to-end way. The experiments show that our network is competitive with some state-of-the-art detection networks.
Lukui Shi, Linyi Kuang, Bin Pan, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Remote Sensing Novel View Synthesis With Implicit Multiplane Representations
abstract
Novel view synthesis of remote sensing scenes is of great significance for scene visualization, human-computer interaction, and various downstream applications. Despite the recent advances in computer graphics and photogrammetry technology, generating novel views is still challenging particularly for remote sensing images due to its high complexity, view sparsity and limited view-perspective variations. In this paper, we propose a novel remote sensing view synthesis method by leveraging the recent advances in implicit neural representations. Considering the overhead and far depth imaging of remote sensing images, we represent the 3D space by combining implicit multiplane images (MPI) representation and deep neural networks. The 3D scene is reconstructed under a self-supervised optimization paradigm through a differentiable multiplane renderer with multi-view input constraints. Images from any novel views thus can be freely rendered on the basis of the reconstructed model. As a by-product, the depth maps corresponding to the given viewpoint can be generated along with the rendering output. We refer to our method as Implicit Multiplane Images (ImMPI). To further improve the view synthesis under sparse-view inputs, we explore the learning-based initialization of remote sensing 3D scenes and proposed a neural network based Prior extractor to accelerate the optimization process. In addition, we propose a new dataset for remote sensing novel view synthesis with multi-view real-world google earth images. Extensive experiments demonstrate the superiority of the ImMPI over previous state-of-the-art methods in terms of reconstruction accuracy, visual fidelity, and time efficiency. Ablation experiments also suggest the effectiveness of our methodology design.
Yongchang Wu, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Deep Autoencoder for Hyperspectral Unmixing via Global-Local Smoothing
abstract
Hyperspectral unmixing is to decompose the mixed pixels into pure spectral signatures (endmembers) and their proportions (abundances). Recently, deep learning-based methods have been applied to enhance the representation ability of unmixing models by extracting joint spatial–spectral characteristics of the hyperspectral data. However, most deep learning based-unmixing methods usually conduct global smoothing by convolutions on the whole hyperspectral imagery, which may ignore the variations within the imagery and result in oversmoothing. In this article, we propose a deep network for hyperspectral unmixing based on a new global–local smoothing autoencoder (GLA). GLA is an unsupervised model, which aims at exploring the local homogeneity and the global self-similarity of hyperspectral imagery. The proposed GLA network mainly includes two modules: a Local Continuous conditional random field Smoothing (LCS) module and a global recurrent smoothing (GRS) module. In LCS, we propose a conditional random field-based smoothing strategy to describe the joint spatial–spectral information within a local homogeneity region, which also reduces the risk of abundance maps boundary blurry. In GRS, we follow the self-similarity assumption for hyperspectral imagery and develop a recurrent neural network structure to exploit potential long-distance dependency relationships among pixels. The GLA is compared with several state-of-the-art unmixing methods on both real and synthetic data, and the abundance estimation results indicate that our method is promising. We will publish the code of GLA if this article has the honor to be accepted.
Xinyu Song 0004, Tao Li 0022, Zhenwei Shi 0001, Bin Pan
IEEE Trans. Geosci. Remote. Sens.4
2022 Hierarchical Similarity Alignment for Domain Adaptive Ship Detection in SAR Images
abstract
Ship detection from synthetic aperture radar (SAR) images is a hot topic, but the difficulty in collecting labeled SAR images may hinder the development of deep-learning-based detection methods. Inspired by the idea of domain adaptation, in this article, we propose a hierarchical similarity alignment neural network (HSANet) for ship detection in SAR images, which is a domain adaptive (DA) approach with optical remote sensing images as training samples. The kernel target of HSANet is to mine and align both the global structure and the local instance information between SAR and optical images, where two modules, structural alignment module (SAM) and prototype alignment module (PAM), are designed to, respectively, conduct two hierarchies of alignment process. In general, SAM attempts to extract the global structure similarity which exists in image-level feature representation, while PAM tends to extract the local shape similarity which is instance-level representation. To be specific, SAM is developed by Fourier-based feature alignment, which tries to describe the similar structural relationship between optical and SAR images. Meanwhile, PAM is proposed based on the conjoint confidence analysis where the instance-level ship representations of the source and target domains is aligned. SAM and PAM work together to construct a hierarchical domain adaptation network for SAR ship detection. Experiments on several public datasets may indicate the effectiveness of the proposed method.
Jun Zhang 0050, Yongfeng Dong, Bin Pan, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 An Open Set Domain Adaptation Algorithm via Exploring Transferability and Discriminability for Remote Sensing Image Scene Classification
abstract
Remote sensing image scene classification aims to automatically assign semantic labels for remote sensing images. Recently, to overcome the distribution discrepancy of training data and test data, domain adaptation has been applied to remote sensing image scene classification. Most domain adaptation approaches usually explore transferability under the assumption that the source domain and target domain have common classes. However, in real applications, new categories may appear in the target domain. Besides, only considering the transferability will degrade the classification performance due to the strong interclass similarity of remote sensing images. In this article, we present an open set domain adaptation algorithm via exploring transferability and discriminability (OSDA-ETD) for remote sensing image scene classification. To be specific, we propose the transferability technology, which aims at the high interdomain variations and high intraclass diversity of remote sensing images. The purpose of transferability is to reduce the global distribution difference of domains and the local distribution discrepancy of the same classes in different domains. For high interclass similarity in remote sensing images, we adopt the discriminability strategy. The discriminability intends to enlarge the distribution discrepancy of different classes in different domains. To further promote the effectiveness of scene classification, we integrate the transferability and the discriminability into a framework. Moreover, we prove that the algorithm has a unique optimizer.
Jun Zhang 0050, Jiao Liu 0003, Bin Pan, Herman Z. Q. Chen, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 High-Resolution Remote Sensing Image Captioning Based on Structured Attention
abstract
Automatically generating language descriptions of remote sensing images has become an emerging research hot spot in the remote sensing field. Attention-based captioning, as a representative group of recent deep learning-based captioning methods, shares the advantage of generating the words while highlighting corresponding object locations in the image. Standard attention-based methods generate captions based on coarse-grained and unstructured attention units, which fails to exploit structured spatial relations of semantic contents in remote sensing images. Although the structure characteristic makes remote sensing images widely divergent to natural images and poses a greater challenge for the remote sensing image captioning task, the key of most remote sensing captioning methods is usually borrowed from the computer vision community without considering the domain knowledge behind. To overcome this problem, a fine-grained, structured attention-based method is proposed to utilize the structural characteristics of semantic contents in high-resolution remote sensing images. Our method learns better descriptions and can generate pixelwise segmentation masks of semantic contents. The segmentation can be jointly trained with the captioning in a unified framework without requiring any pixelwise annotations. Evaluations are conducted on three remote sensing image captioning benchmark data sets with detailed ablation studies and parameter analysis. Compared with the state-of-the-art methods, our method achieves higher captioning accuracy and can generate high-resolution and meaningful segmentation masks of semantic contents at the same time.
Rui Zhao 0019, Zhenwei Shi 0001, Zhengxia Zou
IEEE Trans. Geosci. Remote. Sens.2
2022 Castle in the Sky: Dynamic Sky Replacement and Harmonization in Videos
abstract
We propose a vision-based framework for dynamic sky replacement and harmonization in videos. Different from previous sky editing methods that either focus on static photos or require real-time pose signal from the camera's inertial measurement units, our method is purely vision-based, without any requirements on the capturing devices, and can be well applied to either online or offline processing scenarios. Our method runs in real-time and is free of manual interactions. We decompose the video sky replacement into several proxy tasks, including motion estimation, sky matting, and image blending. We derive the motion equation of an object at infinity on the image plane under the camera's motion, and propose "flow propagation", a novel method for robust motion estimation. We also propose a coarse-to-fine sky matting network to predict accurate sky matte and design image blending to improve the harmonization. Experiments are conducted on videos diversely captured in the wild and show high fidelity and good generalization capability of our framework in both visual quality and lighting/motion dynamics. We also introduce a new method for content-aware image augmentation and proved that this method is beneficial to visual perception in autonomous driving scenarios. Our code and animated results are available at https://github.com/jiupinjia/SkyAR.
Zhengxia Zou, Rui Zhao 0019, Tianyang Shi, Zhenwei Shi 0001
IEEE Trans. Image Process.5
2021 Stylized Neural Painting
abstract
This paper proposes an image-to-painting translation method that generates vivid and realistic painting artworks with controllable styles. Different from previous image-to-image translation methods that formulate the translation as pixel-wise prediction, we deal with such an artistic creation process in a vectorized environment and produce a sequence of physically meaningful stroke parameters that can be further used for rendering. Since a typical vector render is not differentiable, we design a novel neural renderer which imitates the behavior of the vector renderer and then frame the stroke prediction as a parameter searching process that maximizes the similarity between the input and the rendering output. We explored the zero-gradient problem on parameter searching and propose to solve this problem from an optimal transportation perspective. We also show that previous neural renderers have a parameter coupling problem and we re-design the rendering network with a rasterization network and a shading network that better handles the disentanglement of shape and color. Experiments show that the paintings generated by our method have a high degree of fidelity in both global appearance and local textures. Our method can be also jointly optimized with neural style transfer that further transfers visual style from other images. Our code and animated results are available at https://jiupinjia.github.io/neuralpainter/.
Zhengxia Zou, Tianyang Shi, Yi Yuan 0002, Zhenwei Shi 0001
CVPR5
2021 V2RNet: An Unsupervised Semantic Segmentation Algorithm for Remote Sensing Images via Cross-Domain Transfer Learning
abstract
The dependence on large-scale pixel-level annotations brings great challenge to semantic segmentation task for remote sensing images (RSIs). To alleviate this issue, we propose V2RNet, an unsupervised semantic segmentation method which introduces adversarial learning into segmentation network. Our method creatively transfers the segmentation model from the synthetic GTA-V data to the real optical remote sensing data via domain adaptation. Additionally, to unify the source domain semantic structures and target domain image style, we design a semantic segmentation discriminator as auxiliary to optimize the domain adaptation efficiency. Thus the proposed method is effective on typical remote sensing targets such densely arranged, intertwined road. Experimental results on Massachusetts Road data set demonstrate our unsupervised semantic segmentation model achieves comparable segmentation accuracy, which also validates the effectiveness of the proposed method.
Danpei Zhao, Bo Yuan 0009, Zhenwei Shi 0001
IGARSS4
2021 Cross-Domain Transfer for Ship Instance Segmentation in SAR Images
abstract
Considering insufficient data and difficulty of labeling in Synthetic Aperture Radar (SAR) images, we propose a method for SAR ship instance segmentation based on cross-domain transfer learning. Compared with optical images, transfer learning in SAR images faces the difficulties of insufficient data to pre-train and lacking detail features. The proposed method, containing sample transfer module and knowledge transfer module, simulates images from optics to SAR and pre-train the ship detection part of the instance segmentation network with simulation images. In addition, we design a Res-Pyramid network to prevent the deep network from being unable to extract efficient features of SAR images. The method proposed combines the content of the optics and the style of the SAR and incorporates multiscale features in backbone, which improves performance in ship instance segmentation in SAR images. Experiments show that it has achieved 1.3 and 1.1 points higher Average Precision (AP) in detection and segmentation tasks on SAR dataset of HRSID when using cross-domain transfer learning, which has exceeded state-of-the-art methods.
Chunbo Zhu, Danpei Zhao, Xinhu Qi, Zhenwei Shi 0001
IGARSS5
2021 Selective focus saliency model driven by object class-awareness
abstract
Abstract Current many salient object detection (SOD) models only focus on highlighting visual conspicuous region but fail to make saliency detection for specific targets. In this paper, a selective focus saliency model driven by object class‐awareness (SF‐OCA) to run saliency detection is proposed. The framework consists of a visual saliency detection flow, a segmentation‐classification flow, and a class‐awareness selection module. It combines bottom‐up visual perception with a top‐down task‐driven manner, which is capable of detecting specific category salient targets and eliminating the interference from other saliency areas, providing a new idea for saliency detection. Experimental results show that the method achieves comparable performance with state‐of‐the‐art models on four public saliency datasets. In addition, a new dataset was also built to test the proposed framework for the selective focus saliency detection. Compared with other SOD methods, the method not only highlights visual saliency regions but can choose more important or more noteworthy targets in a class‐awareness manner. The method also shows better robustness under a variety of conditions including multi‐targets, small targets and complex background.
Danpei Zhao, Bo Yuan 0009, Zhenwei Shi 0001, Zhiguo Jiang 0001
IET Image Process.3
2021 Single-shot weakly-supervised object detection guided by empirical saliency model
Danpei Zhao, Zhichao Yuan, Zhenwei Shi 0001, Fengying Xie
Neurocomputing3
2021 Multiscale Methods for Optical Remote-Sensing Image Captioning
abstract
Recently, the optical remote-sensing image-captioning task has gradually become a research hotspot because of its application prospects in the military and civil fields. Many different methods along with data sets have been proposed. Among them, the models following the encoder–decoder framework have better performance in many aspects like generating more accurate and flexible sentences. However, almost all these methods are of a single fixed receptive field and could not put enough attention on grabbing the multiscale information, which leads to incomplete image representation. In this letter, we deal with the multiscale problem and propose two multiscale methods named multiscale attention (MSA) method and multifeat attention (MFA) method, to obtain better representations for the captioning task in the remote-sensing field. The MSA method extracts features from different layers and uses the multihead attention mechanism to obtain the context feature, respectively. The MFA method combines the target-level features and the scene-level features by using the target-detection task as the auxiliary task to enrich the context feature. The experimental results demonstrate that both of them perform better with regard to the metrics like BLEU, METEOR, ROUGE_L, and CIDEr than the benchmark method.
Rui Zhao 0019, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.3
2021 An End-to-End Network for Remote Sensing Imagery Semantic Segmentation via Joint Pixel- and Representation-Level Domain Adaptation
abstract
It requires pixel-by-pixel annotations to obtain sufficient training data in supervised remote sensing image segmentation, which is a quite time-consuming process. In recent years, a series of domain-adaptation methods was developed for image semantic segmentation. In general, these methods are trained on the source domain and then validated on the target domain to avoid labeling new data repeatedly. However, most domain-adaptation algorithms only tried to align the source domain and the target domain in the pixel level or the representation level, while ignored their cooperation. In this letter, we propose an unsupervised domain-adaptation method by Joint Pixel and Representation level Network (JPRNet) alignment. The major novelty of the JPRNet is that it achieves joint domain adaptation in an end-to-end manner, so as to avoid the multisource problem in the remote sensing images. JPRNet is composed of two branches, each of which is a generative-adversarial network (GAN). In one branch, pixel-level domain adaptation is implemented by the style transfer with the Cycle GAN, which could transfer the source domain to a target domain. In the other branch, the representation-level domain adaptation is realized by adversarial learning between the transferred source-domain images and the target-domain images. The experimental results on the public data sets have indicated the effectiveness of the JPRNet.
Lukui Shi, Bin Pan, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.4
2021 DCL-Net: Augmenting the Capability of Classification and Localization for Remote Sensing Object Detection
abstract
Deep learning-based remote sensing object detectors are usually composed of two branches: classification and localization. Recently proposed object detectors often follow the pipeline that classification and localization branches share the same feature maps, which leads to a strong coupling relationship between them. However, when tackling remote sensing images, this strong coupling relationship may impair the performance of the detectors because the top-view perspective of remote sensing images may result in conflicts between classification and location branches. To address this issue, we propose a decoupled classification localization network (DCL-Net) by considering the different characteristics between the two branches. Two modules are developed to suppress the strong coupling: receptive field aggregation module (RFAM) and bottom-up path aggregation module (PAM). For the classification branch, RFAM can learn the relationship between objects and context information by simulating the human receptive field and improve the robustness of the classification branch to rotational distortions. For the localization branch, PAM can enhance the entire feature hierarchy by transferring the rich detailed information of low-level features, which helps the detector to achieve precise bounding box regression. Compared with existing methods, the major contribution of DCL-Net is that the independence of the classification and localization branches can be significantly enhanced, which may be beneficial to the detection accuracy for the objects in remote sensing images. Experiments on public data sets validate the effectiveness of our detector.
Enhai Liu, Yu Zheng 0031, Bin Pan, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2021 Simultaneously Multiobjective Sparse Unmixing and Library Pruning for Hyperspectral Imagery
abstract
Sparse hyperspectral unmixing has attracted increasing investigations during the past decade. Recent research has indicated that library pruning algorithms can significantly improve the unmixing accuracies by reducing the mutual coherence of the spectral library. Inspired by the good performance of library pruning, in this article we propose a new hyperspectral unmixing algorithm which integrates the idea of library pruning and sparse representation. An obvious challenge for pruning algorithms is that the real endmembers must be preserved after pruning. Unfortunately, recent proposed pruning algorithms, such as multiple signal classification are actually prepruning strategies, which cannot guarantee that the endmembers exactly exist in the selected spectral subset when the image noise is strong. To overcome this difficulty, we develop a simultaneous optimization approach which involves the pruning operation into the optimization process. Compared with existing prepruning-based unmixing methods, the proposed algorithm can gradually compress the search space of sparse representation, which may relieve the loss of spectral information caused by the rapid compression of the library. Instead of simply designing a regularizer, in this article we utilize a multiobjective-based framework where reconstruction error, sparsity error, and the pruning projection function are considered as three parallel objectives, so as to avoid the manually settings of regularization parameters. Moreover, we have provided theoretical analysis and proof for the reasonability of our pruning objective. Experiments on synthetic hyperspectral data may indicate the superiority of the proposed method under high-noise conditions.
Bin Pan, Herman Z. Q. Chen, Zhenwei Shi 0001, Tao Li 0022
IEEE Trans. Geosci. Remote. Sens.4
2021 Adversarial Training for Solving Inverse Problems in Image Processing
abstract
Inverse problems are a group of important mathematical problems that aim at estimating source data x and operation parameters z from inadequate observations y . In the image processing field, most recent deep learning-based methods simply deal with such problems under a pixel-wise regression framework (from y to x ) while ignoring the physics behind. In this paper, we re-examine these problems under a different viewpoint and propose a novel framework for solving certain types of inverse problems in image processing. Instead of predicting x directly from y , we train a deep neural network to estimate the degradation parameters z under an adversarial training paradigm. We show that if the degradation behind satisfies some certain assumptions, the solution to the problem can be improved by introducing additional adversarial constraints to the parameter space and the training may not even require pair-wise supervision. In our experiment, we apply our method to a variety of real-world problems, including image denoising, image deraining, image shadow removal, non-uniform illumination correction, and underdetermined blind source separation of images or speech signals. The results on multiple tasks demonstrate the effectiveness of our method.
Zhengxia Zou, Tianyang Shi, Zhenwei Shi 0001, Jieping Ye
IEEE Trans. Image Process.3
2020 Deep Adversarial Decomposition: A Unified Framework for Separating Superimposed Images
abstract
Separating individual image layers from a single mixed image has long been an important but challenging task. We propose a unified framework named "deep adversarial decomposition" for single superimposed image separation. Our method deals with both linear and non-linear mixtures under an adversarial training paradigm. Considering the layer separating ambiguity that given a single mixed input, there could be an infinite number of possible solutions, we introduce a "Separation-Critic" - a discriminative network which is trained to identify whether the output layers are well-separated and thus further improves the layer separation. We also introduce a "crossroad L1" loss function, which computes the distance between the unordered outputs and their references in a crossover manner so that the training can be well-instructed with pixel-wise supervision. Experimental results suggest that our method significantly outperforms other popular image separation frameworks. Without specific tuning, our method achieves the state of the art results on multiple computer vision tasks, including the image deraining, photo reflection removal, and image shadow removal.
Zhengxia Zou, Sen Lei, Tianyang Shi, Zhenwei Shi 0001, Jieping Ye
CVPR4
2020 DSSNet: A Simple Dilated Semantic Segmentation Network for Hyperspectral Imagery Classification
abstract
Deep learning-based methods have presented a promising performance in the task of hyperspectral imagery classification (HSIC). However, recent methods usually are considered HSIC as a patchwise image classification problem and addressed it by giving a single label to the patch surrounding a pixel. In this letter, we propose a new semantic segmentation network that can directly label each pixel in an end-to-end manner. Compared with patchwise models, our method can significantly improve training effectiveness and reduce some manual parameters. Another challenge in HSIC is that the spatial resolution of hyperspectral imagery is relatively low; in that case, the pooling operation may result in resolution and coverage loss. To address this issue, we introduce dilated convolution to our model and construct a dilated semantic segmentation network (DSSNet). Different from some existing works, DSSNet is specially designed for HSIC without complicated architecture, and no pretrained models are required. The joint spatial-spectral information can be extracted via an end-to-end manner and, thus, avoid various preprocessing or postprocessing operations. Experiments on two public data sets have demonstrated the effectiveness of our improvements compared with some of the latest deep learning-based HSIC models.
Bin Pan, Zhenwei Shi 0001, Huanlin Luo, Xianchao Lan
IEEE Geosci. Remote. Sens. Lett.3
2020 Local Attention Networks for Occluded Airplane Detection in Remote Sensing Images
abstract
Despite the great progress of deep learning and target detection in recent years, the accurate detection of the occluded targets in remote sensing images still remains a challenge. In this letter, we propose a new detection method called local attention networks to improve the detection of occluded airplanes. Following the idea of “divide and conquer,” the proposed method is designed by first dividing an airplane target into four visual parts: head, left/right wings, body, and tail, and then considering the detection as the prediction of the individual key points in each of the visual parts. We further introduce an additional attention branch in the standard detection pipeline to enhance the features and make the model focus on individual parts of a target even if it is only partially visible in the image. Detection results and ablation studies on three remote sensing target detection data sets (including two publicly available ones) demonstrate the effectiveness of our method, especially for occluded airplane targets. In addition, our method outperforms the other state-of-the-art detection methods on these data sets.
Zhengxia Zou, Zhenwei Shi 0001, Wen-Jun Zeng, Jie Gui
IEEE Geosci. Remote. Sens. Lett.3
2020 Coupled Adversarial Training for Remote Sensing Image Super-Resolution
abstract
Generative adversarial network (GAN) has made great progress in recent natural image super-resolution tasks. The key to its success is the integration of a discriminator which is trained to classify whether the input is a real high-resolution (HR) image or a generated one. Arguably, learning a strong discriminative prior is essential for generating high-quality images. However, in remote sensing images, we discover, through extensive statistical analysis, that there are more low-frequency components than natural images, which may lead to a “discrimination-ambiguity” problem, i.e., the discriminator will become “confused” to tell whether its input is real or not when dealing with those low-frequency regions, and therefore, the quality of generated HR images may be deeply affected. To address this problem, we propose a novel GAN-based super-resolution algorithm named coupled-discriminated GANs (CDGANs) for remote sensing images. Different from the previous GAN-based super-resolution models in which their discriminator takes in a single image at one time, in our model, the discriminator is specifically designed to take in a pair of images: a generated image and its HR ground truth, to make better discrimination of the inputs. We further introduce a dual pathway network architecture, a random gate, and a coupled adversarial loss to learn better correspondence between the discriminative results and the paired inputs. Experimental results on two public data sets demonstrate that our model can obtain more accurate super-resolution results in terms of both visual appearance and local details compared with other state of the arts. Our code will be made publicly available.
Sen Lei, Zhenwei Shi 0001, Zhengxia Zou
IEEE Trans. Geosci. Remote. Sens.2
2020 Deep Matting for Cloud Detection in Remote Sensing Images
abstract
Cloud detection, as an important preprocessing operation for remote sensing (RS) image analysis, has received increasing attention in recent years. Most of the previous cloud detection methods consider the detection as a pixel-wise image classification problem (cloud versus background), which inevitably leads to a category-ambiguity when dealing with the detection of thin clouds. In this article, starting from the RS imaging mechanism on cloud images, we re-examine the cloud detection under a totally different point of view, i.e., to formulate cloud detection as a mixed energy separation between foreground and background images. This process can be further equivalently implemented under a deep learning-based image matting framework with a clear physical significance. More importantly, the proposed method is capable to deal with three different but related tasks, i.e., “cloud detection,” “cloud removal,” and “cloud cover assessment,” under a unified framework. The experimental results on the three satellite image data sets demonstrate the effectiveness of our method, especially for those hard but common examples in RS images, such as the thin and wispy cloud.
Wenyuan Li 0002, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2020 Domain Adaptation Based on Correlation Subspace Dynamic Distribution Alignment for Remote Sensing Image Scene Classification
abstract
Remote sensing image scene classification refers to assigning semantic labels according to the content of the remote sensing scenes. Most machine learning-based scene classification methods assume that training and testing data share the same distributions. However, in real application scenarios, this assumption is difficult to guarantee. Domain adaptation (DA) is a promising approach to address this problem by aligning the feature distribution of training and testing data. Inspired by the idea DA, in this article, we propose a correlation subspace dynamic distribution alignment (CS-DDA) method for remote sensing image scene classification. Aiming at the characteristics of remote sensing scenes, we introduce two strategies to balance the effects of source and target domains: subspace correlation maximization (SCM) and dynamic statistical distribution alignment (DSDA). On the one hand, SCM tries to avoid mapping source domain data into irrelevant subspace to preserve the representation information of the source domain. On the other hand, DSDA is proposed to reduce the data distribution discrepancy between aligned source and target domains. Specifically, DSDA is a dynamic adjustment process where an adaptive factor is learned to balance the interclass and intraclass distribution between domains. Moreover, we integrate SCM and DSDA into a uniform optimization framework, and the optimal solution can be converted to the generalized eigendecomposition problem by derivation. The experimental results indicate that the proposed method can generate better results when compared with other feature distribution alignment methods.
Jun Zhang 0050, Jiao Liu 0003, Bin Pan, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.4
2019 Face-to-Parameter Translation for Game Character Auto-Creation
abstract
Character customization system is an important component in Role-Playing Games (RPGs), where players are allowed to edit the facial appearance of their in-game characters with their own preferences rather than using default templates. This paper proposes a method for automatically creating in-game characters of players according to an input face photo. We formulate the above "artistic creation" process under a facial similarity measurement and parameter searching paradigm by solving an optimization problem over a large set of physically meaningful facial parameters. To effectively minimize the distance between the created face and the real one, two loss functions, i.e. a "discriminative loss" and a "facial content loss", are specifically designed. As the rendering process of a game engine is not differentiable, a generative network is further introduced as an "imitator" to imitate the physical behavior of the game engine so that the proposed method can be implemented under a neural style transfer framework and the parameters can be optimized by gradient descent. Experimental results demonstrate that our method achieves a high degree of generation similarity between the input face photo and the created in-game character in terms of both global appearance and local details. Our method has been deployed in a new game last year and has now been used by players over 1 million times.
Tianyang Shi, Yi Yuan 0002, Changjie Fan, Zhengxia Zou, Zhenwei Shi 0001, Yong Liu 0007
ICCV5
2019 Generative Adversarial Training for Weakly Supervised Cloud Matting
abstract
The detection and removal of cloud in remote sensing images are essential for earth observation applications. Most previous methods consider cloud detection as a pixel-wise semantic segmentation process (cloud v.s. background), which inevitably leads to a category-ambiguity problem when dealing with semi-transparent clouds. We re-examine the cloud detection under a totally different point of view, i.e. to formulate it as a mixed energy separation process between foreground and background images, which can be equivalently implemented under an image matting paradigm with a clear physical significance. We further propose a generative adversarial framework where the training of our model neither requires any pixel-wise ground truth reference nor any additional user interactions. Our model consists of three networks, a cloud generator G, a cloud discriminator D, and a cloud matting network F, where G and D aim to generate realistic and physically meaningful cloud images by adversarial training, and F learns to predict the cloud reflectance and attenuation. Experimental results on a global set of satellite images demonstrate that our method, without ever using any pixel-wise ground truth during training, achieves comparable and even higher accuracy over other fully supervised methods, including some recent popular cloud detectors and some well-known semantic segmentation frameworks.
Zhengxia Zou, Wenyuan Li 0002, Tianyang Shi, Zhenwei Shi 0001, Jieping Ye
ICCV4
2019 Unsupervised Oil Tank Detection by Shape-Guide Saliency Model
abstract
In this letter, a novel oil tank detection framework based on a shape-guide saliency (SGS) model is proposed. Beyond the low-level visual stimuli, SGS focuses more on simulating the selective visual searching, which is dominated by the goal in human minds. Using a top–down strategy, SGS breaks the limitation of the low-level visual features and introduces the high-level task concept to measure saliency. For the oil tank detection, SGS model skillfully extracts the contour shape cue (CSC) as the target-oriented information and uses CSC to guide the selective saliency value calculation. Specifically, a sparse reconstruction with the target-specific dictionary is implemented to generate the saliency map. This saliency map only assigns high values to oil tank regions instead of highlighting all high-contrast regions. Consequently, SGS model is capable of accurately locating oil tanks and eliminating the interferences of high-contrast backgrounds. Experimental results on a remote sensing data set demonstrate that the proposed SGS model outperforms five class-independent saliency models. Comparisons with the state-of-the-art oil tank detection approaches demonstrate the effectiveness of the proposed method.
Minhao Jing, Danpei Zhao, Yue Gao 0008, Zhiguo Jiang 0001, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.6
2019 CoinNet: Copy Initialization Network for Multispectral Imagery Semantic Segmentation
abstract
Remote sensing imagery semantic segmentation refers to assigning a label to every pixel. Recently, deep convolutional neural networks (CNNs)-based methods have presented an impressive performance in this task. Due to the lack of sufficient labeled remote sensing images, researchers usually utilized transfer learning (TL) strategies to fine tune networks which were pretrained in huge RGB-scene data sets. Unfortunately, this manner may not work if the target images are multispectral/hyperspectral. The basic assumption of TL is that the low-level features extracted by the former layers are similar in most data sets, hence users only require to train the parameters in the last layers that are specific to different tasks. However, if one should use a pretrained deep model in RGB data for multispectral /hyperspectral imagery semantic segmentation, the structure of the input layer has to be adjusted. In this case, the first convolutional layer has to be trained using the multispectral /hyperspectral data sets which are much smaller. Apparently, the feature representation ability of the first convolutional layer will decrease and it may further harm the following layers. In this letter, we propose a new deep learning model, COpy INitialization Network (CoinNet), for multispectral imagery semantic segmentation. The major advantage of CoinNet is that it can make full use of the initial parameters in the pretrained network's first convolutional layer. Comparison experiments on a challenging multispectral data set have demonstrated the effectiveness of the proposed improvement. The demo and a trained network will be published in our homepage.
Bin Pan, Zhenwei Shi 0001, Tianyang Shi, Xinzhong Zhu
IEEE Geosci. Remote. Sens. Lett.2
2019 Multiobjective-Based Sparse Representation Classifier for Hyperspectral Imagery Using Limited Samples
abstract
Recent studies about hyperspectral imagery (HSI) classification usually focus on extracting more representative features or combining joint spectral-spatial information. However, besides feature extraction, developing more powerful classifiers can also contribute to the accuracies of HSI classification. In this paper, we propose a multiobjective-based sparse representation classifier (MSRC) for HSI data, which mainly tries to address two problems: 1) pixel mixing and 2) lacking abundant labeled samples. MSRC is motivated by the SRC, and further integrating the idea of hyperspectral unmixing. Different from the traditional SRC-based methods, the novelty of MSRC consists of the optimization process, i.e., we directly handle the L0-norm problem without any relaxation. The sparse term is not considered as a regularization operation. Instead, we transform the problem of weight vector estimation to subset selection, and propose a multiobjective-based method to optimize the L0-norm sparse problem. The residual term and sparse term are regarded as two parallel objective functions that are optimized simultaneously. We further utilize the linear mixing model to represent test pixels based on the selected atoms. The final class labels are determined according to the abundance estimation results by nonnegative least squares. Owing to the characteristics of the multiobjective method and the binary property of the sparse solution vector, MSRC does not require too many training samples to build the dictionary. Moreover, theoretically, MSRC can be easily improved to extended version such as combining spatial information.
Bin Pan, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2019 Analysis for the Weakly Pareto Optimum in Multiobjective-Based Hyperspectral Band Selection
abstract
Band selection refers to finding the most representative channels from hyperspectral images. Usually, certain objective functions are designed and combined via regularization terms. A possible drawback of these methods is that they can only generate one solution in a single run with a given band number. To overcome this problem, multiobjective (MO)-based methods, which were able to simultaneously obtain a series of subsets with different band numbers, were investigated for band selection. However, because the range of band selection problem is discrete, recently proposed weighted Tchebycheff (WT)-based MO methods may suffer weakly Pareto optimal problem. In this case, the solutions for each band number will be nonunique and no optimal solution exists. Decision makers have to manually select a unique solution for each band number. In this paper, we provide a theoretical analysis about the weakly Pareto optimal problem in band selection, and quantitatively give the boundary conditions. Moreover, we further summarize the suggestions which will help users avoid the weakly Pareto optimal problem. According to these criteria, we develop a new adaptive-penalty-based boundary intersection (APBI) framework to improve the MO algorithm in hyperspectral band selection. APBI mainly includes two advantages: 1) avoiding weakly Pareto optimum and 2) reducing the sensibility of the penalty factor. The theoretical analysis is further validated by contrast experiments. The results demonstrate that the weakly Pareto optimal solutions really exist in WT methods, while APBI can overcome this problem.
Bin Pan, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2019 A Classification-Based Model for Multi-Objective Hyperspectral Sparse Unmixing
abstract
Sparse unmixing has become a popular tool for hyperspectral imagery interpretation. It refers to finding the optimal subset of a spectral library to reconstruct the image data and further estimate the proportions of different materials. Recently, multi-objective based sparse unmixing methods have presented promising performance because of their advantages in addressing combinatorial problems. A spectral and multi-objective based sparse unmixing (SMoSU) algorithm was proposed in our previous work, which solves the decision-making problem well. However, it does not show outstanding advantages in strong noise cases. To solve the problem, in this paper, SMoSU is improved based on the estimation of distribution algorithms (EDAs). The machine learning based EDAs have been a reliable approach in solving multi-objective problems. However, most of them are for special problems and relatively weak in theoretical foundations. Thus, it is unreliable to extend it directly to sparse unmixing. Here, we improve EDA on the basis of classification and propose a classification-based model for individual generating under the framework of SMoSU (CM-MoSU). In CM-MoSU, the whole population is divided to be positive and negative. Then, the macroinformation of positive individuals is used to guide the generation of new individuals. Therefore, the optimization task could pay more attention to the feasible space with high quality. Moreover, some theoretical analyses are presented to prove the reliability of CM-MoSU. In experiments, several state-of-the-art sparse unmixing algorithms are compared. Both synthetic and real-world experiments demonstrate the effectiveness of CM-MoSU.
Zhenwei Shi 0001, Bin Pan, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2018 Attention-Based Convolutional Networks for Ship Detection in High-Resolution Remote Sensing Images
Wenyuan Li 0002, Zhenwei Shi 0001
PRCV (4)3
2018 Robust Sparse Unmixing for Hyperspectral Imagery
abstract
A linear sparse unmixing method based on spectral library has been widely used to tackle the hyperspectral unmixing problem, under the assumption that the spectrum of each pixel in the hyperspectral scene can be expressed as a linear combination of pure endmembers in the spectral library. However, because of the ion (atom) substitution in the geological process, there often exists spectral variability between the measured endmembers in the real environment and corresponding ones in the spectral library, which poses a significant challenge to linear sparse unmixing. Physically, the substitution leads to the variation of absorption peaks of endmembers, making the spectral variation of sparse property. To address the above problem, we introduce redundant spectrum to represent the spectral variation caused by ion (atom) substitution and develop a sparse redundant unmixing model by adding the redundant regularization into the classical sparse regression formulation. Based on the alternating direction method of multipliers, we develop a unified algorithm called sparse redundant unmixing to obtain the solution. Both simulation experiment and real data experiment demonstrate that the proposed method can effectively use the redundant spectrum to address the spectral variation problem caused by the ion (atom) substitution.
Dan Wang 0005, Zhenwei Shi 0001, Xinrui Cui
IEEE Trans. Geosci. Remote. Sens.2
2018 Random Access Memories: A New Paradigm for Target Detection in High Resolution Aerial Remote Sensing Images
abstract
We propose a new paradigm for target detection in high resolution aerial remote sensing images under small target priors. Previous remote sensing target detection methods frame the detection as learning of detection model + inference of class-label and bounding-box coordinates. Instead, we formulate it from a Bayesian view that at inference stage, the detection model is adaptively updated to maximize its posterior that is determined by both training and observation. We call this paradigm "random access memories (RAM)." In this paradigm, "Memories" can be interpreted as any model distribution learned from training data and "random access" means accessing memories and randomly adjusting the model at detection phase to obtain better adaptivity to any unseen distribution of test data. By leveraging some latest detection techniques e.g., deep Convolutional Neural Networks and multi-scale anchors, experimental results on a public remote sensing target detection data set show our method outperforms several other state of the art methods. We also introduce a new data set "LEarning, VIsion and Remote sensing laboratory (LEVIR)", which is one order of magnitude larger than other data sets of this field. LEVIR consists of a large set of Google Earth images, with over 22 k images and 10 k independently labeled targets. RAM gives noticeable upgrade of accuracy (an mean average precision improvement of 1% ~ 4%) of our baseline detectors with acceptable computational overhead.
Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Image Process.2
2017 Object Detection with Proposals in High-Resolution Optical Remote Sensing Images
Huoping Ding, Qinhan Luo, Zhengxia Zou, Cuicui Guo, Zhenwei Shi 0001
IDEAL5
2017 Super-Resolution for Remote Sensing Images via Local-Global Combined Network
abstract
Super-resolution is an image processing technology that recovers a high-resolution image from a single or sequential low-resolution images. Recently deep convolutional neural networks (CNNs) have made a huge breakthrough in many tasks including super-resolution. In this letter, we propose a new single-image super-resolution algorithm named local-global combined networks (LGCNet) for remote sensing images based on the deep CNNs. Our LGCNet is elaborately designed with its “multifork” structure to learn multilevel representations of remote sensing images including both local details and global environmental priors. Experimental results on a public remote sensing data set (UC Merced) demonstrate an overall improvement of both accuracy and visual performance over several state-of-the-art algorithms.
Sen Lei, Zhenwei Shi 0001, Zhengxia Zou
IEEE Geosci. Remote. Sens. Lett.2
2017 Fully Convolutional Network With Task Partitioning for Inshore Ship Detection in Optical Remote Sensing Images
abstract
Ship detection in optical remote sensing imagery has drawn much attention in recent years, especially with regards to the more challenging inshore ship detection. However, recent work on this subject relies heavily on hand-crafted features that require carefully tuned parameters and on complicated procedures. In this letter, we utilize a fully convolutional network (FCN) to tackle the problem of inshore ship detection and design a ship detection framework that possesses a more simplified procedure and a more robust performance. When tackling the ship detection problem with FCN, there are two major difficulties: 1) the long and thin shape of the ships and their arbitrary direction makes the objects extremely anisotropic and hard to be captured by network features and 2) ships can be closely docked side by side, which makes separating them difficult. Therefore, we implement a task partitioning model in the network, where layers at different depths are assigned different tasks. The deep layer in the network provides detection functionality and the shallow layer supplements with accurate localization. This approach mitigates the tradeoff of FCN between localization accuracy and feature representative ability, which is of importance in the detection of closely docked ships. The experiments demonstrate that this framework, with the advantages of FCN and the task partitioning model, provides robust and reliable inshore ship detection in complex contexts.
Haoning Lin, Zhenwei Shi 0001, Zhengxia Zou
IEEE Geosci. Remote. Sens. Lett.2
2017 A New Unsupervised Hyperspectral Band Selection Method Based on Multiobjective Optimization
abstract
Unsupervised band selection methods usually assume specific optimization objectives, which may include band or spatial relationship. However, since one objective could only represent parts of hyperspectral characteristics, it is difficult to determine which objective is the most appropriate. In this letter, we propose a new multiobjective optimization-based band selection method, which is able to simultaneously optimize several objectives. The hyperspectral band selection is transformed into a combinational optimization problem, where each band is represented by a binary code. More importantly, to overcome the problem of unique solution selection in traditional multiobjective methods, we develop a new incorporated rank-based solution set concentration approach in the process of Tchebycheff decomposition. The performance of our method is evaluated under the application of hyperspectral imagery classification. Three recently proposed band selection methods are compared.
Zhenwei Shi 0001, Bin Pan
IEEE Geosci. Remote. Sens. Lett.2
2017 Hierarchical Guidance Filtering-Based Ensemble Classification for Hyperspectral Images
abstract
Joint spectral and spatial information should be fully exploited in order to achieve accurate classification results for hyperspectral images. In this paper, we propose an ensemble framework, which combines spectral and spatial information in different scales. The motivation of the proposed method derives from the basic idea: by integrating many individual learners, ensemble learning can achieve better generalization ability than a single learner. In the proposed work, the individual learners are obtained by joint spectral-spatial features generated from different scales. Specially, we develop two techniques to construct the ensemble model, namely, hierarchical guidance filtering (HGF) and matrix of spectral angle distance (mSAD). HGF and mSAD are combined via a weighted ensemble strategy. HGF is a hierarchical edge-preserving filtering operation, which could produce diverse sample sets. Meanwhile, in each hierarchy, a different spatial contextual information is extracted. With the increase of hierarchy, the pixels spectra tend smooth, while the spatial features are enhanced. Based on the outputs of HGF, a series of classifiers can be obtained. Subsequently, we define a low-rank matrix, mSAD, to measure the diversity among training samples in each hierarchy. Finally, an ensemble strategy is proposed using the obtained individual classifiers and mSAD. We term the proposed method as HiFi-We. Experiments are conducted on two popular data sets, Indian Pines and Pavia University, as well as a challenging hyperspectral data set used in 2014 Data Fusion Contest (GRSS_DFC_2014). An effectiveness analysis about the ensemble strategy is also displayed.
Bin Pan, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2017 Can a Machine Generate Humanlike Language Descriptions for a Remote Sensing Image?
abstract
This paper investigates an intriguing question in the remote sensing field: “can a machine generate humanlike language descriptions for a remote sensing image?” The automatic description of a remote sensing image (namely, remote sensing image captioning) is an important but rarely studied task for artificial intelligence. It is more challenging as the description must not only capture the ground elements of different scales, but also express their attributes as well as how these elements interact with each other. Despite the difficulties, we have proposed a remote sensing image captioning framework by leveraging the techniques of the recent fast development of deep learning and fully convolutional networks. The experimental results on a set of high-resolution optical images including Google Earth images and GaoFen-2 satellite images demonstrate that the proposed method is able to generate robust and comprehensive sentence description with desirable speed performance.
Zhenwei Shi 0001, Zhengxia Zou
IEEE Trans. Geosci. Remote. Sens.1
2016 Collaborative sparse unmixing of hyperspectral data using L2, P norm
abstract
Sparse unmixing is a popular method in remotely sensed hyperspectral imagery interpretation. Recently, the collaborative sparse unmixing model has shown its advantage over the traditional single channel sparse unmixing method since it can utilize the subspace nature of the high-dimensional hyperspectral data to alleviate the difficulty caused by the usually high mutual coherence of the spectral library. However, the existing collaborative sparse unmixing model is constructed in the convex ℓ1norm framework while it is known that using the ℓp(02;p(02;p(0 <; p <; 1) norm collaborative sparse unmixing model.
Dan Wang 0005, Zhenwei Shi 0001, Wei Tang 0016
IGARSS2
2016 Hyperspectral Image Classification Based on Nonlinear Spectral-Spatial Network
abstract
Recently, for the task of hyperspectral image classification, deep-learning-based methods have revealed promising performance. However, the complex network structure and the time-consuming training process have restricted their applications. In this letter, we construct a much simpler network, i.e., the nonlinear spectral-spatial network (NSSNet), for hyperspectral image classification. NSSNet is developed from the basic structure of a principal component analysis network. Nonlinear information is included in NSSNet, to generate a more discriminative feature expression. Moreover, spectral and spatial features are combined to further improve the classification accuracy. Experimental results indicate that our method achieves better performance than state-of-the-art deep-learning-based methods.
Bin Pan, Zhenwei Shi 0001, Shaobiao Xie
IEEE Geosci. Remote. Sens. Lett.2
2016 No-Reference Assessment on Haze for Remote-Sensing Images
abstract
Assessment on haze can filter out images with dense haze to improve the reliability of remote-sensing image interpretation. In this letter, a novel no-reference haze assessment method based on haze distribution is proposed for remote-sensing images. First, range channel of an image is defined and the haze distribution map (HDM) is extracted from the hazy image. Then, the haze assessment metric HDM-based haze assessment (HDMHA) is designed according to the HDM. Finally, the degree of haze in remote-sensing images is predicted using the proposed metric. In order to objectively verify the effectiveness of the proposed metric HDMHA, a method of simulating hazy remote-sensing images based on the haze imaging model is proposed in this letter, and the simulated hazy images are greatly similar to real ones in vision. A series of experiments are done on both real images and simulated images, and the results show that the proposed metric achieves good consistency when compared with subjective experiments and outperforms typical blind image quality assessment methods.
Xiaoxi Pan, Fengying Xie, Zhiguo Jiang 0001, Zhenwei Shi 0001, Xiaoyan Luo
IEEE Geosci. Remote. Sens. Lett.4
2016 Hierarchical Suppression Method for Hyperspectral Target Detection
abstract
Target detection is an important application in the hyperspectral image processing field, and several detection algorithms have been proposed in the past decades. Some traditional detectors are built based on the statistical information of the target and background spectra, and their performances tend to be affected by the spectral quality. Some previous methods cope with this problem by refining the target spectra to make the detector robust. In this paper, instead of doing similar to this, we propose a new hierarchical method to suppress the backgrounds while preserving the target spectra, with the purpose of boosting the performance of traditional hyperspectral target detector. The proposed method consists of different layers of classical constrained energy minimization (CEM) detectors. In each layer of detection, the CEM's output of each spectrum is transformed by a nonlinear suppression function and then considered as a coefficient to impose on this spectrum for the next round of iteration. To our knowledge, such hierarchical structure is proposed for the first time. Theoretically, we prove the convergence of the proposed algorithm, and we also give a theoretical explanation on why we can obtain the gradually increasing detection performance through the hierarchical suppression process. Experimental results on two real hyperspectral images and one synthetic image suggest that our method significantly improves the performance of the original CEM detection algorithm and also outperforms other classical and recently proposed hyperspectral target detection algorithms.
Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2016 Ship Detection in Spaceborne Optical Image With SVD Networks
abstract
Automatic ship detection on spaceborne optical images is a challenging task, which has attracted wide attention due to its extensive potential applications in maritime security and traffic control. Although some optical image ship detection methods have been proposed in recent years, there are still three obstacles in this task: 1) the inference of clouds and strong waves; 2) difficulties in detecting both inshore and offshore ships; and 3) high computational expenses. In this paper, we propose a novel ship detection method called SVD Networks (SVDNet), which is fast, robust, and structurally compact. SVDNet is designed based on the recent popular convolutional neural networks and the singular value decompensation algorithm. It provides a simple but efficient way to adaptively learn features from remote sensing images. We evaluate our method on some spaceborne optical images of GaoFen-1 and Venezuelan Remote Sensing Satellites. The experimental results demonstrate that our method achieves high detection robustness and a desirable time performance in response to all of the above three problems.
Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2016 Real-Time Traffic Light Detection With Adaptive Background Suppression Filter
abstract
Traffic light detection plays an important role in intelligent transportation system, and many detection methods have been proposed in recent years. However, illumination variation effect is still of its major technical problem in real urban driving environments. In this paper, we propose a novel vision-based traffic light detection method for driving vehicles, which is fast and robust under different illumination conditions. The proposed method contains two stages: the candidate extraction stage and the recognition stage. On the candidate extraction stage, we propose an adaptive background suppression algorithm to highlight the traffic light candidate regions while suppressing the undesired backgrounds. On the recognition stage, each candidate region is verified and is further classified into different traffic light semantic classes. We evaluate our method on video sequences (more than 5000 frames and labels) captured from urban streets and suburb roads in varying illumination and compared with other vision-based traffic detection approaches. The experiment shows that the proposed method can achieve a desired detection result with high quality and robustness; simultaneously, the whole detection system can meet the real-time processing requirement of about 15 fps on video sequences.
Zhenwei Shi 0001, Zhengxia Zou, Changshui Zhang
IEEE Trans. Intell. Transp. Syst.1
2015 Quadratic Constrained Energy Minimization for hyperspectral target detection
abstract
In this paper, we propose a simple but effective algorithm, Quadratic Constrained Energy Minimization (QCEM) detector for hyperspectral image target detection. QCEM is a nonlinear version of classical Constrained Energy Minimization (CEM) detector, and it exploits the nonlinear characteristics of data by adding quadratic term on CEM model. Experimental results on one real hyperspectral images and one synthetic image suggest our method significantly improves the performance of the original CEM detection algorithm.
Zhengxia Zou, Zhenwei Shi 0001
IGARSS2
2015 A novel unsupervised approach to discovering regions of interest in traffic images
Zhenyu An, Zhenwei Shi 0001, Ying Wu 0001, Changshui Zhang
Pattern Recognit.2
2015 Sparse Unmixing of Hyperspectral Data Using Spectral A Priori Information
abstract
Given a spectral library, sparse unmixing aims at finding the optimal subset of endmembers from it to model each pixel in the hyperspectral scene. However, sparse unmixing still remains a challenging task due to the usually high mutual coherence of the spectral library. In this paper, we exploit the spectral a priori information in the hyperspectral image to alleviate this difficulty. It assumes that some materials in the spectral library are known to exist in the scene. Such information can be obtained via field investigation or hyperspectral data analysis. Then, we propose a novel model to incorporate the spectral a priori information into sparse unmixing. Based on the alternating direction method of multipliers, we present a new algorithm, which is termed sparse unmixing using spectral a priori information (SUnSPI), to solve the model. Experimental results on both synthetic and real data demonstrate that the spectral a priori information is beneficial to sparse unmixing and that SUnSPI can exploit this information effectively to improve the abundance estimation.
Wei Tang 0016, Zhenwei Shi 0001, Ying Wu 0001, Changshui Zhang
IEEE Trans. Geosci. Remote. Sens.2
2015 Robust Hyperspectral Image Target Detection Using an Inequality Constraint
abstract
In real hyperspectral images, there exist variations within spectra of materials. The inherent spectral variability is one of the major obstacles for the successful hyperspectral image target detection. Although several hyperspectral image target detection algorithms have been proposed, there are few algorithms considering the spectral variability. Under such circumstances, in this paper, we propose a hyperspectral image target detection algorithm that is robust to the target spectral variability. The proposed algorithm utilizes an inequality constraint to guarantee that the outputs of target spectra, which vary in a certain set, are larger than one, so that these target spectra could be detected. The proposed algorithm transforms the target detection to a convex optimization problem and uses a kind of interior point method named barrier method to solve the formulated optimization problem effectively. Two synthetic hyperspectral images and two real hyperspectral images are used to conduct experiments. The experimental results demonstrate the proposed algorithm is robust to the target spectral variability and performs better than other classical algorithms.
Zhenwei Shi 0001, Wei Tang 0016
IEEE Trans. Geosci. Remote. Sens.2
2014 Single Remote Sensing Image Dehazing
abstract
Remote sensing images are widely used in various fields. However, they usually suffer from the poor contrast caused by haze. In this letter, we propose a simple, but effective, way to eliminate the haze effect on remote sensing images. Our work is based on the dark channel prior and a common haze imaging model. In order to eliminate halo artifacts, we use a low-pass Gaussian filter to refine the coarse estimated atmospheric veil. We then redefine the transmission, with the aim of preventing the color distortion of the recovered images. The main advantage of the proposed algorithm is its fast speed, while it can also achieve good results. The experimental results demonstrate that our algorithm produces visually appealing dehazing images and retains the very fine details. Moreover, for images containing partly clear and partly hazy areas, our algorithm can also achieve good results.
Jiao Long, Zhenwei Shi 0001, Wei Tang 0016, Changshui Zhang
IEEE Geosci. Remote. Sens. Lett.2
2014 Subspace Matching Pursuit for Sparse Unmixing of Hyperspectral Data
abstract
Sparse unmixing assumes that each mixed pixel in the hyperspectral image can be expressed as a linear combination of only a few spectra (endmembers) in a spectral library, known a priori. It then aims at estimating the fractional abundances of these endmembers in the scene. Unfortunately, because of the usually high correlation of the spectral library, the sparse unmixing problem still remains a great challenge. Moreover, most related work focuses on the l1convex relaxation methods, and little attention has been paid to the use of simultaneous sparse representation via greedy algorithms (GAs) (SGA) for sparse unmixing. SGA has advantages such as that it can get an approximate solution for the l0problem directly without smoothing the penalty term in a low computational complexity as well as exploit the spatial information of the hyperspectral data. Thus, it is necessary to explore the potential of using such algorithms for sparse unmixing. Inspired by the existing SGA methods, this paper presents a novel GA termed subspace matching pursuit (SMP) for sparse unmixing of hyperspectral data. SMP makes use of the low-degree mixed pixels in the hyperspectral image to iteratively find a subspace to reconstruct the hyperspectral data. It is proved that, under certain conditions, SMP can recover the optimal endmembers from the spectral library. Moreover, SMP can serve as a dictionary pruning algorithm. Thus, it can boost other sparse unmixing algorithms, making them more accurate and time efficient. Experimental results on both synthetic and real data demonstrate the efficacy of the proposed algorithm.
Zhenwei Shi 0001, Wei Tang 0016, Zhana Duren, Zhiguo Jiang 0001
IEEE Trans. Geosci. Remote. Sens.1
2014 Ship Detection in High-Resolution Optical Imagery Based on Anomaly Detector and Local Shape Feature
abstract
Ship detection in high-resolution optical imagery is a challenging task due to the variable appearances of ships and background. This paper aims at further investigating this problem and presents an approach to detect ships in a “coarse-to-fine” manner. First, to increase the separability between ships and background, we concentrate on the pixels in the vicinities of ships. We rearrange the spatially adjacent pixels into a vector, transforming the panchromatic image into a “fake” hyperspectral form. Through this procedure, each produced vector is endowed with some contextual information, which amplifies the separability between ships and background. Afterward, for the “fake” hyperspectral image, a hyperspectral algorithm is applied to extract ship candidates preliminarily and quickly by regarding ships as anomalies. Finally, to validate real ships out of ship candidates, an extra feature is provided with histograms of oriented gradients (HOGs) to generate a hypothesis using AdaBoost algorithm. This extra feature focuses on the gray values rather than the gradients of an image and includes some information generated by very near but not closely adjacent pixels, which can reinforce HOG to some degree. Experimental results on real database indicate that the hyperspectral algorithm is robust, even for the ships with low contrast. In addition, in terms of the shape of ships, the extended HOG feature turns out to be better than HOG itself as well as some other features such as local binary pattern.
Zhenwei Shi 0001, Xinran Yu, Zhiguo Jiang 0001
IEEE Trans. Geosci. Remote. Sens.1
2014 Regularized Simultaneous Forward-Backward Greedy Algorithm for Sparse Unmixing of Hyperspectral Data
abstract
Sparse unmixing assumes that each observed signature of a hyperspectral image is a linear combination of only a few spectra (endmembers) in an available spectral library. It then estimates the fractional abundances of these endmembers in the scene. The sparse unmixing problem still remains a great difficulty due to the usually high correlation of the spectral library. Under such circumstances, this paper presents a novel algorithm termed as the regularized simultaneous forward-backward greedy algorithm (RSFoBa) for sparse unmixing of hyperspectral data. The RSFoBa has low computational complexity of getting an approximate solution for the l0problem directly and can exploit the joint sparsity among all the pixels in the hyperspectral data. In addition, the combination of the forward greedy step and the backward greedy step makes the RSFoBa more stable and less likely to be trapped into the local optimum than the conventional greedy algorithms. Furthermore, when updating the solution in each iteration, a regularizer that enforces the spatial-contextual coherence within the hyperspectral image is considered to make the algorithm more effective. We also show that the sublibrary obtained by the RSFoBa can serve as input for any other sparse unmixing algorithms to make them more accurate and time efficient. Experimental results on both synthetic and real data demonstrate the effectiveness of the proposed algorithm.
Wei Tang 0016, Zhenwei Shi 0001, Ying Wu 0001
IEEE Trans. Geosci. Remote. Sens.2
2013 A quasi-Newton-based spatial multiple materials detector for hyperspectral imagery
Zhenwei Shi 0001, Zhiguo Jiang 0001
Neural Comput. Appl.2
2013 Efficient sparse unmixing analysis for hyperspectral imagery based on random projection
Zhenwei Shi 0001, Xinya Zhai, Zhiguo Jiang 0001
Neural Comput. Appl.1
2013 Nonnegative matrix factorization-based hyperspectral and panchromatic image fusion
Zhou Zhang 0001, Zhenwei Shi 0001
Neural Comput. Appl.2
2012 Blind Separation of Superimposed Moving Images Using Image Statistics
abstract
We address the problem of blind separation of multiple source layers from their linear mixtures with unknown mixing coefficients and unknown layer motions. Such mixtures can occur when one takes photos through a transparent medium, like a window glass, and the camera or the medium moves between snapshots. To understand how to achieve correct separation, we study the statistics of natural images in the Labelme data set. We not only confirm the well-known sparsity of image gradients, but also discover new joint behavior patterns of image gradients. Based on these statistical properties, we develop a sparse blind separation algorithm to estimate both layer motions and linear mixing coefficients and then recover all layers. This method can handle general parameterized motions, including translations, scalings, rotations, and other transformations. In addition, the number of layers is automatically identified, and all layers can be recovered, even in the underdetermined case where mixtures are fewer than layers. The effectiveness of this technology is shown in experiments on both simulated and real superimposed images.
Kun Gai, Zhenwei Shi 0001, Changshui Zhang
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 A Hierarchical Connection Graph Algorithm for Gable-Roof Detection in Aerial Image
abstract
In this letter, we present a hierarchical connection graph (HCG) algorithm based on a self-avoiding polygon (SAP) model for detecting and extracting gable roofs from aerial imagery. The SAP model is a deformable shape model that is capable of representing gable roofs of various shapes and appearances. The model is composed of a sequence of roof-corner templates that are connected into a SAP, which serves as a flexible shape prior. An energy function that combines features from three channels (corner, boundary, and interior area) is defined over the sequence to quantify the variability in appearances of gable roofs. To infer the most probable state of the corner sequence for an input image, we use an efficient algorithm-called HCG algorithm. The algorithm converts the solution space of a SAP model into a directed graph (which we call “HCG”) and searches for the best path using dynamic programming (DP). It is efficient for two reasons: 1) By constructing an HCG, the algorithm can quickly prune out a large amount of invalid solutions using only geometric constraints, which are inexpensive to compute, and 2) by employing DP, the algorithm decomposes the searching problem into smaller overlapping subproblems and reuses energy scores, which are expensive to compute. Experimental results on a set of challenging gable roofs show that our algorithm has good performance and is computationally effective.
Qiongchen Wang, Zhiguo Jiang 0001, Junli Yang, Danpei Zhao, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.5
2011 Blind Source Separation Using Quadratic form Innovation
Zhenwei Shi 0001, Hongjuan Zhang, Xueyan Tan, Zhiguo Jiang 0001
Neural Process. Lett.1
2010 Interactive localized content based image retrieval with multiple-instance active learning
Dan Zhang 0007, Fei Wang 0001, Zhenwei Shi 0001, Changshui Zhang
Pattern Recognit.3
2009 Gable Roof Description by Self-Avoiding Polygon
Qiongchen Wang, Zhiguo Jiang 0001, Junli Yang, Danpei Zhao, Zhenwei Shi 0001
ACCV (3)5
2009 Blind separation of superimposed images with unknown motions
abstract
We consider the blind separation of source layers from superimposed mixtures thereof, involving unknown motions and unknown mixing coefficients of layers in each mixture. Previous blind separation approaches for such problems assume motions to be uniform translations, and hence are limited for real world applications. In this paper, we develop a sparse blind separation algorithm to estimate both parameterized motions and mixing coefficients. Then, a novel reconstruction approach is presented to recover all layers, by utilizing not only the mixing model but also the statistical properties of natural images. The whole method can handle more general motions than translations, including scalings, rotations and other transformations. In addition, the number of layers is automatically identified, and all layers can be recovered even in the under-determined case where mixtures are fewer than layers. The effectiveness of this technology is shown in the experiments on two simulated mixtures of four layers, real photos containing transparency and reflections, and real crossfade images from videos.
Kun Gai, Zhenwei Shi 0001, Changshui Zhang
CVPR2
2009 Fast nonlinear autocorrelation algorithm for source separation
Zhenwei Shi 0001, Changshui Zhang
Pattern Recognit.1
2008 Blindly separating mixtures of multiple layers with spatial shifts
abstract
We address the problem of blindly separating mixtures of multiple layer images with unknown spatial shifts and mixing coefficients. Our proposed method can handle the over-determined, determined and under-determined cases where mixtures are more than, as many as and fewer than layers, respectively. The method is fast in over-determined and determined cases, with the same complexity as the fast Fourier transform (FFT), and can separate more layers from fewer mixtures in the under-determined case. It consists of two main steps. First, a novel sparse blind separation algorithm is applied, to estimate the spatial shifts, the mixing coefficients and the edge image of each layer. Second, all layers are reconstructed, by large scale linear programming in the under-determined case, or by least-squares solutions in other cases. The effectiveness of this technology is shown in the experiments on two simulated mixtures of four layers with spatial shifts, real mixture photos containing transparency and reflections, and real mixture images in a dissolve from a video.
Kun Gai, Zhenwei Shi 0001, Changshui Zhang
CVPR2
2008 Localized content based image retrieval by multiple instance active learning
abstract
In this paper, we propose two general multiple instance active learning (MIAL) algorithms, multiple-instance active learning with a simple margin strategy (S-MIAL) and multiple- instance active learning with fisher information (F-MIAL), and apply them to the relevance feedback in localized content based image retrieval (LCBIR). S-MIAL considers the most ambiguous picture as the most valuable one, while F-MIAL can utilize the fisher information and analyze the value of the unlabeled pictures by assigning different labels to them. We show that F-MIAL can be integrated more naturally into the multiple instance learning scenario. In experiments, we will show their superior performances on some real-world image datasets.
Dan Zhang 0007, Fei Wang 0001, Zhenwei Shi 0001, Changshui Zhang
ICIP3
2008 MACBSE: Extracting signals with linear autocorrelations
Zhenwei Shi 0001, Dan Zhang 0007, Changshui Zhang
Neurocomputing1
2008 Unsupervised Single-Channel Music Source Separation by Average Harmonic Structure Modeling
abstract
Source separation of musical signals is an appealing but difficult problem, especially in the single-channel case. In this paper, an unsupervised single-channel music source separation algorithm based on average harmonic structure modeling is proposed. Under the assumption of playing in narrow pitch ranges, different harmonic instrumental sources in a piece of music often have different but stable harmonic structures; thus, sources can be characterized uniquely by harmonic structure models. Given the number of instrumental sources, the proposed algorithm learns these models directly from the mixed signal by clustering the harmonic structures extracted from different frames. The corresponding sources are then extracted from the mixed signal using the models. Experiments on several mixed signals, including synthesized instrumental sources, real instrumental sources, and singing voices, show that this algorithm outperforms the general nonnegative matrix factorization (NMF)-based source separation algorithm, and yields good subjective listening quality. As a side effect, this algorithm estimates the pitches of the harmonic instrumental sources. The number of concurrent sounds in each frame is also computed, which is a difficult task for general multipitch estimation (MPE) algorithms.
Zhiyao Duan, Yungang Zhang, Changshui Zhang, Zhenwei Shi 0001
IEEE Trans. Speech Audio Process.4
2007 Localized Content-Based Image Retrieval Using Semi-Supervised Multiple Instance Learning
Dan Zhang 0007, Zhenwei Shi 0001, Yangqiu Song, Changshui Zhang
ACCV (1)2
2007 Multi-Pitch Estimation Based on Partial Event and Support Transfer
abstract
This paper proposes a method for the multi-pitch estimation of polyphonic music signals. Instead of on the frame level, the estimation is based on the partial event, which is defined like the note event in MIDI. All partial events in a piece of music are extracted dynamically in the process of the frame by frame short time Fourier transform (STFT). For each event, net support degree received from other events is calculated and the events with the highest support degrees are selected to be the fundamental frequency (FO) events. From another point of view, the support is transferred from higher frequency partial events to lower ones and finally concentrated on the FO events. This method can estimate the number of concurrent sounds, the onset and offset times of the notes. Experiments on both randomly mixed chord signals and synthesized ensemble music signals in "wav" format are conducted and the results are promising.
Zhiyao Duan, Dan Zhang 0007, Changshui Zhang, Zhenwei Shi 0001
ICME4
2007 A Fixed-Point Algorithm for Blind Separation of Temporally Correlated Sources
abstract
In this paper we develop a new method for blind separation of temporally correlated sources, possibly dependent signals from linear mixtures of them. The proposed algorithm is based on the mutual independency of the innovations of source signals instead of original signals. This algorithm takes into account both the temporal structure and the high-order statistics of source signals and in contrast to the most known blind separation algorithms only exploiting the second order statistics or the non-Gaussianity. In this framework, a fixed-point algorithm is introduced. The fixed-point algorithm is computationally very simple, converge fast, and does not need choose any learning step sizes. Extensive computer simulations with speech signals and images confirm the validity and high performance of the proposed algorithm.
Zhenwei Shi 0001, Dan Zhang 0007, Changshui Zhang
ICME1
2007 Semi-blind source extraction for fetal electrocardiogram extraction by combining non-Gaussianity and time-correlation
Zhenwei Shi 0001, Changshui Zhang
Neurocomputing1
2007 Nonlinear innovation to blind source separation
Zhenwei Shi 0001, Changshui Zhang
Neurocomputing1
2007 Blind Source Extraction Using Generalized Autocorrelations
abstract
This letter addresses blind (semiblind) source extraction (BSE) problem when a desired source signal has temporal structures, such as linear or nonlinear autocorrelations. Using the temporal characteristics of sources, we develop objective functions based on the generalized autocorrelations of primary sources. Maximizing the objective functions, we propose simple fixed-point source extraction algorithms. We give the stability analysis and prove convergence properties of the algorithms as the generalized autocorrelation function is linear or nonlinear. Especially, as the generalized autocorrelation function is linear, the algorithm has interesting character of "one-iteration" convergence under some conditions. Computer simulations and real-data application experiments show that the algorithms are appealing BSE methods for temporal signals of interest by capturing the linear or nonlinear autocorrelations of the desired sources.
Zhenwei Shi 0001, Changshui Zhang
IEEE Trans. Neural Networks1
2006 Gaussian moments for noisy complexity pursuit
Zhenwei Shi 0001, Changshui Zhang
Neurocomputing1