EDBT 2026 Demo / reviewers in the wild / expert
Qian Shi 0001
dblp:45/2408-1
· DBLP profile ↗
48ranked-venue papers
10as first author
34since 2021 · last 2026
0000-0002-1276-0352ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 42 · 8 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Fine-Grained Oriented Ship Detection for Remote Sensing Imagery via Controllable Generative PretrainingabstractFine-grained ship recognition in remote sensing imagery is essential for maritime applications. However, its development is hindered by two challenges: 1) the limited granularity of existing ship detection datasets, and 2) the disturbance of complex maritime conditions as well as the arbitrary ship orientations and distributions. To address the first issue, we annotated a large-scale fine-grained ship instance detection dataset (LAFI), comprising 48,717 ship instances worldwide with 49 categories. To tackle the challenges of marine disturbance and diverse ship status, we proposed a controllable generative knowledge-driven ship detection framework (COSD). It employs a controllable diffusion model guided by ship-marine textual prompt to generate millions of synthetic images that not only preserve ship structures but also cover diverse sea and weather conditions for robust pretraining. The pretraining stage then utilizes masked reconstruction to learn component-level cues under occlusion, clutter, fog, and illumination changes. Furthermore, a heterogeneous feature alignment decoder is designed to align multi-modal metrics of orientation and distribution features in the latent space, allowing for accurate representation of diverse ship status. Extensive experiments on two benchmark datasets showed that our method respectively increased 0.011 and 0.030 mean average precision (mAP@50) over SOTA methods, particularly in scenarios involving small, densely packed and arbitrary oriented ships. Da He, Xikun Hu, Ping Zhong 0001, Qian Shi 0001, Xiaoping Liu 0001, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 8 |
| 2025 | BTCDNet: Bayesian Tile Attention Network for Hyperspectral Image Change DetectionabstractHyperspectral images (HSI) provide detailed spectral information, which are effective for change detection (CD). Prior knowledge has been proven to improve the robustness of models in HSI processing. However, current CD methods do not fully utilize prior knowledge and research on hyperspectral mangroves CD is limited. In this letter, we propose a general hyperspectral CD model with Bayesian prior guided module (BPGM) and tile attention block (TAB) called BTCDNet. BPGM leverages prior information to steer the model training process under limited labeled samples condition, while TAB can reduce complexity and improve performance by tile attention. Moreover, a novel and restricted hyperspectral CD dataset Shenzhen has been annotated for hyperspectral mangroves CD reference. Experiments demonstrate that our proposal achieves state-of-the-art (SOTA) performances on this dataset and two other public benchmark datasets. Our code and datasets are available at https://github.com/JeasunLok/BTCDNet. Junshen Luo, Jiahe Li 0017, Xinlin Chu, Sai Yang, Lingjun Tao, Qian Shi 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2025 | PSODNet: Pretrained Scene-Aware Object Detection for Optical Remote Sensing ImageryabstractWith the widespread application of remote sensing images in military and civilian fields, remote sensing object detection (RSOD) has become an important research direction. However, limited generalization and complex background interference have long been persistent challenges that hinder the development of RSOD. To address these issues, we propose a Pre-trained Scene-aware Object Detection Network (PSODNet). Firstly, we design an Enhanced Object Network (EON), which leverages a multi-head pretraining strategy to jointly train data from diverse sources, thereby expanding the scale of dataset and improve the generalization ability. Secondly, we introduce scenario-object relationship module to learn a multi-scale relationship map between objects and scenes, which is used to constrain the solution space of object detection, thereby enhancing performance in complex scenarios. Lastly, by using Label Smoothing Loss, PSODNet leverages mutual information of label to prevent extreme distributions of classification probabilities and reduce the risk of overfitting. In the experiment part, PSODNet was pretrained on multiple datasets and then fine-tuned on three datasets for validation. Results on three public datasets demonstrate that PSODNet outperforms existing models in detection performance by up to 2.4%, achieving a maximum mAP of up to 95%. Through visual interpretation of the relationship map, we found that PSODNet is able to bridge the semantic relevance between objects and scenes, demonstrating its potential in object detection in complex scenarios. Code is available at: https://github.com/creature-compound/PSODNet. Da He, Qian Shi 0001, Xiaoping Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | LGTC: Local-Global Tri-Consistency Network for Semi-Supervised Change Detection of Remote Sensing ImagesabstractRecently, semi-supervised change detection (SSCD) has attracted considerable attention due to its remarkable capability to enhance model performance with limited annotated data. However, most existing SSCD methods rely primarily on pixel-level consistency learning to leverage unlabeled data. Although this approach can provide basic prediction consistency constraints, it exhibits notable limitations in capturing spatial continuity and global semantic coherence in structural change regions, and is easily affected by noise. To address this issue, we propose a novel SSCD framework, termed the local-global tri-consistency network (LGTC). LGTC incorporates a tri-consistency learning strategy at the pixel, region, and image levels, which provides complementary supervisory signals from fine-grained to global semantic scales, significantly enhancing representation ability and robustness on unlabeled data. Furthermore, a local-global interaction block (LGIBlock) is introduced to integrate local feature details with global context. Within it, the GFTBlock employs frequency domain attention to achieve efficient modeling of global information while effectively suppressing background noise, further enhancing the recognition performance of the model. Experimental results demonstrate that LGTC achieves state-of-the-art performance, attaining F1 scores of 87.12%, 88.65% and 83.25% on the LEVIR-CD, WHU-CD and CDD-CD dataset with 5% labeled data, respectively, fully verifying the effectiveness of the proposed method. Our source codes are available at https://github.com/sherryxu21/LGTC. Rui Xu 0031, Fulin Luo, Chuan Fu, Tan Guo, Qian Shi 0001, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | SCMVC: Semantic Constraint-Based Spatial-Spectral Multiview Clustering for Hyperspectral ImagesabstractCross-view consensus representation plays a crucial role in hyperspectral image (HSI) clustering. Recently, multi-view contrastive cluster (MVCC) methods have leveraged contrastive loss to extract contextual consensus representations. However, these methods suffer from a critical limitation: MVCC frameworks often regard similar heterogeneous views as positive sample pairs while treating dissimilar homogeneous views as negative sample pairs. This misalignment leads to intra-class inconsistency and inter-class confusion. To address this problem, we propose a novel multi-view clustering method, termed Semantic Constraint-based Spatial-Spectral Multi-view Clustering (SCMVC). First, spatial views are designed to capture diverse features for contrastive clustering. Meanwhile, globally relevant information from the spectral view is extracted using a Transformer, which serves to enhance the representation of similar samples in the spatial multi-view. Then, SCMVC employs a semantic constraint-based joint loss function, comprising a semantic contrast loss and a semantic similarity consistency loss. The semantic contrast loss captures high-level, domain-invariant features from hyperspectral images, while the semantic similarity consistency loss enforces stricter constraints on the similarity of semantically related samples in feature space. Finally, SCMVC utilizes anchor points to guide similarity clustering, reducing randomness by predefining these points. This approach captures the directional characteristics of data, leading to a more stable clustering process. Abundant experiment studies on numerous benchmarks verify the superiority of SCMVC in comparison to some state-of-the-art clustering methods. The codes are available at SCMVC. Fulin Luo, Yi Liu 0038, Tan Guo, Chuan Fu, Qian Shi 0001, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Intercomparison of Ku- and C-Band Backscatter Feature Parameters for Arctic Sea Ice Using Spaceborne FengYun-3E WindRAD ScatterometerabstractThis study exploits the unique capabilities of the FY-3E WindRAD scatterometer, the first spaceborne dual-frequency (Ku- and C-band) and dual-polarization (hhandvv) rotating fan-beam scanning measurements, to investigate the backscatter characteristics of open water (OW), first-year ice (FYI), and multi-year ice (MYI) under different seasonal, wavelength, and polarization conditions throughout 2022 in the Arctic. Four types of feature parameters were defined for systematic analysis based on WindRAD swath data. It is concluded that the mean backscatter coefficient σp,λand the wavelength gradient ratioGRpare key indicators for distinguishing between FYI and MYI, with the Ku-band exhibiting superior performance outside the melt season due to enhanced volume scattering from desalinated ice and bubble structures. During melting, however, both ice types become indistinguishable as meltwater increases dielectric loss and reduces penetration depth. Furthermore, the standard deviation of the backscatter coefficient Δσp,λand the polarization ratio γλprove highly effective in separating sea ice from OW with the C-band showing particular advantage owing to a wider incidence angle range and stronger angular sensitivity of Bragg scattering over water. The γλapproaches 1 for both FYI and MYI due to depolarizing rough surfaces, whereas OW exhibits lower values dominated by Bragg scattering. This study provides a systematic observational basis for exploring the benefits of dual-frequency joint detection in enhancing sea ice monitoring capabilities, providing vital support for the development and refinement of algorithms for FY-3E WindRAD operational sea ice products. Xiaochun Zhai, Shengrong Tian, Jian Shang, Guangzhen Cao, Minghu Ding, Xiao Cheng 0001, Lei Zheng 0016, Qian Shi 0001, Yufang Ye, Zhaojun Zheng, Yixuan Shou, Na Xu 0001, Xiuqing Hu, Lin Chen 0017 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | An Improved TES Method to Retrieve Urban Surface TemperatureabstractThe urban complex material and geometry characteristics result in a 3D thermal heterogeneity and that limits the urban surface temperature (UST) retrieval. In this study, we improved the temperature and emissivity separation (TES) algorithm by incorporating thermal heterogeneity within mixed pixel (MP). The improvement was based on the Discrete Anisotropic Radiative Transfer (DART) model and applied to retrieve land surface temperature (LST) from SDGSAT-1. The TES-MP algorithm was validated with ECOSTRESS and the data simulated by DART model, and the results show that it can reach good accuracy under complex urban conditions. Based on the simulated scenes from the Sheung Wan building in Hong Kong, the RMSE of TES-MP algorithm is 0.85K under thermal homogeneous conditions and 1.13K under thermal heterogeneous conditions. Additionally, new high-reflectivity construction materials are common in urban areas, i.e. metal materials. It shows that the relationship between MMD and minimums emissivity(εmin) is not applicable to these materials. Thus, the impacts of such materials on the UST retrieval were evaluated. The results show that the higher the reflectivity and the fractional abundance of such materials, the larger the LST underestimation. Under nadir observation conditions, the proportion of high-reflectivity walls does not cause significant LST retrieval errors. The geometry and adjacency effects on retrieved LST were evaluated, and results show that the TES-MP algorithm has some resistance to geometry and adjacency effects, thereby reducing errors in LST retrieval. This study provides a new view on retrieving LST of urban MPs, and also suggests that three or more bands should be considered when setting up thermal infrared sensors. Lili Zhu, Jinxin Yang, Xiaoying OuYang, Qian Shi 0001, Yong Xu 0002, Man Sing Wong, Massimo Menenti |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Explicable Sub-Pixel Mapping Based on Nested Self-Attention Network with Spatial Correlation Learnable MechanismabstractConvolutional-based sub-pixel mapping (SPM) is a cutting-edge approach to solve the mixed pixel problem in land cover mapping. However, the feature modeling process of convolutional neural networks (CNN) lacks interpretability, and thus, it is difficult to learn spatial correlation that is helpful for sub-pixel location reasoning, and usually fail in reconstruction of fragmented land parcels. Therefore, this study constructs a SPM network that combines data-driven and model-driven approaches (DEMON). It uses nested self-attention mechanisms to simulate spatial dependency modeling, which learns spatial correlations between sub-pixels and pixels of different land covers, the learned spatial correlations can then be used to explicitly infer the spatial positions to achieve interpretability. Three public datasets are used for validation, and we found that our proposed method outperforms SOTA methods by circa 6%, and the visualized spatial correlations verifies the interpretability of the model. Da He, Qian Shi 0001, Jingqian Xue, XueXiaoping Liu |
IGARSS | 2 |
| 2024 | A Novel Multiplatform Spatiotempoal Data Fusion Approach for Remote Sensing Imagery Based on Parameter SelectionabstractSpatiotemporal fusion is an important means to reconstruct the medium spatial resolution remote sensing image series. Presently, many spatiotemporal fusion approaches have been developed and adopted in research on agriculture, ecology, environment, and so on. Although these approaches have achieved remarkable performance in experiments and applications, most of them are designed to fuse all involved bands using the same model with the same parameters, which ignores the band difference. The ignorance may limit the fusion quality for some bands. To address this problem, we propose a novel spatiotemporal data fusion approach based on parameter selection (PSDFA) in this article. The core idea of the newly proposed PSDFA is producing the synthetic image pairs using available data via three means first and then selecting a similar image pair for each band to provide the parameters that are needed for their fusion. The PSDFA can not only be applied in local computers, and its simplified version can also be implemented in Google Earth Engine (GEE), which is a powerful and widely used cloud platform for remote sensing data computing. To test the PSDFA, we conduct two experiments, one in local computers and another in GEE. In local computers, the PSDFA is compared with five state-of-the-art fusion methods on two public Landsat–Moderate Resolution Imaging Spectroradiometer (MODIS) datasets. In GEE, it is used to produce the monthly 30-m image series in two study sites in the USA and compared with another GEE-based fusion approach. The experimental results demonstrate the outstanding performance of the proposed PSDFA in both local computers and GEE. Yunfei Li 0006, Liangli Meng, Zhenjie Liu, Qian Shi 0001, Jun Li 0009 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | A Memory-Guided Network and a Novel Dataset for Cropland Semantic Change DetectionabstractThe frequent occurrence of nonagriculturalization events has posed significant challenges to global food security and sustainable development. Despite the emerging deep learning (DL) algorithms demonstrating effective capability in capturing changes from remote sensing imagery, they have not consistently maintained favorable performance in cropland semantic change detection (CropSCD) tasks. The primary challenge lies in the natural contradiction between the diverse change classes and the sparse availability of change samples. Furthermore, the scarcity of CropSCD datasets also restricts the capabilities of data-driven models. Therefore, in order to encode diverse semantics from a small amount of change pixels, a memory-guided network (MeGNet) dedicated to CropSCD tasks is proposed. In particular, a class-aware memory module is introduced in the MeGNet to preserve change semantics, which can guide the model to distinguish different change classes. Moreover, a high-resolution CropSCD dataset is also constructed to alleviate the issue of insufficient dataset. The CropSCD dataset comprises 4141 pairs of images, each with a size of$512\times 512$, and is annotated with corresponding labels for eight cropland change classes. Comparative experiments have substantiated the superiority of MeGNet over current state-of-the-art (SOTA) methods, with the highest mean-F1 and mean intersection over union (mIoU) of 71.44% and 58.01% on the high-resolution semantic change detection (HRSCD) dataset, and those of 55.42% and 42.14% on the CropSCD dataset. These results have validated the feasibility and potential of the proposed MeGNet and CropSCD dataset in CropSCD tasks. Mengxi Liu 0001, Simin Lin, Yutong Zhong, Qian Shi 0001, Jiaqi Li 0011 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | An Object Fine-Grained Change Detection Method Based on Frequency Decoupling Interaction for High-Resolution Remote Sensing ImagesabstractChange detection is a prominent research direction in the field of remote sensing image processing. However, most current change detection methods focus solely on detecting changes without being able to differentiate the types of changes, such as “appear” or “disappear” of objects. Accurate detection of change types is of great significance in guiding decision-making processes. To address this issue, this article introduces the object fine-grained change detection (OFCD) task and proposes a method based on frequency decoupling interaction (FDINet). Specifically, in order to enhance the model’s ability to detect change types and improve its robustness to temporal information, a temporal exchange framework is designed. Additionally, to better capture spatial–temporal correlation in bi-temporal features, a wavelet interaction module (WIM) is proposed. This module utilizes wavelet transform for frequency decoupling, separating features into different components based on their frequency magnitudes. Then the module applies different interaction methods according to the characteristics of these frequency components. Finally, to aggregate complementary information from different-scale feature maps and enhance the representational capabilities of the extracted features, a feature aggregation and upsampling module (FAUM) is adopted. A series of experiments show the superiority of FDINet over most state-of-the-art methods, achieving good results on three different datasets. Yingjie Tang, Shou Feng, Chunhui Zhao 0003, Yuanze Fan, Qian Shi 0001, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Incorporating Reliability in Graph Information Propagation by Fluid Dynamics Diffusion: A case of Multimodal Semisupervised Deep LearningabstractClassic graph neural networks show some limitations in information extraction performance when applied to multimodal datasets. This is primarily due to such datasets having high volume, variety, and variability. In this paper, we propose structuring graph neural networks on a new graph representation based on fluid dynamics diffusion that allows us to incorporate the reliability of the features used to characterise each sample within the graph structure itself. This approach aims to address some of the major limitations of the classic graph-based learning structures, so to improve accuracy and robustness of the estimates. We show how this approach can help to strongly improve the quality of the analysis of classic graph neural networks. Experimental results are reported to support this point. Andrea Marinoni, Marine Mercier, Qian Shi 0001, Sivasakthy Selvakumaran, Mark A. Girolami |
ICASSP | 3 |
| 2023 | Using High-Resolution Nighttime Remote Sensing Data to Identify Light Sources in Hong KongabstractAlthough artificial light at night (ALAN) is essential for nighttime activities, any unregulated and abusive usage can lead to severe degradation in quality of life. In this study, we used high-resolution nighttime remote sensing data of 1-meter spatial resolution to investigate light sources in an urban area of Hong Kong with one million residents. We classified ALAN sources into three categories based on their origins: Building, Park, and Street. We found that 42% of light was from Building, a large fraction of which was unnecessary decorative lighting such as signboards. Lighting from Street accounted for 41%, whereas an unexpectedly high proportion was associated with Park (17%), with sport facilities-related lighting being the dominant contributor. We also detailed one case study which shows the disruptive effects of unregulated usage of LED signboards to the neighboring residential apartments. Shengjie Liu 0001, Chu Wing So, Hung Chak Ho, Qian Shi 0001, Chun Shing Jason Pun |
IGARSS | 4 |
| 2022 | A Deep Learning Method for Fined-Grained Urban Green Space MappingabstractIn view of the challenges on urban green space (UGS) mapping from high-resolution images (HRIs), including insufficient dataset as well as the intra-class difference and inter-class similarity of UGS in HRIs, we propose a novel network for UGS extraction (UGSNet) and collect an large urban green space dataset (UGSet) with 4,454 samples of size $512\times 512$ in this paper. The UGSNet integrates the attention mechanism to improve the discrimination of UGS, and employs a point head with point rending strategy for precise edge recovery. Comparison experiments with the state-of-the-art (SOTA) semantic segmentation models show that the UGSNet can achieve the highest F1 of 77.30% on UGSet. Mengxi Liu 0001, Zeteng Li, Qian Shi 0001 |
IGARSS | 4 |
| 2022 | Estimating PM2.5 and PM10 on Zhuhai-1 Hyperspectral ImageryabstractParticulate matter (PM), such as PM2.5 and PM10, was the major pollutant in a severe air pollution episode in 2013 eastern China. Limited by the coverage of stations, fine-scale monitoring at every corner in the city is difficult, if not impossible. Hyperspectral imagery can capture the ground and air information, from which we can estimate the concentrations of PM. In this study, we develop a multitask learning method to estimate the concentrations of PM based on the 10-m hyperspectral data from the newly-launched Zhuhai-1 satellites. We first convert the raw radiance to top-of-atmosphere (TOA) reflectance using the 1985 Wehrli solar irradiance spectrum. Then, we train a multitask network to simultaneously estimate PM2.5 and PM10 concentrations based on the TOA hyperspectral data. Results show that our method leads to estimations of an R-squared of 0.77 for PM2.5 and an R-squared of 0.42 for PM10. Shengjie Liu 0001, Qian Shi 0001 |
IGARSS | 2 |
| 2022 | EC-SAGINs: Edge-Computing-Enhanced Space-Air-Ground-Integrated Networks for Internet of VehiclesabstractEdge-computing-enhanced Internet of Vehicles (EC-IoV) enables ubiquitous data processing and content sharing among vehicles and terrestrial edge computing (TEC) infrastructures (e.g., 5G base stations and roadside units) with little or no human intervention, and plays a key role in the intelligent transportation systems. However, EC-IoV is heavily dependent on the connections and interactions between vehicles and TEC infrastructures, thus will break down in some remote areas where TEC infrastructures are unavailable (e.g., desert, isolated islands, and disaster-stricken areas). Driven by the ubiquitous connections and global-area coverage, space–air–ground-integrated networks (SAGINs) efficiently support seamless coverage and efficient resource management, and represent the next frontier for edge computing. In light of this, we first review the state-of-the-art edge computing research for SAGINs in this article. After discussing several existing orbital and aerial edge computing architectures, we propose a framework of edge computing-enabled SAGINs to support various Internet of Vehicles (EC-IoV) services for the vehicles in remote areas. The main objective of the framework is to minimize the task completion time and satellite resource usage. To this end, a preclassification scheme is presented to reduce the size of action space, and a deep imitation learning-driven offloading and caching algorithm is proposed to achieve real-time decision making. The simulation results show the effectiveness of our proposed scheme. Finally, we also discuss some technology challenges and future directions. Shuai Yu 0001, Xiaowen Gong, Qian Shi 0001, Xiaofei Wang 0001, Xu Chen 0004 |
IEEE Internet Things J. | 3 |
| 2022 | Spectral-Spatial Fusion Sub-Pixel Mapping Based on Deep Neural NetworkabstractSub-pixel mapping (SPM) has been widely adopted to alleviate the mixed pixel problem in hyperspectral image, as an extension of spectral unmixing (SU), providing a way to observe the spatial location of the endmember within mixed pixel. However, most of the SPM methods are unmixing-then-mapping (UTM), i.e., SPM process relies on the abundance images generated from SU, in which process uncertainty inherently exists and would be propagated to SPM. Furthermore, the prior knowledge toward the sub-pixel scale distribution is mainly model-driven/handcrafted, which has limitation for geographical-realistic distribution representation. In this letter, we proposed spectral–spatial fusion SPM based on deep neural network (SSNET), to realize the integrative modeling of SU and SPM problem in a unified network fashion to avoid uncertainty accumulation in UTM process, and it can simultaneously generate SU result and SPM result. Besides, SSNET provides a supervised manner to learn prior knowledge with external exemplar pairs of low- and high-resolution images for a geographical-realistic distribution representation. The experiment with two hyperspectral images validated the superiority of the proposed SSNET. Da He, Qian Shi 0001, Xiaoping Liu 0001, Yanfei Zhong, Xiaoding Liu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | PA-Former: Learning Prior-Aware Transformer for Remote Sensing Building Change DetectionabstractBuilding change detection (BCD) is significant for urban planning and environmental protection. In view of the inter-class similarity and intra-class difference of building changes in complex built-up area, specialized solutions have been introduced in BCD. Mainstream methods include extracting building prior information in advance and enhancing long-range context information. These methods often require additional processing, and ignore the construction of cross-temporal context information, resulting in deficiencies on CD performance and efficiency. Therefore, an end-to-end PA-Former for BCD is proposed in this letter, which combines prior extraction and contextual fusion together by learning prior-aware Transformer. Specifically, the PA-Former adopts a prior-feature extractor to capture prior and deep features from the bi-temporal images, in which a prior interpreter is integrated to obtain priori structural information of buildings. Besides, a prior-aware Transformer module (PATM) is designed to obtain contextual tokens with spatiotemporal information from the prior features, and integrate into the deep features. Extensive experiments with state-of-the-art methods are conducted for comparison. Particularly, the PA-Former surpasses the baselines with an F1 of 88.79% on BCDD dataset and that of 85.32% on Google dataset. Mengxi Liu 0001, Qian Shi 0001, Zhuoqun Chai |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | An Adversarial Domain Adaptation Framework With KL-Constraint for Remote Sensing Land Cover ClassificationabstractLand cover classification plays a crucial role in land resource monitoring and planning. Recently, deep learning-based methods are becoming the dominating method for precise land cover mapping. However, the large-scale application of them is deeply hindered by the domain shift between different images, which is easily caused by illumination, climate, regional divergence, and so on. With the aim to cope with the problem of domain shift, many domain adaptation (DA) methods have been provided and great achievements have been made, especially the newborn adversarial DA, which usually contains a generator and a discriminator. Among these methods, the pixel-level methods are of high memory consumption, whereas feature-level methods are found hard to decode the structured information for semantic segmentation tasks due to the lack of low-dimensional information. Therefore, we propose an adversarial domain adaptation framework with Kullback–Leibler constraint (KL-ADDA) for remote sensing land cover classification. A state-of-the-art (SOTA) semantic segmentation network is utilized as the generator, which directly outputs the segmentation results to the discriminator to retain more low-level information. Besides, a Kullback–Leibler (KL)-divergence is calculated to improve the discriminative ability of the discriminator and thus enhance the generator’s performance. Experiments on the international society for photogrammetry and remote sensing (ISPRS) data set and two simulated target data sets have shown the effectiveness of KL-ADDA for DA. Mengxi Liu 0001, Pengyuan Zhang, Qian Shi 0001, Mengwei Liu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Learning Token-Aligned Representations With Multimodel Transformers for Different-Resolution Change DetectionabstractDifferent-resolution change detection (DRCD) is now becoming an urgent problem to be solved, which is of great potential in rapid monitoring, such as disaster assessment, urban expansion, etc. In DRCD tasks, bi-temporal inputs are given in the form of different resolutions, thus conventional CD methods cannot be applied directly. Previous studies have attempted to deal with this problem by reconstructing the low-resolution (LR) image into a high-resolution (HR) one, including interpolation and super-resolution (SR). However, these solutions are limited by the availability of training data, making it hard to meet different kinds of needs. Besides, these image-level strategies have also ignored the interaction and alignment of high-level features. Therefore, we propose a new approach based on multi-model Transformers (MM-Trans), which solves the resolution gaps of bi-temporal inputs in DRCD tasks from the perspective of feature alignment. In the MM-Trans, a weight-unshared feature extractor is first utilized to precisely capture the features of the different-resolution inputs; then a spatial-aligned Transformer (sp-Trans) is introduced to align the LR-image features to the same size of the HR-image ones, which can be optimized in a learnable way by an auxiliary token loss; after that, a semantic-aligned Transformer (se-Trans) is adopted, in which the bi-temporal features can be further interacted and aligned semantically; finally, a prediction head is employed to obtain fine-grained change results. Experiments conducted on three common CD datasets, CDD, S2Looking, and HTCD dataset, have shown the advancement of the MM-Trans and fully demonstrated its potential in DSCD tasks. Mengxi Liu 0001, Qian Shi 0001, Zhuoqun Chai |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Super-Resolution-Based Change Detection Network With Stacked Attention Module for Images With Different ResolutionsabstractChange detection (CD) aims to distinguish surface changes based on bitemporal images. Since high-resolution (HR) images cannot be typically acquired continuously over time, bitemporal images with different resolutions are often adopted for CD in practical applications. Traditional subpixel-based methods for CD using images with different resolutions may lead to substantial error accumulation when the HR images are employed, which is because of intraclass heterogeneity and interclass similarity. Therefore, it is necessary to develop a novel method for CD using images with different resolutions that are more suitable for the HR images. To this end, we propose a super-resolution-based change detection network (SRCDNet) with a stacked attention module (SAM). The SRCDNet employs a super-resolution (SR) module containing a generator and a discriminator to directly learn the SR images through adversarial learning and overcome the resolution difference between the bitemporal images. To enhance the useful information in multiscale features, a SAM consisting of five convolutional block attention modules (CBAMs) is integrated to the feature extractor. The final change map is obtained through a metric learning-based change decision module, wherein a distance map between bitemporal features is calculated. Ablation study and comparative experiments on two large datasets, building change detection dataset (BCDD) and season-varying change detection dataset (CDD), and a real-image experiment on the Google dataset fully demonstrate the superiority of the proposed method. The source code of SRCDNet is available athttps://github.com/liumency/SRCDNet. Mengxi Liu 0001, Qian Shi 0001, Andrea Marinoni, Da He, Xiaoping Liu 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | LCS: A Collaborative Optimization Framework of Vector Extraction and Semantic Segmentation for Building ExtractionabstractIn the field of building extraction, many CNN-based methods have been developed to solve the problem of the irregular boundaries in their predictions. The prevailing approach is to build an additional edge segmentation branch or obtain accurate vector components of buildings. However, pixel-based methods still cannot obtain accurate location of the boundary, while vector extraction will bring the problem of sample imbalance and missing detection. In this work, we utilize the complementarity of the two types of methods and propose the line segment collaborate segmentation (LCS) framework. In the proposed LCS framework, semantic segmentation provides location guidance for vector extraction, while vector extraction provides precise positioning for semantic segmentation. By this way, the two tasks can leverage their respective strengths. The results on three datasets show that the performance of vector extraction and semantic segmentation is improved simultaneously using the LCS framework, which proves the effectiveness of our method. At the same time, our framework is flexible and can be embedded in other vector extraction methods to improve performance. Qian Shi 0001, Jinpei Ou |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Deeply Supervised Attention Metric-Based Network and an Open Aerial Image Dataset for Remote Sensing Change DetectionabstractChange detection (CD) aims to identify surface changes from bitemporal images. In recent years, deep learning (DL)-based methods have made substantial breakthroughs in the field of CD. However, CD results can be easily affected by external factors, including illumination, noise, and scale, which leads to pseudo-changes and noise in the detection map. To deal with these problems and achieve more accurate results, a deeply supervised (DS) attention metric-based network (DSAMNet) is proposed in this article. A metric module is employed in DSAMNet to learn change maps by means of deep metric learning, in which convolutional block attention modules (CBAM) are integrated to provide more discriminative features. As an auxiliary, a DS module is introduced to enhance the feature extractor’s learning ability and generate more useful features. Moreover, another challenge encountered by data-driven DL algorithms is posed by the limitations in change detection datasets (CDDs). Therefore, we create a CD dataset, Sun Yat-Sen University (SYSU)-CD, for bitemporal image CD, which contains a total of 20 000 aerial image pairs of size$256\times256$. Experiments are conducted on both the CDD and the SYSU-CD dataset. Compared to other state-of-the-art methods, our network achieves the highest accuracy on both datasets, with an F1 of 93.69% on the CDD dataset and 78.18% on the SYSU-CD dataset. Qian Shi 0001, Mengxi Liu 0001, Shengchen Li, Xiaoping Liu 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Global Snow Depth Retrieval From Passive Microwave Brightness Temperature With Machine Learning ApproachabstractCurrent global snow retrieval algorithms based on spaceborne microwave measurements inherit noticeable biases and uncertainties regarding spatial distribution and temporal variations. In this article, we present an improved spatiotemporally dynamic global snow depth retrieval algorithm to account for the heterogeneity of snowpacks in different seasons worldwide. The proposed model adopts nonlinear machine learning to retrieve snow depths from passive microwave measurements and other auxiliary information. We indirectly characterized the variation in snow grain size using the daily profiles of the temperature gradient within the snowpack. In addition, a zoning and multitemporal modeling strategy was employed to reduce the bias and uncertainty caused by snow heterogeneity across different ecoregions and seasons. The proposed model was implemented to retrieve the global daily snow depth from 2001 to 2010. The results were validated byin situobservations and compared with the NASA Advanced Microwave Scanning Radiometer for EOS (AMSR-E) snow water equivalent product (AE_DySno). Satisfactory accuracy was achieved for different ecoregions with regard to daily, monthly, and yearly validations (the root-mean-square error (RMSE) varied from ~7.5 to ~12 cm; the Pearson correlation coefficient$R$ranged from 0.75 to 0.85). The results of ten trials indicated the promising stability of the proposed model in different ecoregions with small variations in RMSE and$R$values. Compared with the AE_DySno products, the estimation results did not exhibit the overestimation problem and provided snow depth patterns with greater spatial heterogeneity, showing RMSEs ~5 cm lower and$R$values ~0.3 higher than those of the AE_DySno products. Xiaocong Xu, Xiaoping Liu 0001, Xia Li 0001, Qian Shi 0001, Yimin Chen 0001, Bin Ai |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Spatial Downscaling of IMERG Considering Vegetation Index Based on Adaptive Lag PhaseabstractHigh spatial resolution precipitation data are important for hydrological modeling and meteorological applications, especially at regional scales. Statistical downscaling methods for satellite precipitation products using the normalized difference vegetation index (NDVI) have been carried out in many regions to provide high spatial resolution precipitation. These methods generally use NDVI and precipitation at the same time, assuming that there is a real-time response of vegetation to precipitation. However, this assumption does not hold in many scenarios. It is known that different vegetation types exhibit different response times to precipitation, i.e., there is a possible lag in the response of vegetation to precipitation depending on the vegetation/landcover type. Therefore, it is not appropriate to estimate precipitation using NDVI collected at the same time. To better represent the relationship between precipitation and vegetation, this article develops a new vegetation index based on adaptive lag phase (VIAL) estimated from a new growth rate that is adaptive to landcover type. Based on VIAL, a new local precipitation downscaling method called LPVIAL is proposed, which essentially considers the nonstationary relationship between precipitation and VIAL. The performance of LPVIAL is assessed by downscaling Integrated Multi-satellitE Retrievals for Global Precipitation Measurement (IMERG) from 0.1° to 1-km spatial resolution over the Pearl River Basin in Southern China from 2010 to 2017 at 16-day temporal resolution, and the downscaled products are validated against ground observations. Results indicate that the high-resolution precipitation data obtained from the new downscaling approach perform well, and the accuracy is higher than traditional approaches. With the enhancement of spatial resolution, LPVIAL downscaled products show more detailed spatial information of precipitation with smooth distribution, and the downscaled products have slightly higher accuracy compared with IMERG. It is, therefore, suggested that the adaptive lag phase should be considered in the satellite precipitation product downscaling process. Zhaozhao Zeng, Haonan Chen 0001, Qian Shi 0001, Jun Li 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Rethinking the High Frequency Components in Deep Sub-Pixel Mapping NetworkabstractDeep sub-pixel mapping network (DSMNet) is a state-of-the-art approach in the field of sub-pixel mapping (SPM, also called super resolution mapping), combining deep learning theory, to solve the mixed pixel problem, which is ubiquitous in remote sensing images due to the spatial-resolving limitation. However, traditional DSMNet usually do not consider the multi-scale distribution characteristics of the real geographical distribution exposed in urban landscape. Furthermore, the heterogeneous distribution characteristics (high-frequency components) are the most important for SPM, but are difficult to learn and usually ignored in the tradition network models. In this paper, the high-frequency component aware (HFCA) module was proposed, based on the hierarchical supervised deep sub-pixel mapping network (HiDSMNet). HiDSMNet establishes a hierarchical supervised architecture for explicit multi-scale supervision to prompt the network to learn a multi-scale representation. Besides, HFCA module is integrated to prompt the network to intensify the learning of the high-frequency representation. The experimental results with three public datasets validated the superiority of the proposed HiDSMNet. Da He, Yanfei Zhong, Qian Shi 0001, Xiaoping Liu 0001 |
IGARSS | 3 |
| 2021 | DSAMNet: A Deeply Supervised Attention Metric Based Network for Change Detection of High-Resolution ImagesabstractIn view of the insufficient of current change detection, we propose a deeply-supervised attention metric-based network (DSAMNet) for bi-temporal image change detection. The DSAMNet contains a CBAM integrated change decision module to learn a change map directly from features from feature extractor, and an auxiliary deep supervision module to generate intermediate change results to help the training of hidden layers. We also provide a new benchmark-SYSU-CD-with totally 20000 image pairs for the training and testing of deep learning based CD methods. Comparative experiments on the SYSU-CD dataset have proved the effectiveness of the proposed method. Mengxi Liu 0001, Qian Shi 0001 |
IGARSS | 2 |
| 2021 | Multi-Label Local Climate Zone Mapping as Scene Classification Using Very High Resolution Imagery: Preliminary Result of Hong KongabstractWith the near completion of WUDAPT (World Urban Database and Access Portal Tools) Level 0 data, one of the next goals is to generate more accurate and detailed local climate zone (LCZ) maps. An important issue is how to integrate building height information into LCZ maps. We here present a multi-label classification method using very high resolution (VHR) imagery to implicitly integrate building height information. Since we humans can tell whether a place is high-rise or not based on the shading of buildings and the surrounding context, it is possible to extract such information using deep learning methods. We use Hong Kong as a case study and show the potential of LCZ mapping with VHR imagery in distinguishing small-scale landscape features like city parks. The multi-label LCZ maps also provide a solution to generate fine-grained subclass LCZ mapping, in which a place can be classified as a combination of multiple LCZs, e.g., compact low-rise with open high-rise. Shengjie Liu 0001, Qian Shi 0001 |
IGARSS | 2 |
| 2021 | Active Ensemble Deep Learning for Polarimetric Synthetic Aperture Radar Image ClassificationabstractAlthough deep learning has achieved great success in the image-classification tasks, its performance is subject to the quantity and quality of the training samples. For the classification of the polarimetric synthetic aperture radar (PolSAR) images, it is nearly impossible to annotate the images from visual interpretation. Therefore, it is urgent for remote-sensing scientists to develop new techniques for PolSAR image classification under the condition of very few training samples. In this letter, we take the advantage of active learning and propose active ensemble deep learning (AEDL) for PolSAR image classification. We first show that only 35% of the predicted labels of the deep-learning model's snapshots near its convergence were exactly the same. The disagreement between the snapshots is nonnegligible. From the perspective of multiview learning, the snapshots together serve as a good committee to evaluate the importance of the unlabeled instances. Using the snapshot committee to give out the informativeness of the unlabeled data, the proposed AEDL achieved better performance on two real PolSAR images than the standard active learning strategies. It achieved the same classification accuracy with only 86% and 55% of the training samples compared to the breaking tie active learning and random selection for the Flevoland data set. Shengjie Liu 0001, Haowen Luo, Qian Shi 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Scene-Driven Multitask Parallel Attention Network for Building Extraction in High-Resolution Remote Sensing ImagesabstractThe application of convolutional neural networks has been shown to significantly improve the accuracy of building extraction from very high-resolution (VHR) remote sensing images. However, there exist so-called semantic gaps among different kinds of buildings due to the large intraclass variance of buildings, and most of the present-day methods are ineffective in extracting various buildings in large areas that cover different scenes, for example, urban villages and high-rise buildings, because existing building extraction strategies are the same for various scenes. With the improvement of the resolution of remote sensing images, it is feasible to improve the image interpretation based on the scene prior. However, this idea has not been fully utilized in building extraction from VHR remote sensing imagery. This study proposes a scene-driven multitask parallel attention convolutional network (MTPA-Net) to resolve these limitations. The proposed approach classifies the input image into multilabel scenes and further separately maps the buildings in pixel level under different scenes. In addition, a simple postprocessing method is applied to integrate the building extraction results and scene prior. Our proposed method does not require multimodel training and the network can learn in an end-to-end manner. The performance of our proposed method is evaluated on a data set that includes various urban and rural scenes with diverse landscapes. The experimental results show that the proposed MTPA-Net outperforms state-of-the-art algorithms by reducing misclassification areas and maintaining improved robustness. Qian Shi 0001, Bo Du 0001, Liangpei Zhang 0001, Dongzhi Wang, Huaxiang Ding |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Correction to "Scene-Driven Multitask Parallel Attention Network for Building Extraction in High-Resolution Remote Sensing Images"abstractPresents corrections to the above named paper. Qian Shi 0001, Bo Du 0001, Liangpei Zhang 0001, Dongzhi Wang, Huaxiang Ding |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Deep Subpixel Mapping Based on Semantic Information Modulated Network for Urban Land Use MappingabstractMixed pixel problem is omnipresent in remote sensing images for urban land use interpretation due to the hardware limitations. Subpixel mapping (SPM) is a usual way to solve this problem by improving the observation scale and realizing a finer spatial resolution land cover mapping. Recently, deep learning-based subpixel mapping network (DLSMNet) was proposed, benefited from its strong representation and learning ability, to restore a visually pleasing finer mapping. However, the spatial context features of artifacts are usually aggregated and progressively lost during the forward pass of the network without sufficient representation, which make it difficult to be learned and restored. In this article, a semantic information modulated (SIM) deep subpixel mapping network (SIMNet) is proposed, which uses low-resolution semantic images as prior, to reinforce the representation of spatial context features. In SIMNet, SIM module is proposed to parametrically incorporate the semantic prior into the state-of-the-art (SOTA) feed forward network architecture in an end-to-end training fashion. Furthermore, stacked SIM module with residual blocks (SIM_ResBlock) is adopted to pass the representation of spatial context feature to the deep layers, to get it fully learned during backpropagation. Experiments have been implemented on three public urban scenario data sets, and the SIMNet generates a clearer outline of artificial facilities with sufficient spatial context, and is distinctive for even individual building, which is challenging for other SOTA DLSMNet. The results demonstrate that the proposed SIMNet is a promising way for high-resolution urban land use mapping from easily available lower resolution remote sensing images.Mixed pixel problem is omnipresent in remote sensing images for urban land use interpretation due to the hardware limitations. Subpixel mapping (SPM) is a usual way to solve this problem by improving the observation scale and realizing a finer spatial resolution land cover mapping. Recently, deep learning-based subpixel mapping network (DLSMNet) was proposed, benefited from its strong representation and learning ability, to restore a visually pleasing finer mapping. However, the spatial context features of artifacts are usually aggregated and progressively lost during the forward pass of the network without sufficient representation, which make it difficult to be learned and restored. In this article, a semantic information modulated (SIM) deep subpixel mapping network (SIMNet) is proposed, which uses low-resolution semantic images as prior, to reinforce the representation of spatial context features. In SIMNet, SIM module is proposed to parametrically incorporate the semantic prior into the state-of-the-art (SOTA) feed forward network architecture in an end-to-end training fashion. Furthermore, stacked SIM module with residual blocks (SIM_ResBlock) is adopted to pass the representation of spatial context feature to the deep layers, to get it fully learned during backpropagation. Experiments have been implemented on three public urban scenario data sets, and the SIMNet generates a clearer outline of artificial facilities with sufficient spatial context, and is distinctive for even individual building, which is challenging for other SOTA DLSMNet. The results demonstrate that the proposed SIMNet is a promising way for high-resolution urban land use mapping from easily available lower resolution remote sensing images. Da He, Qian Shi 0001, Xiaoping Liu 0001, Yanfei Zhong, Xinchang Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Few-Shot Hyperspectral Image Classification With Unknown Classes Using Multitask Deep LearningabstractCurrent hyperspectral image classification assumes that a predefined classification system is closed and complete, and there are no unknown or novel classes in the unseen data. However, this assumption may be too strict for the real world. Often, novel classes are overlooked when the classification system is constructed. The closed nature forces a model to assign a label given a new sample and may lead to overestimation of known land covers (e.g., crop area). To tackle this issue, we propose a multitask deep learning method that simultaneously conducts classification and reconstruction in the open world (named MDL4OW) where unknown classes may exist. The reconstructed data are compared with the original data; those failing to be reconstructed are considered unknown based on the assumption that they are not well represented in the latent features due to the lack of labels. A threshold needs to be defined to separate the unknown and known classes; we propose two strategies based on the extreme value theory for few- and many-shot scenarios. The proposed method was tested on real-world hyperspectral images; state-of-the-art results were achieved, e.g., improving the overall accuracy by 4.94% for the Salinas data. By considering the existence of unknown classes in the open world, our method achieved more accurate hyperspectral image classification, especially under the few-shot context. Shengjie Liu 0001, Qian Shi 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Hyperspectral Image Denoising Using a 3-D Attention Denoising NetworkabstractHyperspectral image (HSI) denoising plays an important role in image quality improvement and related applications. Convolutional neural network (CNN)-based image denoising methods have been predominant due to advances made in the field of deep learning in recent years. Spatial and spectral information are crucial to HIS denoising, along with their correlations. However, existing methods fail to consider the global dependence and correlation between spatial and spectral information. Accordingly, in this article, we propose a novel dual-attention denoising network to overcome these limitations. We design two parallel branches to process the spatial and spectral information separately. The position attention module is applied to the spatial branch to formulate the interdependencies on the feature map, while the channel attention module is applied to the spectral branch to simulate the spectral correlation before the two branches are combined. A multiscale structure is also employed to extract and fuse the multiscale features following the fusion of spatial and spectral information. Experimental results on simulated and real data substantiate the superiority of our method both visually and quantitatively when compared with state-of-the-art methods. Qian Shi 0001, Xiaopei Tang, Taoru Yang, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Siamese Generative Adversarial Network for Change Detection Under Different ScalesabstractChange detection methods based on low-resolution (LR) images with higher temporal resolution often lead to fuzzy results, while high-resolution images (HRIs) can provide more detailed information to solve this problem. However, it's hard to obtain two tiles of HRIs with high-quality for rapid change detection in actual production due to low temporal resolution and high cost. Therefore, it is necessary to explore a change detection method combing low- and high-resolution images to acquire urban change areas more accurately and quickly. In this paper, an end-to-end siamese generative adversarial network (SiamGAN) integrating a super resolution network and the siamese structure was proposed for change detection under different scales. The super-resolution network is used to reconstruct low-resolution images into high-resolution images, while the siamese structure is adopted as the classification network to detect changes. In the experiments, SiamGAN achieved an F1 of 76.06% and an IoU of 61.52% in the test set, which is respectively 5.68% and 6.92% higher than the CNN-based methods using LR images after bicubic interpolation. The results show that our proposed method can effectively overcome difference in scale between low- and high-resolution images and perform change detection more precisely and rapidly. Mengxi Liu 0001, Qian Shi 0001, Penghua Liu |
IGARSS | 2 |
| 2020 | Improved Guided Source Separation Integrated with a Strong Back-End for the CHiME-6 Dinner Party Scenario
Hangting Chen, Pengyuan Zhang, Qian Shi 0001, Zuozhen Liu |
INTERSPEECH | 3 |
| 2020 | Object-Oriented Mangrove Species Classification Using Hyperspectral Data and 3-D Siamese Residual NetworkabstractMangrove species classification is of particular importance for coastal conservation and restoration. However, it is challenging to distinguish species-level differences with limited training data. In this letter, we propose an object-oriented classification method for mangrove forests by using the hyperspectral image (HSI) and the 3-D Siamese residual network. First, superpixel segmentation is utilized to obtain objects with various shapes and scales. Second, 3-D patches of each object are extracted from the original HSI, and those patches containing training samples are adopted to pairwise train the network. The 3-D spatial pyramid pooling (3-D-SPP) is added in the network to extract features in multiple scales. Finally, the abstract features of test samples are learned by the trained network, and the labels are determined by the nearest neighbor classifier within the metric space. Experiments on real mangrove hyperspectral data demonstrate the effectiveness of the proposed method in species classification of mangroves. Zhi He, Qian Shi 0001, Kai Liu 0003, Jingjing Cao, Wen Zhan, Beifen Cao |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | Multitask Deep Learning With Spectral Knowledge for Hyperspectral Image ClassificationabstractIn this letter, we propose a multitask deep learning method for the classification of multiple hyperspectral data in a single training. Deep learning models have achieved promising results on hyperspectral image classification, but their performance highly relies on sufficient labeled samples that are scarce on hyperspectral images. However, samples from multiple data sets might be sufficient to train one deep learning model, thereby improving its performance. To do so, we trained an identical feature extractor for all data, and the extracted features were fed into corresponding softmax classifiers. Spectral knowledge was introduced to ensure that the shared features were similar across domains. Four hyperspectral data sets were used in the experiments. We achieved higher classification accuracies on three data sets (Pavia University, Pavia Center, and Indian Pines) and competitive results on the Salinas Valley data compared with the baseline. Spectral knowledge was useful to prevent the deep network from overfitting when the data shared similar spectral response. The proposed method tested on two deep CNNs successfully shows its ability to utilize samples from multiple data sets and to enhance networks' performance. Shengjie Liu 0001, Qian Shi 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | Domain Adaption for Fine-Grained Urban Village Extraction From Satellite ImagesabstractUrban villages (UVs) are distinctive products formed in the process of rapid urbanization. The fine-grained mapping of UVs from satellite images has always been a considerable challenge because of the complex urban structures and the insufficiency of labeled samples. In this letter, we propose using the domain adaptation strategy to tackle the domain shift problem by employing adversarial learning to tune the semantic segmentation network so as to adaptively obtain similar outputs for input images from different domains. The proposed method was coupled with several segmentation networks, including U-Net, RefineNet, and DeepLab v3+, and the results show that domain adaptation can significantly improve the pixel-level mapping of UVs. Qian Shi 0001, Mengxi Liu 0001, Xiaoping Liu 0001, Penghua Liu, Pengyuan Zhang, Jinxing Yang, Xia Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2019 | Accessibility-Free Active Learning for Hyperspectral Image ClassificationabstractThis work proposes a new collaborative active and semi-supervised learning approach, named accessibility-free active learning (AFAL), for hyperspectral imaging classification. The proposed approach aims to tackle an existing problem in traditional active learning methods, that is, the fact that some selected samples are not accessible by oracles for assigning them pseudo labels, i.e., confident predictions for the classifier. The proposal specifically addresses this problem using superpixels in a self-training context. Specifically, AFAL first generates a set of candidates locally around the labeled pixels and then expands them to other subregions via a density peak-based augmentation strategy, in order to guarantee the confidence of pseudo labels. Our experimental results, obtained on two real and well-used hyperspectral images, reveal that the proposed scheme can lead to state-of-the-art performance. Chenying Liu 0001, Jun Li 0009, Mercedes Eugenia Paoletti, Juan Mario Haut, Antonio Plaza, Qian Shi 0001 |
IGARSS | 6 |
| 2019 | Regularised transfer learning for hyperspectral image classificationabstractThis study presents a transfer learning method for addressing the insufficient sample problem in hyperspectral image classification. In order to find common feature representation for both the source domain and target domain, we introduce a regularisation based on Bregman divergence into the objective function of the subspace learning algorithm, which can minimise the Bregman divergence between the distribution of training samples in the source domain and the test samples in the target domain. Hyperspectral image with biased sampling is used to evaluate the effectiveness of the proposed method. The results show that the proposed method can achieve a higher classification accuracy than traditional subspace learning methods under the condition of biased sampling. Qian Shi 0001, Xiaoping Liu 0001, Kefei Zhao |
IET Comput. Vis. | 1 |
| 2019 | Domain Adaptation With Discriminative Distribution and Manifold Embedding for Hyperspectral Image ClassificationabstractHyperspectral remote sensing image classification has drawn a great attention in recent years due to the development of remote sensing technology. To build a high confident classifier, the large number of labeled data is very important, e.g., the success of deep learning technique. Indeed, the acquisition of labeled data is usually very expensive, especially for the remote sensing images, which usually needs to survey outside. To address this problem, in this letter, we propose a domain adaptation method by learning the manifold embedding and matching the discriminative distribution in source domain with neural networks for hyperspectral image classification. Specifically, we use the discriminative information of source image to train the classifier for the source and target images. To make the classifier can work well on both domains, we minimize the distribution shift between the two domains in an embedding space with prior class distribution in the source domain. Meanwhile, to avoid the distortion mapping of the target domain in the embedding space, we try to keep the manifold relation of the samples in the embedding space. Then, we learn the embedding on source domain and target domain by minimizing the three criteria simultaneously based on a neural network. The experimental results on two hyperspectral remote sensing images have shown that our proposed method can outperform several baseline methods. Zengmao Wang, Bo Du 0001, Qian Shi 0001, Weiping Tu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2018 | An Active Relearning Framework for Remote Sensing Image ClassificationabstractClassification is an important technique for remote sensing data interpretation. In order to enhance the performance of a supervised classifier and ensure the lowest possible cost of the training samples used in the process, active learning (AL) can be used to optimize the training sample set. At the same time, integrating spatial information can help to enhance the separability between similar classes, which can in turn reduce the need for training samples in AL. To effectively integrate spatial information into the AL framework, this paper proposes a new active relearning (ARL) model for remote sensing image classification. In particular, our model is used to relearn the spatial features on the classification map, which contributes significantly to enhancing the performance of the classifier. We integrate the relearning model into the AL framework, with the aim to accelerate the convergence of AL and further reduce the labeling cost. Under the newly developed ARL framework, we propose two spatial–spectral uncertainty criteria to optimize the procedure for selecting new training samples. Furthermore, an adaptive multiwindow ARL model is also introduced in this paper. Our experiments with two hyperspectral images and two very high resolution images indicate that the ARL model exhibits faster convergence speed with fewer samples than traditional AL methods. Our results also suggest that the proposed spatial–spectral uncertainty criteria and the multiwindow version can further improve the performance when implementing ARL. Qian Shi 0001, Xiaoping Liu 0001, Xin Huang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Active learning approach for remote sensing imagery classification using spatial informationabstractIn the last few years, integrating spatial information into active learning framework has been gaining growing interest in the remote sensing community to optimize the collection of training sample set for supervised image classification. We address this problem from two directions. One of the directions focus on improving the classifier's performance to reduce the need for training samples. For this purpose, relearning model is introduced to combine the active learning framework to form mutually reinforcing process. In the meantime, another direction focus on the way to select most informative samples. For this purpose, new uncertainty criterion is proposed to favor the selection of samples not only with most spectral uncertainty, but also located in most uncertain spatial regions. Experiments on hyperspectral image show the effectiveness of proposed active learning framework. Qian Shi 0001, Xin Huang 0002, Jiayi Li 0001, Liangpei Zhang 0001 |
IGARSS | 1 |
| 2015 | Domain Adaptation for Remote Sensing Image Classification: A Low-Rank Reconstruction and Instance Weighting Label Propagation Inspired AlgorithmabstractThis paper presents a framework for a semisupervised domain adaptation method for remote sensing image classification. Most of the representation-based domain adaptation methods attempt to find a total transformation matrix for all the samples from the source domain; however, they ignore the individual changes in each class, which often leads to the misalignment of the samples in each class between the two domains. This paper attempts to find new representations for the samples in different classes from the source domain by multiple linear transformations, which corresponds to the practical changes in each class to a higher degree. Furthermore, to avoid the influence of outliers and noise in the source domain samples, low-rank reconstruction is further applied to make the domain adaptation method more robust. In addition, in the stage of predicting the unlabeled samples by label propagation (LP), the proposed LP with instance weighting can effectively further reduce the negative effect of misleading samples from the source domain. The results obtained with a QuickBird data set and a hyperspectral data set confirm the effectiveness and reliability of the proposed method. Qian Shi 0001, Bo Du 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Spatial Coherence-Based Batch-Mode Active Learning for Remote Sensing Image ClassificationabstractBatch-mode active learning (AL) approaches are dedicated to the training sample set selection for classification, regression, and retrieval problems, where a batch of unlabeled samples is queried at each iteration by considering both the uncertainty and diversity criteria. However, for remote sensing applications, the conventional methods do not consider the spatial coherence between the training samples, which will lead to the unnecessary cost. Based on the above two points, this paper proposes a spatial coherence-based batch-mode AL method. First, mean shift clustering is used for the diversity criterion, and thus the number of new queries can be varied in the different iterations. Second, the spatial coherence is represented by a two-level segmentation map which is used to automatically label part of the new queries. To get a stable and correct second-level segmentation map, a new merging strategy is proposed for the mean shift segmentation. The experimental results with two real remote sensing image data sets confirm the effectiveness of the proposed techniques, compared with the other state-of-the-art methods. Qian Shi 0001, Bo Du 0001, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | Spatial correlated information based batch mode active learning method for remote sensing image classificationabstractBatch-mode active learning approaches are dedicated to the problem of training sample set selection, where a batch of unlabeled samples is queried at each iteration by considering both uncertainty and diversity criteria. However, the current batch-mode approaches do not consider spatial correlation between adjacent queries pixels, thus they spend some unnecessary time costs and are accompanied by relatively high annotation costs. This paper employs mean shift segmentation to describe the spatial correlation information which is used to select most diverse samples in the geographic space and to automatically label part of the pixels that need querying. As a result, the labeling costs can be lowered sharply. Meanwhile, the number of new queries in each iteration is adaptive to the distribution of the uncertain samples, which can reduce the iterations. Experimental results obtained in the classification of a hyperspectral image confirm the effectiveness of the proposed technique. Qian Shi 0001, Liangpei Zhang 0001, Bo Du 0001 |
IGARSS | 1 |
| 2013 | Semisupervised Discriminative Locally Enhanced Alignment for Hyperspectral Image ClassificationabstractThis paper proposes a new semisupervised dimension reduction (DR) algorithm based on a discriminative locally enhanced alignment technique. The proposed DR method has two aims: to maximize the distance between different classes according to the separability of pairwise samples and, at the same time, to preserve the intrinsic geometric structure of the data by the use of both labeled and unlabeled samples. Furthermore, two key problems determining the performance of semisupervised methods are discussed in this paper. The first problem is the proper selection of the unlabeled sample set; the second problem is the accurate measurement of the similarity between samples. In this paper, multilevel segmentation results are employed to solve these problems. Experiments with extensive hyperspectral image data sets showed that the proposed algorithm is notably superior to other state-of-the-art dimensionality reduction methods for hyperspectral image classification. Qian Shi 0001, Liangpei Zhang 0001, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |