VLDB 2026 Research / reviewers in the wild / expert
Qiqi Zhu
dblp:49/2637
· DBLP profile ↗
43ranked-venue papers
17as first author
27since 2021 · last 2025
0000-0002-5339-0829ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 39 · 15 first-author · 25 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MLPET: a Multi-Level Prompt-Enhanced Transformer for Unified Molecular Property and Drug-Drug Interaction Event PredictionabstractAccurate prediction of molecular properties and drug-drug interaction (DDI) events is crucial for drug discovery and pharmacological safety assessment. However, existing approaches typically focus on single-level molecular representation and lack generalizable strategies for multi-task knowledge transfer. In this work, we propose a unified Multi-Level PromptEnhanced Transformer framework (MLPET), jointly modeling both multi-level molecular features and inter-drug relational semantics. To capture multi-level structural knowledge, we design three complementary self-supervised pretraining tasks: (1) atom prediction, which predicts randomly masked atom types to learn local chemical context; (2) bond prediction, which infers the existence of chemical bonds between atom pairs to capture topological dependencies; and (3) distance prediction, which regresses the 3D spatial distance between atoms to encode molecular geometry. In the downstream stage, we introduce a lightweight prompt fusion mechanism that integrates task-specific prompts into a unified vector to guide fine-tuning. This enables flexible and efficient knowledge transfer to multiple tasks, including molecular property prediction and multi-class DDI event classification. Extensive experiments on MoleculeNet, Ryu, and Deng benchmarks demonstrate that MLPET consistently outperforms state-of-the-art baselines, particularly in low-resource and longtail scenarios. Our results highlight the potential of promptguided multi-task pretraining as a generalizable paradigm for molecular representation learning. The source code of MLPET is available at https://github.com/fuhaitao95/MLPET. Qiqi Zhu, Haitao Fu |
BIBM | 1 |
| 2025 | DistRMI: a deep distance-aware neural network for explainable RNA loop motif-small molecule interaction predictionabstractRNA participates in the occurrence and development of various diseases by regulating gene expression. Owing to its potential to circumvent the limitations of traditional "undruggable" protein targets, it is regarded as a core direction for next-generation precision therapy. Against this backdrop, the accurate and interpretable prediction of RNA-small molecule interactions has become a key link in accelerating the discovery of RNA-targeted drugs. However, existing methods suffer from insufficient prediction accuracy and interpretability, failing to effectively guide lead compound screening or elucidate the mechanism of action. This study presents DistRMI, which integrates Transformers and graph neural networks to capture, respectively, the sequence information of RNA loop motifs and the chemical topological features of small molecules while introducing distance priors between them and leveraging a distance-aware attention mechanism to capture their interaction information. The results show that DistRMI outperforms baseline models, and its performance remains robust even when confronted with unknown RNA loop motifs and small molecules. Visualization of attention weights reveals that bases near the paired bases of RNA loop motifs contribute significantly. Furthermore, retrospective case studies validate the model's reliability. Predicting the binding preferences between RNA loop motifs and small molecules while providing interpretability facilitates an in-depth understanding of RNA-small molecule interactions, promotes in-depth research on RNA and related drugs, and opens up new avenues for disease treatment. Zhaoxiang Liu, Qiqi Zhu, Qingyan Tian, Yingxiang Deng, Dengguo Wei, Haitao Fu |
Briefings Bioinform. | 2 |
| 2025 | VideoMamba++: Integrating state space model with dual attention for enhanced video understanding
Wang Tian, Qiqi Zhu, Xianglong Zhang |
Image Vis. Comput. | 3 |
| 2025 | Geographical Dual-Prior Guided Few-Shot Network for Cross-Domain Hyperspectral Image ClassificationabstractCross-domain hyperspectral image (HSI) classification (HSIC) addresses the challenge of real-time labeling of new regions. To mitigate the performance decline caused by unseen classes, a few-shot learning (FSL) method is used. However, these methods fail to fully consider the problem of sample scarcity and classification imbalance due to FSL methods. In addition, the issue of category confusion stemming from localized spectral fluctuations within the same class is commonly overlooked. To solve these problems, a geographical dual-prior guided few-shot network (Gprior-FSN) is proposed. In Gprior-FSN, combining prior knowledge of the first law of geography, a geographical prior guided bicorrelated (G-B) sample enhancement mechanism is proposed which includes geospatially correlated enhancement (GCE) and spectral feature correlated enhancement (SFCE). GCE uses a hierarchical sampling strategy to tackle the inherent imbalance problem for FSL methods. Subsequently, GCE mitigates sample scarcity via neighborhood sample expansion while identifying candidate pseudosamples with geospatial correlation. To make the acquired pseudosamples of the same category bicorrelated in both geospatial and spectral features, G-B combining spectral feature clustering and probabilistic statistics mechanism is designed. Inspired by the second law of geography, Gprior-FSN uses a spatial constraint mechanism to effectively enhance intraclass similarity by reducing spatial local heterogeneity, while improving global interclass discriminability. Finally, to further capture representative spatial-spectral feature, a weighted dual feature fusion network is designed. Experimental results from three distinct HSI datasets show that Gprior-FSN outperforms advanced HSIC methods in both efficiency and accuracy. In addition, the Gprior-FSN demonstrates strong generalization performance on real GF-5 image. Weihuan Deng, Qiqi Zhu, Qingfeng Guan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | A Road-Detail Preserving Framework for Urban Road Extraction From VHR Remote Sensing ImageryabstractAutomatic road extraction has gained significant attention in urban navigation, sustainable transport, and disaster response. Conventional convolutional neural networks (CNNs) operate within the local receptive field, limiting their capacity to represent potential global relations between roads and surroundings. In addition, the edge is important topological information for road targets. Several works focus on predicting precise boundaries to enhance road extraction. However, over fit edges and the course integration between features of different network layers may lead to loss of local details and incorrect road segmentation results. Therefore, the Road-detail Preserving Mapper (RoadDP-Mapper) framework is proposed. First, RoadDP-Mapper employs a hierarchical transformer as the encoder to enable local-to-global reasoning. The asymmetric upsampling layers (APLs) are introduced to enhance the model’s capability to perceive and reconstruct critical road detail information. Second, a road edge-constrained branch with a detail preservation module (DPM) is devised to amplify the distinction between roads and backgrounds by extracting and preserving explicit class boundary details. The proposed joint loss inspires the transformer to capture the contextual spatial relationships while preserving the fine-grained features of the road. We evaluated our framework on the DeepGlobe dataset and self-annotated images from ten representative cities in China. The proposed framework has demonstrated its effectiveness by significantly reducing both missed detections and false alarms in road extraction. Furthermore, spatial transfer experiments have confirmed the generalizability of RoadDP-Mapper for large-scale road mapping. Qiqi Zhu, Sisi Peng, Longli Ran, Lizeng Wang, Jiancheng Luo |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | From Intra-Distinctiveness to Inter-Invariance: A Cycle-Resemblance Few-Shot Transformation Network for Cross-Domain Hyperspectral Image ClassificationabstractFor large-scale mapping applications, cross-domain hyperspectral image classification (HSIC) has emerged as a highly promising research area. However, the classification accuracy decreased significantly when unseen classes emerged. Few shot learning (FSL) methods are adopted in cross-domain HSIC methods to address this problem. Despite this, existing cross-domain HSIC methods still have three key issues that hamper their classification capabilities: 1) previous works struggle to balance incorporating distinctive intradomain knowledge and managing model complexity in the face of significant domain representation differences; 2) previous works inadequately consider the limited capture capacity of interdomain intrinsic mutually invariant structures; and 3) previous works fail to capture the distinct characteristics of both head categories (e.g., urban buildings) and tail categories (e.g., urban corn) simultaneously when applying FSL to deal with unseen classes problem. In this article, we propose a cycle-resemblance few-shot transformation (CF-Trans) network to effectively handle the aforementioned challenges by integrating intradomain distinctiveness with interdomain invariance. To facilitate efficient intradomain feature aggregation for HSI, a novel lightweight intradomain attentive network is introduced. Different from previous works, to reduce the negative impact caused by inaccurate classifier predictions, from the perspective of interdomain knowledge transformation, a cycle-resemblance adversarial network is designed to capture the intrinsic mutually invariant structures. A dynamic label expansion mechanism is designed to capture the distinctive intradomain features of the head and tail classes. Experimental results on six HSI datasets including agricultural, rural-urban and urban datasets show the remarkably performance of our network. Qiqi Zhu, Weihuan Deng, Qingfeng Guan 0001, Jiancheng Luo |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Spectral Correlation-Based Fusion Network for Hyperspectral Image Super-ResolutionabstractTo address the limitations of hyperspectral imaging systems, super-resolution (SR) techniques that fuse low-resolution hyperspectral image (HSI) with high-resolution multispectral image (MSI) are applied. Due to the significant modal difference between HSI and MSI, and the insufficient consideration of HSI band correlation in previous works, issues arise such as spectral distortions and loss of fine texture and boundaries. In this article, an unsupervised spectral correlation-based fusion network (SCFN) is proposed to address the above challenges. A new dense spectral convolution module (DSCM) is proposed to capture the intrinsic similarity dependence between spectral bands in HSI to effectively extract spectral domain features and mitigate spectral aberrations. To preserve the rich texture details in MSI, a global-local aware block (GAB) is designed for joint global contextual information and emphasize critical regions. To address the cross-modal disparity problem, new joint losses are constructed to improve the preservation of high-frequency information during image reconstruction and effectively minimize spectral disparity for more precise and accurate image reconstruction. The experimental results on three hyperspectral remote sensing datasets demonstrate that SCFN outperforms other methods in both qualitative and quantitative comparisons. Results from the fusion of real hyperspectral and multispectral remote sensing data further confirm the applicability and effectiveness of the proposed network. Qiqi Zhu, Meilin Zhang, Guizhou Zheng, Jiancheng Luo |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Local-Global Context-Aware Generative Dual-Region Adversarial Networks for Remote Sensing Scene Image Super-ResolutionabstractRecently, high-resolution (HR) remote sensing images have attracted increasing attention in a number of tasks. Super-resolution (SR) is an efficient method to obtain high-resolution remote sensing images. Due to the influence of imaging distances and angles, remote sensing images significantly differ from natural images in terms of land cover element distribution, ground object scale and scene complexity. This poses a challenge for capturing global and local low- and high-frequency and restoring fine image details for remote sensing image SR. In this article, a local-global context-aware generative dual-region adversarial network (LGC-GDAN) is designed for remote sensing image SR. It is composed of dual region-level discriminators and a dual-path generator with a context-aware network and an edge-assisted network. To capture global and local low- and high-frequency information, the global-aware self-attention (GAS) mechanism and local-aware self-attention (LAS) mechanism are introduced into the context-aware network. The GAS mechanism combines high-pass and low-pass filtering for long-range similarity feature, while LAS uses local aggregation for fine-level feature. The LR images and the corresponding edge maps are input to the edge-assisted network to extract the detailed geometric structure. To address small ground object and complex ground scenes, conventional image-level discriminators exhibit limited performance in capturing detailed information. Unlike previous discriminator, a region-level discriminator is designed to obtain the real/fake label of each local region. Moreover, two task-driven loss functions are designed to produce diverse images for further scene classification. The experiments undertaken on several remote sensing datasets demonstrate that LGC-GDAN outperforms the other state-of-the-art methods. Weihuan Deng, Qiqi Zhu, Qingfeng Guan 0001, Jiancheng Luo |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | A Cross-Domain Object-Semantic Matching Framework for Imbalanced High Spatial Resolution Imagery Water-Body ExtractionabstractLarge-scale information pertaining to surface water bodies is crucial for activities such as flood monitoring. Deep learning algorithms have shown great potential in water-body extraction based on high spatial resolution (HSR) imagery. However, the current reliance on deep learning for HSR imagery water-body extraction necessitates a substantial quantity of manually labeled training samples. The variance in spatial resolution among images and the intricacies of scenes consistently pose challenges to the transferability of deep learning. Moreover, the number of pixels representing water bodies is typically lower compared to the number of background pixels. This imbalance in class prediction probabilities often limits the accuracy of water-body class predictions. In this paper, we propose a cross-domain object-semantic matching (COM) framework for extracting water bodies from unlabeled high-resolution remote sensing imagery. The distinctions in spectra, shapes, and semantic distributions of water bodies across various domains create challenges for certain source domain samples to contribute positively to model training. Therefore, a sample semantic similarity matching mechanism is devised. The proposed object contextual perception network (OCPNet) models multi-scale water body features and object-contextual representations, aiming to achieve a more accurate and comprehensive representation of surface water bodies. Additionally, to prevent the training process from being dominated by easily transferred categories in the target domain, a weighted joint loss is designed to alleviate the imbalance of predicted probabilities and pixel numbers between water and non-water bodies. Experiments on four public datasets of GID, CCF, LoveDA and DeepGlobe demonstrate the effectiveness and generalization of our proposed framework. Zhen Li 0042, Qiqi Zhu, Jianjun Lv, Qingfeng Guan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | KO-Shadow: KnOwledge-Driven Shadow Progressive Removal Framework for Very High Spatial Resolution Remote Sensing ImageryabstractThe formation of shadows in very high spatial resolution (VHR) remote sensing imagery is attributed to light being blocked by objects, reducing spectral radiance in the shadow landscape. An accurate and robust shadow removal method can recover spectral and textural information and, hence, is a crucial preprocessing step for urban image analyses. In this study, we develop a KnOwledge-driven shadow progressive removal (KO-Shadow) framework with three subnets for VHR imagery using a weakly supervised manner. Specifically, the shadow preelimination subnet is proposed to initially address the large chromatic aberration between the real and shadow situations. Then, the prior knowledge-guided refinement subnet is proposed to refine the preelimination results by mining tone and texture information. Moreover, the locality feature discriminator is designed for region-specific evaluation of the generated shadow-free samples to improve the capacity of subnets. Experimental results of six typical cities in the world show that KO-Shadow is superior to the existing methods. Moreover, the generalizability analysis in complex urban scenarios validates the robustness of our method. The shadow recovery score (SRI) is proposed to evaluate the spectral similarities between the recovered area and shadow-related land-cover types (e.g., road, building, and lawn). The results show that KO-Shadow can yield more visually realistic shadow-free images and better quantitative performance. Overall, KO-Shadow provides a new perspective for VHR image shadow removal by mining the prior knowledge of the complex shadows in urban areas. Mingqiang Guo, Qiqi Zhu, Longli Ran, Jiancheng Luo |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Uksd-Net: an Unsupervised Knowledge-Guided Symmetric Deep Network for Forest Burned Areas DetectionabstractForest fires are one of the most widespread, frequent and harmful disasters. Accurate information over burned areas (BA) is essential for the assessment and restoration of economic and ecological losses. However, the feature representation of BA is often complicated. Meanwhile, since the spectral similarity among land cover categories, changes of background information have a greater impact. In addition, there are very few publicly available forest fire datasets. To solve the above problems, an unsupervised knowledge-guided symmetric deep network (UKSD-Net) is proposed for BA detection, where multi-index combination is utilized to incorporate the prior knowledge. The background changes are suppressed with the combination of symmetric deep network and Slow Feature Analysis (SFA). More reliable pseudo samples generated by uncertainty analysis are used to improve the final performance. Besides, we construct two forest fire datasets. The experimental results on the proposed dataset show that the framework is able to extract accurate and complete BA of multi-temporal remote sensing imagery. Qiqi Zhu, Yihui Shao, Qingfeng Guan 0001 |
IGARSS | 2 |
| 2022 | Cross-Domain Local Climate Zone Classification with Distance Metric Based on Domain Adaptation MethodabstractThe local climate zone (LCZ) classification system provides a standardized framework for presenting the morphological and functional characteristics of cities, especially for urban climate studies. The LCZ classification scheme is now widely used in urban heat island climate studies, urban planning, weather forecasting and other studies. Previous studies have generally been based on traditional machine learning and manual feature design for LCZ classification. Although deep learning methods perform well in LCZ classification, the generalization ability is still insufficient when facing large scale mapping, especially for cross-domain classification where the source and target domains have significant differences in feature distribution. In this paper, a subdomain adaptation feature alignment architecture (SAFA) is proposed to reduce inter-domain differences. In SAFA, convolutional neural network as a feature extractor in the framework, subdomain adaptive layers embedded in the network aligning features in the source and target domains. The proposed architecture is evaluated on a publicly available dataset. The experimental results show that the proposed architecture improves the overall classification performance of LCZ and NDVI helps to improve the accuracy of natural categories. Longli Ran, Qiqi Zhu, Qingfeng Guan 0001 |
IGARSS | 2 |
| 2022 | Subdomain Style Compensation Network for Remote Sensing Cross-Domain Scene ClassificationabstractHigh spatial resolution (HSR) imagery scene classification has been the subject of increased interest in recent years, and has great potential for many applications, such as urban planning and land cover classification. Deeping learning has been widely exploited in scene classification and achieved high classification accuracy. However, current the scene classification method based on deep learning depends on a large number of datasets for training, and the feature distribution of the default testing datasets is the same as that of the training datasets. In practical application, it is not only time-consuming and laborious to obtain a large number of training data, but also difficult to meet these needs due to the data shift between different domain. In addition, directly use of features extracted by the convolutional neural network (CNN) will lead to limited performance owing to the influence of domain migration. Therefore, how to effectively reduce the domain offset between the training data and the testing data, and discover discriminative information to scene classification is a challenging task. In this paper, the SSCN is proposed for HSR remote Sensing cross-domain scene classification. In SSCN, we introduce DSAN in which the local maximum mean discrepancy (LMMD) is first introduced for scene classification to capture the fine-grained information for each category and reduce the domain migration between training data and testing data. To extract more discriminative information from HSR images, a style normalization and restitution module (SNR) is developed, and an attention mechanism is added to the module to improve the performance of the model. The experimental results demonstrate that the proposed SSCN framework is superior to the state-of-the-art methods1and performs well for cross-domain scene classification. Qiqi Zhu, Qingfeng Guan 0001 |
IGARSS | 2 |
| 2022 | HFL-Net: Highlight Foreground and Local Scale Features Network for Cross-Domain Ship DetectionabstractShip detection receives increasing concerns as an essential ocean mission. Recently, deep learning (DL) has greatly improved ship detection accuracy from traditional methods. However, most of exist methods only utilize single source data, there are still some challenges that influence the ship detection performance. Therefore, we propose a highlight foreground and local scale features network (HFL-Net) for cross-domain ship detection to solve the problems of complex inshore background interferences and multi-scale ship feature differences. It is evaluated on two datasets with three other state-of-the-art methods and the experimental results confirm the superiority ofHFL-Net. Anqi Wu, Qiqi Zhu |
IGARSS | 2 |
| 2022 | A Spectral-Spatial-Dependent Global Learning Framework for Insufficient and Imbalanced Hyperspectral Image ClassificationabstractDeep learning techniques have been widely applied to hyperspectral image (HSI) classification and have achieved great success. However, the deep neural network model has a large parameter space and requires a large number of labeled data. Deep learning methods for HSI classification usually follow a patchwise learning framework. Recently, a fast patch-free global learning (FPGA) architecture was proposed for HSI classification according to global spatial context information. However, FPGA has difficulty in extracting the most discriminative features when the sample data are imbalanced. In this article, a spectral-spatial-dependent global learning (SSDGL) framework based on the global convolutional long short-term memory (GCL) and global joint attention mechanism (GJAM) is proposed for insufficient and imbalanced HSI classification. In SSDGL, the hierarchically balanced (H-B) sampling strategy and the weighted softmax loss are proposed to address the imbalanced sample problem. To effectively distinguish similar spectral characteristics of land cover types, the GCL module is introduced to extract the long short-term dependency of spectral features. To learn the most discriminative feature representations, the GJAM module is proposed to extract attention areas. The experimental results obtained with three public HSI datasets show that the SSDGL has powerful performance in insufficient and imbalanced sample problems and is superior to other state-of-the-art methods. Qiqi Zhu, Weihuan Deng, Zhuo Zheng, Yanfei Zhong, Qingfeng Guan 0001, Weihua Lin, Liangpei Zhang 0001, DeRen Li |
IEEE Trans. Cybern. | 1 |
| 2022 | A Weakly Pseudo-Supervised Decorrelated Subdomain Adaptation Framework for Cross-Domain Land-Use ClassificationabstractHigh spatial resolution (HSR) remote sensing image scene classification is a crucial way for land-use interpretation. However, most of the current scene classification methods assume that the training and test sets of remote sensing images follow the same feature distribution. In practical application, this assumption is difficult to guarantee. Domain adaptation (DA) is a machine learning paradigm that can effectively alleviate such problems. However, previous works mostly focused on aligning the global distribution of source domain (SD) and target domain (TD), which lose the inter-subdomain contextual relations between both domains, and ignore the redundancy among features. However, most DA methods usually only use the manually designed measurement criteria to establish the relationship between the SD and the TD, which is insufficient or complicated. In this paper, a weakly pseudo-supervised decorrelated subdomain adaptation (WPS-DSA) framework is proposed for HSR cross-domain land-use classification. In WPS-DSA, a feature extractor based on the subdomain adaptation network is used to extract the inter-subdomain characteristics of both domains. To weaken the influence of the features redundancy among remote sensing images, the switchable whitening module is introduced. In addition, a domain hierarchical sampling mechanism is designed to strengthen the connection between SD and TD in a simple way. Moreover, the WH-SH DA Dataset which is sampled from two typical Chinese cities is constructed to verify the generalization of the proposed framework. The experimental results of the cross-domain tasks on three publicly available HSR datasets and WH-SH DA Dataset display considerable performance and generalization ability of WPS-DSA. Qiqi Zhu, Yuwen Sun, Qingfeng Guan 0001, Lizhe Wang 0001, Weihua Lin |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | S3TRM: Spectral-Spatial Unmixing of Hyperspectral Imagery Based on Sparse Topic Relaxation-Clustering ModelabstractHyperspectral unmixing (HU) has been one of the hot spots in hyperspectral remote sensing research and has great potential in many applications. In recent years, the employment of the probabilistic topic model to mine latent topics in hyperspectral images has been an effective way for spectral unmixing. However, these methods fail to fully exploit the potential of topic models in uncovering image semantics and need extra sparsity constraints, which greatly increases the complexity of the model. In addition, the spatial information which can provide the correlation of features in adjacent pixels is usually ignored in topic model-based HU. To solve these problems, a spectral-spatial unmixing framework of hyperspectral imagery based on a sparse topic relaxation-clustering model ($\mathrm{S}^{3}{\mathrm {TRM}}$) is proposed. In$\mathrm{S}^{3}{\mathrm {TRM}}$, the topic model combined with implied sparse prior constraints are introduced, and the sparse characteristics of$\mathrm{S}^{3}{\mathrm {TRM}}$are used to capture the semantic representation of the spectrum. With the proposed relaxation-clustering strategy, multiple possible spectral representations of features can be obtained, which further alleviates the influence caused by endmember variability. Group clustering is used to locate the endmember quickly and accurately. Moreover, superpixel segmentation is considered to supplement the spatial distribution information of features, thereby improving the fractional abundance. Experiments on a synthetic dataset and three well-known real hyperspectral datasets confirm the excellent performance of the proposed framework in both qualitative assessment and quantitative evaluation, compared with the other state-of-the-art methods. Qiqi Zhu, Jiale Chen 0002, Wen Zeng 0003, Yanfei Zhong, Qingfeng Guan 0001, Zhijiang Yang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | CDANet: Contextual Detail-Aware Network for High-Spatial-Resolution Remote-Sensing Imagery Shadow DetectionabstractShadow detection automatically marks shadow pixels in high-spatial-resolution (HSR) imagery with specific categories based on meaningful colorific features. Accurate shadow mapping is crucial in interpreting images and recovering radiometric information. Recent studies have demonstrated the superiority of deep learning in very-high-resolution satellite imagery shadow detection. Previous methods usually overlap convolutional layers but cause the loss of spatial information. In addition, the scale and shape of shadows vary, and the small and irregular shadows are challenging to detect. In addition, the unbalanced distribution of the foreground and the background causes the common binary cross-entropy loss function to be biased, which seriously affects model training. A contextual detail-aware network (CDANet), a novel framework for extracting accurate and complete shadows, is proposed for shadow detection to remedy these issues. In CDANet, a double branch module is embedded in the encoder–decoder structure to effectively alleviate low-level local information loss during convolution. The contextual semantic fusion connection with the residual dilation module is proposed to provide multiscale contextual information of diverse shadows. A hybrid loss function is designed to retain the detailed information of the tiny shadows, which per-pixel calculates the distribution of shadows and improves the robustness of the model. The performance of the proposed method is validated on two distinct shadow detection datasets, and the proposed CDANet reveals higher portability and robustness than other methods. Qiqi Zhu, Xiongli Sun, Mingqiang Guo |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Oil Spill Contextual and Boundary-Supervised Detection Network Based on Marine SAR ImagesabstractOil spills have caused serious harm to the marine environment. Remote sensing technology is one of the important tools for marine environment monitoring. Synthetic aperture radar (SAR) has become an important technology for detecting marine pollution. Identifying dark spots is essential for oil spill detection based on SAR images. Dark spots’ detection can be achieved using image segmentation techniques. However, natural phenomena, such as waves and currents, can also cause dark spots, resulting in consistently uneven intensity, high noise, and blurred boundaries in oil spill images. In addition, existing oil spill detection models often perform well for large targets but have poor detection accuracy for small targets. To solve the above problems, the oil spill contextual and boundary-supervised detection network (CBD-Net) is proposed to extract refined oil spill regions by fusing multiscale features. To improve the internal consistency of oil spill regions, the spatial and channel squeeze excitation (scSE) block is introduced. In CBD-Net, boundary details are enhanced with optimized edge supervision. In addition, a manually labeled dataset is proposed, Deep-SAR Oil Spill (SOS) dataset, aiming to solve the problem of insufficient existing oil spill detection dataset. Experimental results demonstrate that CBD-Net outperforms other comparative models and is able to extract robust and accurate oil spill regions from complex SAR images. The highest mIoU of 83.42% and the highest F1 score of 87.87% were achieved on the SOS dataset. The CBD-Net model proposed in this article can play a guiding role in the marine oil spill decision support system. Qiqi Zhu, Xiaorui Yan, Qingfeng Guan 0001, Yanfei Zhong, Liangpei Zhang 0001, DeRen Li |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Proportion Estimation of Urban Mixed Scenes Based on Nonnegative Matrix Factorization for High Spatial Resolution Remote Sensing ImagesabstractUrban scenes play a pivotal role in urban planning and environmental protection. Thus, depicting urban scenes is an increasingly important research area in remote sensing applications. Researchers have not treated quantitative measurements on scenes in much detail. So far, there has been little discussion about scene decomposition methods for high resolution imagery. This article proposed a framework based on nonnegative matrix factorization of mixed scene (NMFMS) to estimate the mixing ratio of urban scenes. The framework consists of two steps, namely deep feature representation and scene decomposition. The urban mixed scenes can be described by the proportions of each sub-scenes. To confirm the feasibility of the proposed method, experiments with high resolution imagery of Hanyang District in Wuhan and a mixed scene data set are applied. We further utilized root-mean-square error (RMSE) and Vector angular distance (VAD) to quantitatively evaluate the effect of scene decomposition. The experimental results confirmed the high precision of the proposed algorithm. Jiale Chen 0002, Qiqi Zhu, Xiongli Sun, Qingfeng Guan 0001 |
IGARSS | 2 |
| 2021 | EML-GAN: Generative Adversarial Network-Based End-to-End Multi-Task Learning Architecture for Super-Resolution Reconstruction and Scene Classification of Low-Resolution Remote Sensing ImageryabstractHigh spatial resolution remote sensing images (HSR-RSIs) are critical to providing fine land cover/land use information for scene classification. The global low spatial resolution remote sensing images (LSR-RSIs) can be easily obtained at present, whereas it is still a challenge to acquire large-scale HSR-RSIs. In this paper, an algorithmic-based architecture is proposed to improve the spatial resolution of RSIs beyond the limits of imaging sensors. The generative adversarial network-based end-to-end multi-task learning architecture (EML-GAN) is proposed for LSR-RSIs super-resolution reconstruction and scene classification simultaneously. In EML-GAN, the generator network is used to recover the fine geometric structures of LSR-RSIs by fusing the deep contextual, structure, and edge information. In addition, the discriminator network is designed to predict the scene label and distinguish the real/fake of the input data. The proposed architecture is evaluated on a public dataset and two self-made dataset. The experimental results show that the proposed architecture improves the visual effect and classification performance of LSR-RSIs. Weihuan Deng, Qiqi Zhu, Xiongli Sun, Weihua Lin, Qingfeng Guan 0001 |
IGARSS | 2 |
| 2021 | A Siamese Global Learning Framework for Multi-Class Change DetectionabstractChange detection is one of the main tasks in remote sensing field, and is essential for the accurate processing and understanding of the available earth observation data. However, most of the works focus on traditional binary change detection, without considering the change classes information. To make use of the semantic information and analyze change classes comprehensively, we proposed a Siamese Global Learning Framework (Siam-GL) for multi-class change detection. In Siam-GL, the global hierarchically (G-H) sampling strategy is designed to address the imbalanced training sample problem. In addition, the change mask was created and designed to distinguish the changed object for bi-temporal remote sensing images simultaneously. The proposed Siam-GL framework was validated on the Guangzhou dataset of two high spatial resolution remote sensing images in 2013 and 2015, and achieved a better performance than the other state-of-the art methods. Qiqi Zhu, Weihuan Deng, Qingfeng Guan 0001 |
IGARSS | 2 |
| 2021 | A Union Framework with Sparse Topic Relaxion and Group Clustering for Hyperspectral UnmixingabstractHyperspectral unmixing has been one of the hot spots and has great potential in a wide range of hyperspectral applications. In recent years, the employment of the probabilistic topic model to acquire latent topics of hyperspectral image has been an effective way for spectral unmixing. However, these methods fail to fully exploit the potential of topic models in uncovering image semantics and need extra sparsity constraints, which greatly increases the complexity of the model. To solve these problems, a union framework with sparse topic relaxion and group clustering (USTRGC) is proposed. In USTRGC, the implied sparse prior constraints by the sparse topic model are introduced. Through the relaxation of model, the influence caused by the endmember variability on the unmixing accuracy can be alleviated. Furthermore, unmixing models of different category are united to alleviate the ill-posed nature of the model, thereby improving the fractional abundance. Comprehensive evaluations on two real hyperspectral images, and comparisons with some popular HU methods available in existing research verifies the effectiveness and superiority of the proposed methods. Qiqi Zhu, Wen Zeng 0003, Qingfeng Guan 0001 |
IGARSS | 2 |
| 2021 | CADNet: Top-Down Contextual Saliency Detection Network for High Spatial Resolution Remote Sensing Image Shadow DetectionabstractIn order to improve the feature richness of remote sensing images and meet the needs of remote sensing image interpretation, shadow detection has become a hotspot in high-resolution remote sensing (HSR) research. Traditional threshold-based and machine learning-based methods do show their effectiveness, but they may not take the inherent details of the shadow into consideration, and it is difficult to cope with the saliency and intricate distribution pattern of the shadow in HSR images, which results in the lack of robustness in extracting sufficient global and local shadow contexts. To solve those problems, a top-down contextual saliency detection network (CADNet) is proposed. Compared with the traditional shadow detection network, more contextual information can be retained by the double-branch strategy of the encoder and residual dilation upsampling of the decoder in CADNet. The low-level and high-level semantic information can be combined to accurately predict the salient regions through the proposed short connection. The proposed CADNet is evaluated on a public shadow detection dataset, and the experimental results demonstrate the effectiveness of the proposed CADNet. Mingqiang Guo, Qiqi Zhu |
IGARSS | 3 |
| 2021 | UCWater: Unsupervised Content-Adaptive Water-Body Extraction Framework for High-Resolution Satellite ImageryabstractDeep learning techniques have provided significant improvements in water-body extraction in high-resolution satellite imagery (HRSI). The current deep learning based HRSI water-body extraction usually requires large labeling of training samples, which follows a supervised learning framework. In this paper, an unsupervised content-adaptive water-body extraction framework (UCWater) is proposed for HRSI. A content-consistent matching (CCM) selecting samples strategy is first introduced for water-body extraction, by selecting positive source image and their similar pixels. This strategy can alleviate the domain bias caused by those content-irrelevant images in the UCWater framework. A maximum squares (MS) loss is designed to prevent the training process being dominated by easy-to-transfer samples in the target domain, balancing the gradient of well-classified target samples. Additionally, the image-wise weighting ratio is also designed to alleviate the class imbalance in the unlabeled target domain based on the MS loss. Experiments are conducted on a representative water-body migration task, i.e., GID → DeepGlobe, which demonstrate the effectiveness of our proposed approach. Qiqi Zhu, Jianjun Lv, Qingfeng Guan 0001 |
IGARSS | 2 |
| 2021 | A Survey of Hyperspectral Image Super-Resolution TechnologyabstractHyperspectral images (HSIs) have very high spectral resolution, which can reflect the characteristics of different materials well. However, compared with RGB image or multispectral image (MSI), the spatial resolution of HSI is much lower, which limits its applications. Therefore, many super-resolution (SR) techniques have been proposed to reconstruct HSI with high spatial resolution image. To the best of our knowledge, there has not, to date, that been a study aimed at expatiating and summarizing the current research situation. Therefore, this is our motivation in this survey. In view of the promising development prospects in this field, this paper systematically reviews the existing SR methods of HSI. Specifically, two major categories are summarized, one is fusion-based methods, and the other is single HSI SR methods. At the end of the paper, several future development directions for HSI SR are given. Meilin Zhang, Xiongli Sun, Qiqi Zhu, Guizhou Zheng |
IGARSS | 3 |
| 2021 | Oil Spill Detection Based on CBD-Net Using Marine SAR ImageabstractOil spill is one of the most widespread, frequent and harmful marine pollution. Oil spill detection based on synthetic aperture radar (SAR) images detects the oil film by identifying dark spots in the images. Dark spots detection can be achieved using image segmentation techniques. However, natural phenomena such as waves and currents can also cause dark spots, resulting in consistently uneven intensity, high noise and blurred boundaries in oil spill images. In addition, existing oil spill detection models often perform well for large targets, but have poor detection accuracy for small targets, causing part of the oil spill to be ignored. To solve the above problems, oil contextual and boundary-supervised detection network (CBD-Net) is proposed to extract refined oil spill regions by fusing multiscale features, where the scSE attention block is used to model the global context to improve the internal consistency of oil spill regions. In CBD-Net, boundary details are enhanced with optimized edge supervision. In addition, a manually labeled dataset is proposed, Deep-SAR Oil Spill (SOS) dataset, aiming to solve the problem of insufficient existing oil spill detection dataset. Experimental results demonstrate that the proposed model outperforms other comparative models, and is able to extract robust and accurate oil spill regions from complex SAR images. Qiqi Zhu, Qingfeng Guan 0001 |
IGARSS | 2 |
| 2020 | Urban Scenes Change Detection Based on Multi-Scale Irregular Bag of Visual Features for High Spatial Resolution ImageryabstractRemote sensing scene change detection (SCD) is to detect whether and what changes have occurred in the semantic category of corresponding scenes for a long time at the semantic level. This can provide detailed land use/land cover change information for Urban planning and environmental monitoring. Previous studies take regular patches divided by uniform grid sampling as scene units. This may lead to mosaic phenomenon, and use fixed Window to extract features, ignoring the multi-scale features of ground objects, while extracting scene features. To solve the problems, the multi-scale irregular bag of visual features (MIBVF) framework is proposed for high spatial resolution (HSR) imagery SCD. In this paper, we integrate image classification of the physical characteristics from remote sensing data with the socio-economic attributes from open source geographic data. Road network data is used to preserve the geological significance and semantic integrity of urban scenes, and multi-scale window sampling is used to solve the problem of different object sizes. To confirm the feasibility of the proposed method, experiments with multi-temporal images of the Pudong area in Shanghai indicate that the proposed method achieves a clearly higher change detection accuracy than current state-of-the-art methods. Jiale Chen 0002, Qiqi Zhu, Yanfei Zhong, Qingfeng Guan 0001, Liangpei Zhang 0001, DeRen Li |
IGARSS | 2 |
| 2020 | Remote Sensing Scene Classification Using Spatial Transformer Fusion NetworkabstractRemote sensing image scene classification is one of the hottest topics in high spatial resolution remote sensing image understanding.The complexity of spatial distribution and structure patterns of objects in high resolution remote sensing images make the problem challenging.Deep learning methods represented by convolutional neural networks have powerful feature learning capabilities and perform well in scene classification tasks of remote sensing images.Inspired by spatial transformer network (STN), this paper proposes an effective remote sensing image scene classification network-the Spatial Transformer Fusion Network (STFN), which applies the spatial transformer network to remote sensing image scene classification task.The structure of STFN uses spatial transformer network to crop remote sensing images to extract the attention areas.Then, STFN extracts features of original images and the cropped images and finally fuse them.Experiments and evaluations are performed on two public remote sensing image datasets: the UC Merced Land-Use dataset with 21 scene categories and the NWPU-RESISC45 dataset with 45 scene categories.Results of experiments show that the proposed method has a relatively simple network structure and can produce competitive classification performance. Shun Tong, Kunlun Qi, Qingfeng Guan 0001, Qiqi Zhu, Chao Yang 0007, Jie Zheng 0007 |
IGARSS | 4 |
| 2020 | Semi-Automatic Fully Sparse Semantic Modeling Framework for Hyperspectral UnmixingabstractIn order to improve the accuracy of surface classification and meet the needs of sub-pixel-level target detection, spectral unmixing has been one of the hot spots in hyperspectral remote sensing research. The employment of the probabilistic topic model to acquire latent topics of hyperspectral image has been an effective way for spectral unmixing. However, this approach fails to consider the sparsity of the semantic representation and high computational complexity. In addition, the number of endmembers cannot be determined automatically. To solve the problem, in this paper, the novel spectral unmixing method based on semi-automatic fully sparse semantic modeling framework (SFSSM) is proposed. In SFSSM, modestly few arithmetic operations are required to identify the pure spectral signatures (endmembers) and the fractional abundances of the endmembers. Meanwhile, the sparsity and representativeness of the topics generated by SFSSM guarantee that the endmembers can be obtained automatically in low time consumption. The experimental results obtained with two real image confirm that the proposed method significantly improves the performance when compared with the other methods. Qiqi Zhu, Wen Zeng 0003, Yanfei Zhong, Qingfeng Guan 0001, Liangpei Zhang 0001, DeRen Li |
IGARSS | 2 |
| 2020 | A Modified D-Linknet with Transfer Learning for Road Extraction from High-Resolution Remote SensingabstractRoad extraction, which aims to label remote images with a specific road detection, is a fundamental task for understanding remote sensing imagery. Deep learning has strong characteristic learning ability, for example, the state-of-the-art D-Linknet is an effective way to capture the road information. However, existing regularization methods either do not match the performance for large batches, or still exhibit degradation in performance for smaller batches. Besides, the roads in different areas have various characteristics and lack a good transfer. To remedy these issues, we proposed a novel road extraction network which integrated the filter response normalization (FRN) layer with D-Linknet (FND-Linknet). The FRN layer is effective and robust for road extraction task, and can eliminate the dependency on other batch samples. In addition, the multisource road dataset is collected and annotated to improve features transfer. Experimental results on three datasets verify that the proposed FND-Linknet framework outperforms the state-of-the-art methods both in accuracy and connectivity. Qiqi Zhu, Yanfei Zhong, Qingfeng Guan 0001, Liangpei Zhang 0001, DeRen Li |
IGARSS | 2 |
| 2020 | Super Resolution Generative Adversarial Network Based Image Augmentation for Scene Classification of Remote Sensing ImagesabstractHigh spatial resolution remote sensing image (RSI) scene classification, aimed at automatically labelling images with the given semantic categories, has been a hot issue. As it's difficult for RSI to quickly obtain a large number of training samples from a specific area. Traditional scene classification researches were mainly using deep learning models to transfer natural images to RSI. Considering the differences between natural images and RSI, we trained several Super Resolution GAN models by using different resolution RSI data from Google earth image. This paper proposed a novel SRGAN-CNN framework. Through transferring the data with scene classification dataset to obtain high resolution fake RSI. The experimental results demonstrate that the proposed framework can enhance transfer effect and help improve the accuracy of scene classification using low resolution RSI. Qiqi Zhu, Yanfei Zhong, Qingfeng Guan 0001, Liangpei Zhang 0001, DeRen Li |
IGARSS | 1 |
| 2020 | Topic Model for Remote Sensing Data: A Comprehensive ReviewabstractFrom text analysis to image interpretation, the topic model (TM) always plays an important role. With its powerful semantic mining capabilities, it is able to capture the latent spectral and spatial information from remote sensing (RS) images. Recent years have witnessed widespread use of TM to solve the problems in RS image interpretation, i.e., semantic segmentation, target detection, and scene classification. However, there has not yet been a study expatiating and summarizing the current situation of RS applications with TM. This paper intends to systematically summarize the application of TM in RS images and to conduct several typical experiments for comparison. Specifically, the architecture of our work can be explained as follows: 1) the theory of TM; 2) the applications of RS based on TM; 3) experimental analysis of typical TM methods to provide reference for further understanding, and 4) summary and prospects for guiding further research into TM for RS data. Qiqi Zhu, Jiangqin Wan, Yanfei Zhong, Qingfeng Guan 0001, Liangpei Zhang 0001, DeRen Li |
IGARSS | 1 |
| 2019 | Sub-Pixel Mapping with Multiple Shifted Hyperspectral Images Based on Multiobjective Evolutionary AlgorithmabstractSub-pixel mapping (SPM) can interpret the sub-pixel spatial distribution of land-cover classes in hyperspectral image, which is an ill-posed problem due to the inadequate information of a single image. Auxiliary information provided by multiple shifted (MS) images can make SPM problem well-posed and improve mapping accuracy. The maximum a posteriori (MAP) technique can incorporate the auxiliary information of MS images, but it introduces a fixed weight parameter to fuse the auxiliary information and spatial prior information, heavily influencing the mapping result. This paper proposed a novel SPM method to model the auxiliary information and spatial prior information into two objective functions, which can be simultaneously optimized by the devised multiobjective evolutionary algorithm. Therefore, there is no need of weight parameter, and the two objective functions can be intelligently integrated during the evolution. Experimental results and parameter analysis have indicated the superiority of the proposed method. Mi Song, Yanfei Zhong, Ailong Ma, Qiqi Zhu, Liqin Cao, Liangpei Zhang 0001 |
IGARSS | 4 |
| 2019 | High-Resolution Remote Sensing Image Scene Understanding: A ReviewabstractHigh-resolution remote sensing (HRS) image analysis is a fundamental but challenging problem. To bridge the semantic gap, scene understanding has been proposed to achieve higher-level interpretation, through classifying the HRS scene through spatial relationship cognition and semantic induction between the land-cover objects. As a new research field, however, there has not yet been a study expatiating and summarizing the current situation of scene understanding. This paper first defines the concept of scene understanding for HRS imagery, which is different from natural image scene classification. The theory of scene understanding for HRS imagery is investigated, and is classified into four main categories: 1) scene classification based on semantic objects; 2) scene classification based on mid-level features; 3) scene classification based on deep learning; and 4) scene understanding applications based on geographic data mining. Qiqi Zhu, Xiongli Sun, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 1 |
| 2018 | Scene Classification Based on the Sparse Homogeneous-Heterogeneous Topic Feature ModelabstractHigh spatial resolution (HSR) imagery scene classification has been the subject of increased interest in recent years, and has great potential for many applications, such as urban functional analysis. Rooted in natural information processing, the use of the probabilistic topic model (PTM) to capture latent topics to represent HSR images has been an effective way to bridge the semantic gap. However, how to effectively discover discriminative information to recognize the HSR scenes is a challenging task. In this paper, the sparse homogeneous-heterogeneous topic feature model (SHHTFM) is proposed for HSR image scene classification. Differing from the conventional PTM-based scene classification methods, which utilize only heterogeneous features, SHHTFM explores the effect of the homogeneous information. Based on the union of uniform grid sampling and simple linear iterative clustering superpixel sampling, SHHTFM exploits both the heterogeneous and homogeneous information. After separately mining different types of low-level features and latent topics, the sparse topic inference procedure of SHHTFM further improves the fusion of the sparse heterogeneous and homogeneous topics. In addition, multisource geographical data are effectively integrated, where the water and vegetation boundaries define a more accurate way to restrict the boundaries of different scenes, and are then combined with the road network data to further improve the scene annotation performance. This provides more reliable and applicable results for us to better understand the complex scenes. The experimental results obtained with two HSR image classification data sets and an HSR image annotation data set demonstrate that the proposed SHHTFM framework can solve the scene classification problem, with a high classification accuracy as well as a high time efficiency. Qiqi Zhu, Yanfei Zhong, Liangpei Zhang 0001, DeRen Li |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Adaptive Deep Sparse Semantic Modeling Framework for High Spatial Resolution Image Scene ClassificationabstractHigh spatial resolution (HSR) imagery scene classification, which involves labeling an HSR image with a specific semantic class according to the geographical properties, has received increased attention, and many algorithms have been proposed for this task. The employment of the probabilistic topic model to acquire latent topics and the convolutional neural networks (CNNs) to capture deep features for representing HSR images has been an effective ways to bridge the semantic gap. However, the midlevel topic features are usually local and significant, whereas the high-level deep features convey more global and detailed information. In this paper, to discover more discriminative semantics for HSR images, the adaptive deep sparse semantic modeling (ADSSM) framework combining sparse topics and deep features is proposed for HSR image scene classification. In ADSSM, the fully sparse topic model and a CNN are integrated. To exploit the multilevel semantics for HSR scenes, the sparse topic features and deep features are effectively fused at the semantic level. Based on the difference between the sparse topic features and the deep features, an adaptive feature normalization strategy is proposed to improve the fusion of the different features. The experimental results obtained with four HSR image classification data sets confirm that the proposed method significantly improves the performance when compared with the other state-of-the-art methods. Qiqi Zhu, Yanfei Zhong, Liangpei Zhang 0001, DeRen Li |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | Scene Classification Based on the Fully Sparse Semantic Topic ModelabstractIn high spatial resolution (HSR) imagery scene classification, it is a challenging task to recognize the high-level semantics from a large volume of complex HSR images. The probabilistic topic model (PTM), which focuses on modeling topics, has been proposed to bridge the so-called semantic gap. Conventional PTMs usually model the images with a dense semantic representation and, in general, one topic space is generated for all the different features. However, this approach fails to consider the sparsity of the semantic representation, the classification quality, as well as the time consumption. In this paper, to solve the above problems, a fully sparse semantic topic model (FSSTM) framework is proposed for HSR imagery scene classification. FSSTM, with an elaborately designed modeling procedure, is able to represent the image with sparse but representative semantics. Based on this framework, the topic weights of multiple features are exploited by solving a concave maximization problem, which improves the fusion of the discriminative semantic information at the topic level. Meanwhile, the sparsity and representativeness of the topics generated by FSSTM guarantee that the image is adaptive to the change of a topic number. FSSTM can consistently achieve a good performance with a limited number of training samples, and is robust for HSR image scene classification. The experimental results obtained with three different types of HSR image data sets confirm that the proposed algorithm is effective in improving the performance of scene classification, and is highly efficient in discovering the semantics of HSR images when compared with the state-of-the-art PTM methods. Qiqi Zhu, Yanfei Zhong, Liangpei Zhang 0001, DeRen Li |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Bag-of-Visual-Words Scene Classifier With Local and Global Features for High Spatial Resolution Remote Sensing ImageryabstractScene classification has been studied to allow us to semantically interpret high spatial resolution (HSR) remote sensing imagery. The bag-of-visual-words (BOVW) model is an effective method for HSR image scene classification. However, the traditional BOVW model only captures the local patterns of images by utilizing local features. In this letter, a local-global feature bag-of-visual-words scene classifier (LGFBOVW) is proposed for HSR imagery. In LGFBOVW, the shape-based invariant texture index is designed as the global texture feature, the mean and standard deviation values are employed as the local spectral feature, and the dense scale-invariant feature transform (SIFT) feature is employed as the structural feature. The LGFBOVW can effectively combine the local and global features by an appropriate feature fusion strategy at histogram level. Experimental results on UC Merced and Google data sets of SIRI-WHU demonstrate that the proposed method outperforms the state-of-the-art scene classification methods for HSR imagery. Qiqi Zhu, Yanfei Zhong, Gui-Song Xia, Liangpei Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | Scene Classification Based on the Multifeature Fusion Probabilistic Topic Model for High Spatial Resolution Remote Sensing ImageryabstractScene classification has been proved to be an effective method for high spatial resolution (HSR) remote sensing image semantic interpretation. The probabilistic topic model (PTM) has been successfully applied to natural scenes by utilizing a single feature (e.g., the spectral feature); however, it is inadequate for HSR images due to the complex structure of the land-cover classes. Although several studies have investigated techniques that combine multiple features, the different features are usually quantized after simple concatenation (CAT-PTM). Unfortunately, due to the inadequate fusion capacity of k-means clustering, the words of the visual dictionary obtained by CAT-PTM are highly correlated. In this paper, a semantic allocation level (SAL) multifeature fusion strategy based on PTM, namely, SAL-PTM (SAL-pLSA and SAL-LDA) for HSR imagery is proposed. In SAL-PTM: 1) the complementary spectral, texture, and scale-invariant-featuretransform features are effectively combined; 2) the three features are extracted and quantized separately by k-means clustering, which can provide appropriate low-level feature descriptions for the semantic representations; and 3)the latent semantic allocations of the three features are captured separately by PTM, which follows the core idea of PTM-based scene classification. The probabilistic latent semantic analysis (pLSA) and latent Dirichlet allocation (LDA) models were compared to test the effect of different PTMs for HSR imagery. A U.S. Geological Survey data set and the UC Merced data set were utilized to evaluate SAL-PTM in comparison with the conventional methods. The experimental results confirmed that SAL-PTM is superior to the single-feature methods and CAT-PTM in the scene classification of HSR imagery. Yanfei Zhong, Qiqi Zhu, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2014 | Multi-feature probability topic scene classifier for high spatial resolution remote sensing imageryabstractScene classification can obtain the high-level semantic information in high spatial resolution (HSR) imagery. Probability topic model as a typical scene semantic representation has been successfully applied to nature scene by utilizing a single feature. However, it is not completely fit for HSR images due to the complexity of land cover classes. To solve the problem, multi-feature probability topic scene classifier based on Latent Dirichlet allocation (LDA), namely MFPTSC, is proposed for HSR imagery. In MFPTSC, the spectral, texture, and SIFT features as three representative features are firstly integrated. If the traditional multi-features fusion method (VIS-LDA) is used, which each feature vector is usually stacked at the visual word level, abundant information is lost, which leads to an undesirable classification performance. In this paper, a novel feature fusion strategy at the semantic allocation level, named SAL-LDA, is proposed to avoid information loss to a large extent by mining the latent semantics in accordance with the distinctive characteristics of each feature. Experiment results using the image dataset of 21 land-use classes demonstrate that the multi-feature fusion strategies of VIS-LDA and SAL-LDA both improve the classification accuracy, but the proposed SAL-LDA strategy is better than VIS-LDA. Qiqi Zhu, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 1 |
| 2012 | Performance analysis of collaborative design networkabstractIn a collaborative design process, design activities, people and design tools are connected each other for collaboratively finishing design tasks. A method based on complex network theory is proposed in this paper to analyze the performance of the process. Firstly, the collaborative design process is abstracted as a network named collaborative design network (CDN), and its elements of CDN, such as nodes, edges and weight of edges are defined. Then the topological and physical characteristics of the CDN are defined and analyzed for revealing its performances. Finally, a steam turbine rotor design process is studied as an example to illustrate the feasibility and availability of the proposed method. Leijie Fu, Pingyu Jiang, Qiqi Zhu |
CSCWD | 3 |
| 2008 | An architecture to implement dynamic online manual for working capability services of CNC machine toolsabstractIn order to change the serial service in conventional manuals, dynamic online manuals (DOA) are designed. DOMs meet the requirements of sustainable development and benefit users and producers, which provide just-in-time and high-quality services to users, and offer new design ideas to producers. The DOM architecture for CNC mill center is then proposed after services content is presented, and key technologies of the architecture are outlined. Finally, it is concluded that the DOM has some support for future research on product service. Qiqi Zhu, Pingyu Jiang |
CSCWD | 1 |