VLDB 2026 Research / reviewers in the wild / expert
Libao Zhang
dblp:133/8949
· DBLP profile ↗
117ranked-venue papers
45as first author
69since 2021 · last 2027
0000-0002-0888-2330ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 55 · 20 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 54 · 21 first-author · 31 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Prototype region calibration guided federated domain generalization
Wenjie Yao, Suxia Zhu, Libao Zhang, Guanglu Sun, Xinzhong Zhu |
Inf. Process. Manag. | 3 |
| 2026 | DFGLIoT: A Dual-Fusion Graph Learning Framework for Cross-Institutional Mobile IoT Device IdentificationabstractAccurately identifying IoT devices connected to a network is crucial for improving network management and ensuring network security. As IoT devices increasingly move across institutional boundaries, identifying devices that migrate among institutions has become a significant challenge. However, existing studies assume that IoT devices remain stationary. As a result, they are inadequate for addressing the security risks introduced by device mobility across institutions. Therefore, we propose an IoT device identification framework named DFGLIoT. The framework adopts a decentralized fully connected architecture to enable knowledge sharing among institutions. This design avoids single points of failure while effectively supporting the identification of cross-institutional mobile IoT devices. We model the communication traffic between IoT devices and their gateways as communication interaction graphs. These graphs provide a comprehensive view of the interaction process between the communicating parties. Based on this representation, we design a graph classifier that integrates a dual-scale dependency modeling module with a spatial feature extraction module. The classifier captures interaction patterns in the communication traffic and constructs behavior fingerprints of IoT devices, enabling accurate device identification. Experimental results on three public datasets demonstrate the effectiveness of DFGLIoT in identifying IoT devices that move across institutions. The source code can be accessed at https://github.com/traveler-wang/DFGLIoT. Guanglu Sun, Libao Zhang, Wenjie Yao |
IEEE Internet Things J. | 4 |
| 2026 | Federated Chain Context Optimization for Long-Tailed Multi-Label Image ClassificationabstractFederated learning is an emerging machine learning paradigm that effectively alleviates the data silo problem by distributing the model training process to multiple data holders. However, data from real-world mobile applications often has multi-label and presents a long-tailed distribution, where labels are generally non-independent and non-identically distributed, thereby increasing the challenges caused by data heterogeneity. To address the above problems, we propose a Federated Chain Context Optimization (FedCCO) for long-tailed multi-label image classification. Inspired by the success of Chain of Though (CoT) in enhancing the semantic expressive ability of models, this method fine-tunes the CLIP model using semantic descriptive vectors generated by the Chain Context Optimization (ChCoOp) to establish semantic correlations between head and tail classes across clients, which improves the ability of the model to recognize tail classes. The experimental results show that the FedCCO achieves satisfactory performance in long-tailed multi-label image classification in federated learning on VOC-LT and COCO-LT datasets. Libao Zhang, Suxia Zhu, Wenjie Yao, Guanglu Sun |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | SLVR: Super-Light Visual Reconstruction via Blueprint Controllable Convolutions and Exploring Feature Diversity RepresentationabstractRecently, improving the residual structure and designing efficient convolutions have become important branches of lightweight visual reconstruction model design. We have observed that the feature addition mode (FAM) in existing residual structure tends to lead to slow feature learning or stagnation in feature evolution, a phenomenon we define as network inertia. In addition, although blueprint separable convolutions (BSConv) have proved the dominance of intra-kernel correlation, BSConv forces the blueprint to perform scale transformation on all channels, which may lead to incorrect intra-kernel correlation and introduce useless or disruptive features on some channels and hinder the effective propagation of features. Therefore, in this paper, we rethink the FAM and BSConv for super-light visual reconstruction framework design. First, we design a novel linking mode, called feature diversity evolution link (FDEL), which aims to alleviate the phenomenon of network inertia by reducing the retention of previous low-level features, thereby promoting the evolution of feature diversity. Second, we propose blueprint controllable convolutions (B2Conv). The B2Conv can adaptively pick accurate intra-kernel correlation in the depth-axis, effectively preventing the introduction of useless or disruptive features. Based on FDEL and B2Conv, we develop a super-light super-resolution (SR) framework SLVR for visual reconstruction. Both FDEL and B2Conv can serve as efficient plugins. Extensive experimental results demonstrate the effectiveness of our proposed B2Conv, FDEL, and SLVR. Our code will be available at https://github.com/chongningni/SLVR. Ning Ni 0002, Libao Zhang |
CVPR | 2 |
| 2025 | Hazy Low-Quality Satellite Video Restoration Via Learning Optimal Joint Degradation Patterns and Continuous-Scale Super-Resolution Reconstruction
Ning Ni 0002, Libao Zhang |
CVPR | 2 |
| 2025 | Joint Semantic Segmentation of Optical and SAR Image in Hazy Environments via Cross-modal Information Rectification and Cross-attention FusionabstractSemantic segmentation is crucial in remote sensing image processing. In recent years, semantic segmentation using optical and SAR images for multi-modal fusion is gaining attention for its good results. The current research primarily encompasses two problems: 1) Existing fusion methods are designed for clear images and struggle in harsh weather. 2) Current fusion methods insufficiently capture multi-modal information correlation. This paper presents a joint semantic segmentation of optical and SAR in hazy environments network that incorporates channel fusion for feature enhancement and cross-attention for feature fusion, enabling efficient segmentation of hazy optical images. First, we design a channel-guided crossmodal information correction module. This module regards fog as noise and uses high-confidence features of one modality to calibrate the other modality, thereby reducing the impact of fog. Secondly, to address the issue of insufficient fusion of multimodal information, we introduce a cross-attention fusion module to effectively combine the complementary information from two streams leveraging the large receptive field enabled by self-attention. The implementation results show that this method has better results for the semantic segmentation of blurry remote sensing images. Xinyue Fan, Libao Zhang |
ICASSP | 2 |
| 2025 | Hazy Remote Sensing Image Semantic Segmentation with Weak Annotations via Pre-training Optimization and Co-trainingabstractIn recent years, weakly supervised semantic segmentation has emerged as a prominent research topic in the field of remote sensing image semantic segmentation due to its cost-effective labeling advantages. However, the presence of haze in remote sensing images poses significant challenges to accurate semantic segmentation. Despite the numerous haze removal methods developed for remote sensing images, their efficacy in the subsequent task of semantic segmentation remains inadequate. To address these issues, this paper aims to enhance the robustness of the segmentation network against haze interference by proposing a weakly supervised semantic segmentation framework based on pre-training optimization and dual-network co-training. The proposed approach employs a pseudo-label optimization network to effectively filter out noise interference and subsequently trains two parallel segmentation networks that mutually guide each other for enhanced robustness. Additionally, an edge optimization loss is introduced to improve prediction accuracy by incorporating both edge and texture information. Experimental comparisons with alternative methods across multiple datasets validate the superior performance of our proposed method. Junda Xu, Libao Zhang |
ICASSP | 2 |
| 2025 | Clouds and Haze Co-Removal Based on Saliency-Guided Multi-Scale Diffusion Model for Remote Sensing ImagesabstractClouds and haze co-removal from remote sensing images is an important task. However, current algorithms often struggle with complex distributions and uneven illumination. To solve the issues, we propose a multi-scale diffusion model guided by global perceptual saliency. Firstly, the global perceptual saliency block is employed to enhance the model’s ability to perceive clouds and haze distributions. Then, an edge feature extraction block is utilized to generate a texture saliency map, guiding the model to focus on texture information outside the clouds and haze, thereby improving the image structure generation capability. The texture saliency map is then projected into a multi-scale guiding network, where gray-scale conversion and down-sampling are applied to suppress information unrelated to image structure, enhancing the model’s robustness to variations in the input domain. Finally, a loss function based on regressor-constrained output is designed to optimize the model’s performance. Extensive experiments demonstrate that the proposed method outperforms state-of-the-art methods on synthetic and real-world images. Libao Zhang |
ICME | 2 |
| 2025 | Contrastive Adversarial Learning for Region-Aware Weakly Annotated Object Segmentation in Hazy Remote Sensing ImagesabstractIt is essential to generate high quality class activation maps (CAMs) for accurate object segmentation of weakly annotated remote sensing images (RSIs). However, RSIs are highly susceptible to haze interference during capturing, adversely affecting the precision of object location. In addition, the diverse shapes and indistinguishable boundaries of targets in RSIs hinder the generation of ideal CAMs. To address these problems, we propose a region-aware weakly annotated object segmentation (RA-WAOS) model for hazy RSIs based on contrastive adversarial learning, where the haze condition of an image is regarded as its inherent style. The haze style projector (HSP) optimized by contrastive learning is specially designed to obtain haze style embeddings, and the adversarial training of HSP and RA-WAOS is adopted to narrow the gap between different styles in the latent space. Experimental results reveal that our proposal performs superiorly on hazy RSIs compared with other competing methods. Wanning Zhu, Libao Zhang |
ICME | 2 |
| 2025 | CLIP-HNet: Hybrid Network with Cross-Modal Guidance for Self-Supervised Remote Sensing DehazingabstractUnsupervised remote sensing dehazing remains a challenging and ill-posed task due to the absence of reliable supervision signals. Existing dehazing methods with unpaired data often oversimplify haze removal as style transfer, limiting generalization in complex scenarios. Moreover, current unimodal frameworks neglect cross-modal cues that could improve contextual reasoning. To address these issues, we propose a novel cross-modal guided self-supervised dehazing framework called CLIP-HNet, which achieves multi-model feature extraction, boundary-focused reconstruction and adaptive sample filtering. Specifically, to capture global-local contextual features, a hybrid feature interaction network is designed, which bridges the feature representations of multi models with global context-aware module (GCAM) and hybrid feature fusion module (HF2 M). Then, based on the hybrid features, a boundary-aware feature reconstruction (BFRec) is proposed to further refine edge details. Furthermore, a CLIP-guided progressive information distillation scheme is presented to dynamically prioritize training samples and distill useful signals, which predicts haze concentration by CLIP and progressively increases sample difficulty during the training stage. Finally, a frequency-domain texture matching (FTM) strategy refines texture and spectral details, enhancing the model's ability to recover fine details. Experiments on synthetic and real RSIs demonstrate that the proposed CLIP-HNet surpasses state-of-the-art approaches, achieving superior visual quality and quantitative performance. Shan Wang 0009, Weisi Lin, Yun Liu 0002, Libao Zhang |
ACM Multimedia | 4 |
| 2025 | Saliency-Guided Adaptive Random Diffusion for Remote Sensing Images Restoration with Cloud and HazeabstractRemote sensing image restoration under cloud and haze occlusions poses a significant challenge due to severe spectral degradation and spatial distortions. While recent generative models have shown promise in image restoration, they struggle with three key issues: (1) Lack of precise annotations, making supervised methods unreliable; (2) Unintended interference with clear regions, leading to distortion in unaffected areas; (3) Spectral and structural inconsistencies in heavily occluded regions, limiting realistic recovery. To address these challenges, we propose Saliency-Guided Adaptive Random Diffusion Strategy(SG-ARD), a novel blind restoration framework that integrates saliency-aware guidance with adaptive diffusion for enhanced reconstruction. First, we introduce a Saliency-Guided Pseudo-label Generation module (SGPG) to identify degraded regions and generate pseudo-labels for blind restoration. Second, we propose an Adaptive Random Diffusion Correction Strategy (ARDC), which employs a Random-Walk-based Diffusion and an Adaptive Enhancement module to refine local and global texture pseudo-labels. Lastly, we design a Spectral-Aware Consistency Loss (SAC) to improve spectral fidelity, ensuring that the generated content aligns with the real spectral distribution. Extensive experiments on three large-scale remote sensing datasets demonstrate that SG-ARD outperforms state-of-the-art generative restoration models, producing high-fidelity, visually coherent remote sensing images. Wanting Zhang, Libao Zhang |
ACM Multimedia | 3 |
| 2025 | Confidence-Guided Joint Complementary Learning for Weakly Annotated Remote Sensing Object SegmentationabstractObject segmentation from weakly annotated remote sensing images (RSIs) is an essential task that helps substantially reduce pixelwise labeling costs. Although mainstream multistage methods have achieved great performance, they generally suffer from high implementation complexity, which limits their practical application. On the other hand, for a single-stage scheme that can be trained in one cycle, error accumulation during joint optimization results in performance degradation. To address these issues, we propose a confidence-guided joint complementary learning (CGJCL) framework for remote sensing object segmentation under image-level annotations that can be easily trained in a single stage and achieve satisfactory performance. CGJCL integrates the localization and segmentation network into a unified framework that employs a joint complementary learning strategy, collaboratively enhancing the performance of each subnetwork. First, a confidence-guided pseudolabel refinement (CGPLR) module is developed to fuse the complementary semantic information of the localization clues and the object boundary/structure details learned from each subnetwork to generate high-quality pseudolabels (PLs) and alleviate error accumulation. Second, a dual self-supervised multiscale consistency (DSMC) loss is presented to both explicitly and implicitly utilize consistency regularization, enabling the promotion of model robustness on multiscale objects and preventing overfitting on false-annotated pixels in noisy PLs. Third, a locally enhanced feature aggregation network is proposed to integrate the multilevel semantic features and alleviate low-level noise interference, producing precise segmentation masks. Extensive evaluation of three RSI datasets demonstrates that the proposed method yields superior performance compared with recent single-stage techniques and several multistage methods, thus revealing its effectiveness and superiority. Libao Zhang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Object Segmentation Based on Pseudo Supervision Relearning Under Extremely Weak Annotations for Remote Sensing ImagesabstractWeakly annotated object segmentation for remote sensing images (RSIs) has attracted lots of attention due to its low labeling costs. However, annotating a huge amount of multiclass RSIs with image labels is still highly dependent on specific expert knowledge and, thus, requires considerable labeling costs. In this article, intending to further alleviate the labor-intensive labeling costs, we introduce a novel extremely weak annotation condition. In this condition, a substantial portion of samples are unlabeled, while merely a small number of samples are labeled with inexact image-level annotations. To achieve object segmentation under extremely weak annotations, we propose pseudo supervision relearning (PSRL), a novel three-stage framework with the core insight of effectively harnessing the potentially valuable supervision clues stored in abundant unlabeled data. In the first stage, the extremely weak annotations are switched to fine-grained but noisy pseudo supervision with the aid of image-level semantic learning and attention-guided data augmentation. Then, a category-aware dataset resplit strategy based on the masking perturbation mechanism is designed, aiming at adaptively selecting high-quality pixelwise pseudomasks from the artificially generated pseudo supervision and achieving class-balanced reliable-unreliable labels division. Ultimately, we devise a novel dynamic thresholding strategy (DTS)-guided relearning network to take full advantage of the valuable semantic information in the resplit dataset. Experimental results on two public RSI datasets show the effectiveness of the proposed framework. Utilizing less supervised information, the proposed method yields competing results compared to weakly supervised learning-based methods with complete image-level annotations. Wanning Zhu, Libao Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Semantic-Enhanced ULIP for Zero-Shot 3D Shape RecognitionabstractIn recent years, the zero-shot image recognition with semantic knowledge has achieved good performance due to vision-language models. However, because of the complexity of 3D shapes, the model cannot fully use the semantic knowledge of 3D shapes, which results in low accuracy of zero-shot 3D shape recognition. To address this problem, we propose a Semantic-enhanced ULIP for Zero-shot 3D Shape Recognition (SE-ULIP). This method utilizes the contrastive learning to fine-tune the text encoder in two stages, including the domain adaptation fine-tuning and the triplets-based text encoder fine-tuning. In the domain adaptation fine-tuning, we fine-tune the image encoder and the text encoder using the views and the Semantic Descriptive Text (SDT) of each view generated by the Visual Question Answering (VQA) model, which aims to align the view features with the semantic knowledge. In the triplets-based text encoder fine-tuning, we propose an Adaptive Conditional Adjustment Context Optimization (ACACoOp) to learn the optimal context vectors. The optimal context vectors are used as the input to fine-tune the text encoder again, which enhance SE-ULIP to understand the semantic knowledge of 3D shapes. Experiments show that our method achieves the state-of-the-art performance through the fine-tuned text encoder on three 3D backbone networks for both zero-shot and standard 3D shape recognition. Bo Ding 0003, Libao Zhang, Yongjun He 0002 |
IEEE Trans. Multim. | 2 |
| 2024 | Phase Learning Based on Interactive Perception for Limited-Sample Residential Area Semantic SegmentationabstractDue to the rich details of residential areas and the characteristics of remote sensing image sharpness vulnerable to haze, it will not only consume a lot of labor costs but also be very difficult to produce a large-scale dataset with strong labels. Therefore, the limited-sample dataset has become a hotspot in recent years. To address this issue, we proposed a semantic segmentation method for residential areas by phase learning. The main task of the first stage is to generate a joint saliency map by reducing the interference of haze noise through the feature comparison similarity sorting algorithm and combine them to generate initial pixel-level pseudo labels for the next stage of training. In the second stage, we proposed to construct a group feature interactive perception module to achieve image group semantic co-segmentation. Comprehensive evaluations with 2 datasets and the comparison with 7 methods validate the superiority of the proposed model. Xinran Lyu, Libao Zhang |
ICASSP | 2 |
| 2024 | Semantic Segmentation for Multi-Scene Remote Sensing Images with Noisy Labels Based on Uncertainty PerceptionabstractAs the annotation of remote sensing images requires domain expertise, it is difficult to construct a large-scale and accurate annotated dataset. Image-level annotation data learning has become a research hotspot. In addition, due to the difficulty in avoiding mislabeling, label noise cleaning is also a concern. In this paper, a semantic segmentation method for remote sensing images based on uncertainty perception with noisy labels is proposed. The main contributions are three-fold. First, a label cleaning method based on iterative learning is presented to handle noise labels such as missing or incorrect annotations. Second, a two-stage semantic segmentation model is proposed for image-level annotation, which eliminates the need for post-processing steps during testing. Lastly, a complementary uncertainty perception function is introduced to improve the utilization of dataset features and enhance the accuracy of segmentation. The effectiveness of this method was verified through comprehensive evaluation with 7 models on four datasets. Xinran Lyu, Libao Zhang |
ICASSP | 2 |
| 2024 | Hazy Remote Sensing Images Semantic Segmentation for Weakly Annotation Based on Saliency-Aware Alignment StrategyabstractThe technique of semantic segmentation (SS) holds significant importance in the domain of remote sensing image (RSI) processing. The current research primarily encompasses two problems: 1) RSIs are easily affected by clouds and haze; 2) SS based on strong annotation requires vast human and time costs. In this paper, we propose a weakly supervised semantic segmentation (WSSS) method for hazy RSIs based on saliency-aware alignment strategy. Firstly, we design alignment network (AN) and target network (TN) with the same structure. After training the AN with clear images, we extract the class activation maps of the two networks and construct a consistency loss to train the TN with hazy images. Secondly, we design a multi-scale channel-spatial attention module in the two classification networks to solve the problems of unclear foreground-background boundary and blurred target texture in hazy images. Finally, the pseudo-labels generated by the TN are utilized to train a feedback saliency analysis network, which is subsequently employed for obtaining segmentation results during the testing phase. The experiment results demonstrate that our approach achieves superior SS performance for hazy RSIs. Junda Xu, Libao Zhang |
ICASSP | 2 |
| 2024 | Unsupervised Remote Sensing Haze Removal Based on Saliency-Guided Transmission RefinementabstractHaze causes information loss and quality degradation in remote sensing images. Unsupervised learning-based dehazing methods aim to reduce reliance on paired hazy images and their labels. However, complex mapping relationships often increase the difficulty in network convergence, resulting in color distortion and loss of texture details in remote sensing images. To address these issues, we propose an unsupervised haze removal method based on saliency-guided transmission refinement for remote sensing images. Firstly, we introduce a saliency-guided transmission refinement method, which decomposes and recombines two transmission maps obtained under different conditions, guided by saliency information. Secondly, we propose a loss function comprising energy loss and texture loss. The energy loss provides an energy reference based on the coarse transmission estimation, while the texture loss enhances the preservation of texture details. Experimental results demonstrate that our method achieves comparable performance to several supervised methods. Ruohui Zheng, Libao Zhang |
ICASSP | 2 |
| 2024 | SDRNet: Saliency-Guided Dynamic Restoration Network for Rain and Haze Removal in Nighttime ImagesabstractDue to the different physical imaging models, most haze or rain removal methods for daytime images are not suitable for nighttime images. Fog effect produced by the accumulation of rain also brings great challenges to the restoration of low-light nighttime images. To deal well with the multiple noise interference in this complex situation, we propose a saliency-guided dynamic restoration network (SDRNet) that can remove rain and haze in nighttime scenes. First, a saliency-guided detail enhancement preprocessing method is designed to get images with clearer details as the auxiliary input. Second, following a rain removal network (RRN), we design an all-in-one nighttime dehazing network (ANDN) to estimate the spatially variable ambient light and transmission comprehensively by deforming the nighttime haze image model. Finally, an attention-based enhancement network (AEN) with dynamic fusion attention module is proposed to enhance the lowlight background image. Experimental results indicate that SDRNet can obtain clearer images with less fog and distortion compared with other methods. Wanning Zhu, Libao Zhang |
ICASSP | 3 |
| 2024 | Land Use Classification Via Multi-Modal Complementary Feature Fusion and Context Information Enhancement For Optical and Sar ImagesabstractLand use classification by fusing optical and synthetic aperture radar (SAR) images has become a research hotspot since it can greatly improve segmentation accuracy. However, due to their different object expression patterns, existing multimodal algorithms have problems such as insufficient utilization of complementary features and inability to solve class imbalance. In this paper, we develop a multi-modal semantic segmentation method based on complementary features fusion and context information enhancement. First, we propose a dual-branch network with no shared weights to extract features of optical images and SAR images respectively. The multi-channel parallel convolution (MCPC) blocks in the network can improve the receptive field of feature extraction. Then, we propose a multi-modal complementary feature fusion (MCFF) module. The two modalities are encouraged to exchange complementary information and suppress redundant information. Finally, for advanced semantic features, we design the context information enhancement (CIE) module to capture multi-scale semantic information and increase feature utilization efficiency to a greater extent. The result of comparative experiments with state-of-the-arts proves the effectiveness of the network. Xinyue Fan, Libao Zhang |
ICIP | 2 |
| 2024 | Remote Sensing Image Uneven Haze Removal Based On Haze Density Estimation and Saliency-Driven Dual Channel FusionabstractRemote sensing images (RSIs) are usually degraded by haze, losing the spectral fidelity and texture details. Most previous dehazing works adopted a unified dehazing strategy to process the entire image, ignoring the discrepancies of haze density, spectral information and texture complexity among different regions in one RSI, and thus resulted in the distorted spectral and blurred texture. In this paper, we propose a skip-connected network based on haze density estimation and saliency-driven dual channel fusion, realizing a differentiated defogging method. First, we propose a haze density estimation model, generating the haze density map. Second, we design a saliency-driven dual channel fusion to distinguish the spectral and texture features of different regions in the input, generating the saliency map. The two maps above jointly serve as guidance for the network to perform differentiated dehazing. Finally, an attention-based skip-connected block is introduced, which helps to fuse multi-scale information and obtain more refined dehazing results. A mixed loss function is also constructed to retain more texture details. The efficiency of the proposed method is validated by comparing its performance with other SOTA schemes. Yanmeng Liu, Libao Zhang |
ICIP | 2 |
| 2024 | Dynamic Activation Function Based on the Branching Process and its Application in Image ClassificationabstractThe choice of activation function in deep learning is crucial to the performance of neural networks. The activation function used in conventional deep learning remains unchanged for neural networks of different depths, leading to performance degradation as the depth of the model increases. In this paper, we propose a $\operatorname{sigmoid}_{n}$ dynamic activation function that can change with the depth of the neural network. We firstly introduced the dual relationship between the activation function and the probability generating function(PGF) from the perspective of the branching process, and explained the reason why the model performance of different activation functions decreases as the neural network deepens. Then, we use the law of large numbers in the super critical branching process to optimize the PGF and propose the sigmoid ${ }_{n}$ dynamic activation function through the dual relationship between the PGF and the activation function. Finally, to better extract the spatial context information of the image, we add a convolution channel based on the sigmoid ${ }_{n}$ dynamic activation function and propose a two-dimensional Fsigmoid ${ }_{n}$ dynamic activation function. Experiments on CIFAR-10 and CIFAR-100 datasets verify the superiority of the proposed sigmoid ${ }_{n}$ activation function. Wanting Zhang, Libao Zhang |
ICIP | 2 |
| 2024 | Clouds and Haze Co-Removal Based on Weight-Tuned Overlap Refinement Diffusion Model for Remote Sensing ImagesabstractRemote sensing image dehazing is essential for preprocessing, but most methods overlook the joint occlusion by clouds and haze. Furthermore, generative model-based restoration struggles with haze, weakening edge recovery for targets obscured by both elements. In this paper, we propose a new diffusion model based on weight-tuned overlap refinement for clouds and haze co-removal in remote sensing images. Firstly, we integrate the diffusion model into clouds and haze co-removal task, offering a lightweight solution effective even under limited samples. Secondly, we propose a weight-tuned overlap refinement method (WTOR) to guide noise estimation updating in overlaps in the backward-sampling process. It takes the reciprocal of the variance of each overlap as the weight factor, adjusts the backward-sampling update frequency to improve the model’s ability of clouds and haze co-removal. Finally, we introduce the gaussian filter feature extraction block (GFAB) to enhance the learning ability of the model for structural information, which can detect local changes and extract features with precision. The experimental results demonstrate that the proposed method is capable of co-removing clouds and haze with reserving rich color and texture details. Libao Zhang |
ICIP | 2 |
| 2024 | Weakly Supervised Semantic Segmentation of Remote Sensing Images Based on Progressive Mining and Saliency-Enhanced Self-AttentionabstractGiven the high demands of effort in generating pixel-level annotations, weakly supervised semantic segmentation (WSSS) has become an important approach for remote sensing image (RSI) interpretation. However, current methods are mostly borrowed from natural scene studies, regardless of the significant variation in object sizes as well as the highly confusing intraclass heterogeneity and interclass homogeneity which are characteristic of RSIs. In this letter, we propose a WSSS method based on progressive mining and saliency-enhanced self-attention, to efficiently segment RSIs with image-level labels. First, we exploit multiscale orientation patterns to sufficiently extract the rich texture in RSIs which can help to discern between the different classes, and combine this information with contrast and luminance features to generate fine saliency maps. Second, we design a progressive mining process to gradually discover both the large objects, representative of semantics, and the small objects, rich in patterns. Finally, we employ self-attention mechanism to capture global dependencies in RSIs for refining category areas. To inhibit the mis-spread of attention, we use saliency as a mask discerning between the background and the object classes. Experiments on different datasets demonstrate the competence of the proposed method, in terms of both metrical results and visual effects. Ting Hao, Shuya Bai, Libao Zhang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Semantic Segmentation of Weakly Annotated Remote Sensing Images Based on Feature Adversary and Uncertainty PerceptionabstractThe rich details in remote sensing images require the expertise of domain experts for annotation, making precise labeling across multiple scenarios challenging. Currently, there is an increasing interest in low-precision image-level annotations. However, errors and omissions are difficult to avoid during the labeling process. How to achieve accurate semantic segmentation under noisy annotations has become an urgent problem that requires resolution. In this letter, we consider the rich features of remote sensing targets and design a weakly labeled semantic segmentation model for remote sensing images based on feature adversary and uncertainty perception. First, we introduce a confidence model voting method to handle missing or incorrect image-level labels. Subsequently, we perform multilevel feature fusion to obtain initial pixel-level pseudo labels. Finally, leveraging the significant differences in image features across different categories, we design a feature adversarial model and introduce an uncertainty analysis method, improving the utilization of remote sensing image features and enhancing the accuracy of semantic segmentation. The effectiveness of this approach is validated through a comprehensive evaluation of four datasets. Xinran Lyu, Ruohui Zheng, Libao Zhang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Deformable Convolution Alignment and Dynamic Scale-Aware Network for Continuous-Scale Satellite Video Super-ResolutionabstractRecently, due to higher requirements for satellite video resolution, video super-resolution (VSR) has been extensively studied. However, the following problems have not been effectively resolved: 1) Previous satellite VSR methods cannot achieve continuous-scale (integer and non-integer scale) VSR with a single model. 2) Satellite video has complex ground and weak textures, which increases the difficulty of capturing motion information. In addition, existing methods adopt a unified alignment path, which leads to a drop in feature alignment accuracy. 3) During feature fusion, previous methods ignore the correlation of spatio-temporal information in satellite video and cannot make full use of the spatio-temporal information. To address the above problems, in this paper, we propose a novel network for continuous-scale satellite VSR (CSVSR). Specifically, first, for effective motion capture and accurate feature alignment, we design a residual-guided and time-aware dynamic routing alignment module, which can use feature residuals to lock motion areas and then dynamically select the corresponding alignment path based on the temporal distance. Second, we proposed a non-local mask-based feature fusion module to exploit the correlation of the spatio-temporal features and complete effective spatio-temporal feature fusion. Third, to make our network adapt to multi-task learning, we develop a scale-aware convolutional (SA-Conv) layer, which lets our network dynamically extract scale-adaptive features according to the input scale factors. Finally, we propose a continuous-scale upsampling module with a global feature implicit function (GFIF), which can achieve continuous-scale mapping from features to pixel values. In addition, we carefully design a novel training strategy to optimize our network. Comprehensive experiments verify that the proposed CSVSR has superior reconstruction performance on continfuous-scale factors. The code will be available at https://github.com/chongningni/CSVSR. Ning Ni 0002, Libao Zhang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | A Unified Framework for Double-Degradation Remote Sensing Image Restoration Through Saliency-Guided Interaction LearningabstractRemote sensing images (RSIs) are often exposed to various degradation factors such as sensor noise and poor observation environments. These factors can result in the loss of image details, spectral distortion, and blurred scenes, which will severely hinder the performance of subsequent RSI applications. Unfortunately, current methods are only capable of addressing individual degradation factors, and lack the ability to handle multiple degradation factors within a unified framework. To mitigate this problem, we model the complex degradation process as a double-degradation model for RSIs, and propose a unified framework based on saliency-guided interaction learning (SGIL) for double-degradation RSI restoration. The proposed SGIL can simultaneously alleviate the influence of external environment degradation and internal sensor noise degradation, which comprises three parts: pseudo pixel supervision-based saliency analysis (PPS-SA), a task-aware interaction learning (TAIL) model, and a global feature enhancement module (GFEM). In PPS-SA, an explicit PPS-SA method is designed to generate saliency maps to effectively distinguish different texture complexities of RSIs, and a saliency-guided mapping selection mechanism is introduced to adapt to complex external environmental interference factor. In the TAIL model, a dehazing module and a super resolution (SR) module are specialized in alleviating the external environment interference and internal sensor noise, respectively. Instead of simple cascading, the two modules interact and collaborate with each other, which drastically improves the performance of double-degradation image restoration. To further improve the performance of the proposed SGIL, we also propose the GFEM to exploit global features and refine the restored results. Experiments on several RSI datasets demonstrate that the proposed SGIL achieves promising results on complex double-degradation RSIs. Shan Wang 0009, Libao Zhang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | 3D shape classification based on global and local features extraction with collaborative learning
Bo Ding 0003, Libao Zhang, Yongjun He 0002 |
Vis. Comput. | 2 |
| 2023 | Progressive Refinement Learning Based on Feature Cross Perception for Residential Areas Semantic SegmentationabstractDue to the pixel-level accurate annotation of remote sensing images consumes a lot of labor costs, weak annotation semantic segmentation has become a hotspot in recent years. However, due to the lack of label accuracy, these methods often have insufficient expression ability. In this paper, we proposed a semantic segmentation method for residential areas by progressive refinement learning. This method mainly consists of two parts. In the first part, we constructed a classification network and proposed an initial pixel-level label calculation method based on multi-layer category feature awareness. In the second part, we proposed to construct feature cross perceptron module in the structure of the multi-level codec to achieve image pair semantic co-segmentation. In addition, we used confidence maps to modify the loss function to achieve more accurate results. Comprehensive evaluations with GeoEye-1 dataset and the comparison with 7 methods validate the superiority of the proposed model. Xinran Lyu, Libao Zhang |
ICASSP | 2 |
| 2023 | UAV Remote Sensing Image Dehazing Based on Multi-Dimensional Saliency Awareness Unequal NetworkabstractCurrent UAV image haze removal methods often suffer from problems of insufficient dehazing and spectrum distortion, especially in regions with rich spectrum and texture information. In this paper, we propose a multi-dimensional saliency awareness unequal network to avoid texture loss and color distortions. First, we design an unequal network structure, which enhances local color and texture detail feature learning. Specifically, we propose a saliency dense block, which employs a saliency map to guide the network to pay attention to the regions with rich spectrum and texture information unequally. Second, a multi-dimensional saliency detection method is proposed. It fuses the features of different dimensions to extract the salient regions, to obtain the spectrum and texture information. Finally, a loss function combining mean square error and edge loss is defined to enhance the learning of the texture. The experimental results demonstrate that the proposed method is capable of restoring color information as well as reserving texture details, especially for salient regions of UAV hazy images. Ruohui Zheng, Libao Zhang |
ICASSP | 2 |
| 2023 | DSG-PL: ROI Extraction Based on Dual Saliency Guided Progressive Learning for Weakly Labeled Remote Sensing ImagesabstractRegion of Interest (ROI) extraction from weakly annotated remote sensing images (RSIs) can save the huge labor cost of labelling accurate pixel-level annotations. However, weakly supervised approaches with sparseness and incompleteness inevitably result in a performance gap compared with fully supervised counterparts. To tackle this issue, a dual saliency guided progressive learning (DSG-PL) framework is developed, which focuses on progressively enhancing the quality of supervision from the image level to the precise pixel level. To begin, a dual saliency constraint mechanism is created to guide the training of a classification network in both an explicit and implicit manner for generating integral pixel-wise pseudo labels (PLs). Then, to gradually refine the initial PLs, an adaptive label self-correction module is presented, in which the updated labels are used to iteratively train a context-enhanced segmentation network, therefore constantly boosting model performance. Finally, a confidence-aware denoising loss is intended to alleviate the impacts of training with noisy PLs by adaptively reweighting the pixel-wise loss with confidence scores. Comprehensive evaluations and ablation studies verify the superiority of the proposed DSG-PL. Libao Zhang |
ICIP | 2 |
| 2023 | Mutually Supervised Learning via Interactive Consistency for Geographic Object Segmentation from Weakly Labeled Remote Sensing ImageryabstractGeographic object segmentation from weakly annotated remote sensing images has become a research hotspot, since it can greatly reduce the costly annotation burden. Recently, it has made remarkable progress by dividing it into two sequential steps, which first produces pseudo labels (PLs) from a localization model, then uses PLs to train a segmentation network for final results. The one-way knowledge transfer in the above schemes, however, lacks the feedback from the segmentation to localization model which may result in suboptimal performance. In this paper, we develop a mutually supervised learning (MSL) framework for geographic object segmentation under image-wise annotations. First, MSL learns the localization and segmentation model concurrently and employs the output from each of the two models as pseudo supervision for the other one by formulating an interactive consistency loss, which encourages each model to provide positive feedback and guidance to the other. Then, a variance-based uncertainty estimation strategy is introduced to explicitly approximate the uncertainty of the PLs, which helps to alleviate the detrimental effect caused by learning from noisy PLs. Finally, we design a multi-scale activation integration-based localization model to produce high-quality localization maps. Comprehensive evaluations and ablation studies validate the superiority of the MSL framework. Libao Zhang |
ICIP | 2 |
| 2023 | Progressive Refinement Learning Based on Feature Interactive Fusion for Semantic Segmentation of Remote Sensing Limited DatasetabstractDue to the labor cost and the accuracy of manual identification, it is very difficult to make a strong label dataset of remote sensing images with a large amount of data. Therefore, the limited remote sensing dataset has become a research hotspot in recent years. However, due to insufficient precision and the lack of label accuracy, these methods often have insufficient expression ability. In this paper, we proposed a semantic segmentation method for remote sensing images by progressive refinement learning. Firstly, we construct multiple classification networks to vote for label noise cleaning, and select a network to retrain. Then, the method based on hierarchical feature learning is used to realize the pixel-level pseudo label calculation. Secondly, we proposed to construct feature interactive fusion module in the multi-level codec to achieve image group semantic segmentation. Comprehensive evaluations and the comparison with 7 methods validate the superiority of the proposed model. Xinran Lyu, Libao Zhang |
ICIP | 2 |
| 2023 | Residential Extraction Based on Weakly-Supervised Similarity-Aware Multi-Source Alignment Strategy with Limited SAR DataabstractResidential extraction based on deep learning approach is a significant task in Synthetic Aperture Radar (SAR) image processing. However, extremely limited SAR data brings great challenges to the data-driven method: 1) pixel-wise annotations are hard to obtain due to the expensive cost; 2) There is not any universal large-scale SAR image dataset with heterogeneous SAR images. In this paper, a novel residential extraction method based on similarity-aware multi-source alignment strategy is proposed to solve such problems. Firstly, we propose a weakly-supervised Multi-source Similarity-aware Extraction Network (MSENet) to preserve the context dependency of pixels and improve the integrity of the extraction. Then, to tackle with the lack of training samples, a multi-source knowledge alignment strategy is proposed to learn transferrable knowledge from heterogeneous SAR datasets. Finally, affinity-guided optimization is introduced to refine the coarse maps with clear boundaries. Comprehensive experiments demonstrate the efficiency of our method. Sijia Ma, Libao Zhang |
ICIP | 2 |
| 2023 | Sar Target Extraction Based On Saliency-Guided Cross-Domain Discrepancy Alignment StrategyabstractTarget extraction based on deep learning approaches is a significant task in Synthetic Aperture Radar (SAR) image processing. However, the lack of SAR image samples brings great challenges to the data-driven method. In this paper, a Saliency-guided Cross-domain Discrepancy Alignment strategy is proposed to solve this problem. Firstly, we propose a saliency-guided attention module, which utilizes the context-aware saliency knowledge to guide the feature extraction and improve the training efficiency. Secondly, we train the Saliency-guided Cross-domain Alignment Network (SCANet) by large-scale natural optical image dataset and tiny-scale SAR image dataset. Thirdly, based on the guidance of saliency attention, cross-domain representation alignment strategy is proposed to learn a latent representation which aligns feature distribution between the source and target domain. Finally, SCANet is more adaptive for SAR images and extracts targets more accurately. Comparison with state-of-the-arts and ablation experiments demonstrate the efficiency of our method, especially in complex conditions. Sijia Ma, Libao Zhang |
ICIP | 2 |
| 2023 | Nighttime Haze Removal with Spatially Variant Ambient Light and Saliency-Weighted Fused TransmissionabstractDifferent from hazy images captured in the daytime, the ambient illumination of nighttime hazy images is not globally homogeneous. Highlight and lowlight regions have different transmission properties in nighttime hazy scenes. In this paper, we propose a nighttime dehazing model without using the dark channel prior. We estimate the ambient illumination via the Difference of Gaussian (DoG), which selectively retains the high-frequency edges of brightness mutation. We propose a coefficient fusion algorithm in the LAB color space using Homomorphic filtering to estimate the transmission of highlight regions. And we propose a Retinex-like transmission estimation model for lowlight regions. Then we acquire the global saliency-weighted fused transmission. Finally, we get the haze-free results via the nighttime atmospheric scattering model. Experimental results show that our method outperforms other state-of-the-art methods in both color and detail recovery. Ruohui Zheng, Libao Zhang |
ICIP | 3 |
| 2023 | Hazy Remote Sensing Image Restoration Based on Saliency-Guided Transmission Optimization and Texture BoostingabstractRemote sensing images (RSIs) are susceptible to haze, losing the spectral fidelity and texture details. Haze removal is highly desired in the follow-up tasks such as target identification and semantic segmentation. Most previous works adopted a unified dehazing method to process the entire image and could not fully restore the rich texture information contained in RSIs. In this paper, we propose a hazy RSI restoration method based on saliency-guided transmission optimization and texture boosting. First, we design a saliency-guided transmission optimization method, which achieves different degrees of dehazing to areas with different saliency, fully obtaining the texture and spectral information of RSIs. Second, we propose a saliency-guided atmospheric light (AL) correction method, which fuses the AL of salient regions and non-salient regions to avoid excessive energy attenuation. Finally, a saliency-guided RSI texture boosting method is introduced, further enhancing the texture details of the dehazed RSIs. The efficiency of the proposed method is validated by comparing its performance with six state-of-art schemes. Yanmeng Liu, Libao Zhang |
IGARSS | 2 |
| 2023 | Target Extraction Based on Cross-Domain Alignment and Self-Correlation Mechanism With Weak-Labeled SAR DataabstractTarget extraction is a significant task in Synthetic Aperture Radar (SAR) image processing. Recently, SAR target extraction with weak labels has attracted great attention due to the low labeling cost. However, weak-labeled SAR data brings great challenges to the data-driven methods: 1) location and structural information of the targets are lost in weak labels; 2) discrepancy of heterogeneous SAR images restricts the training efficiency of the model. In this paper, a novel Cross-domain Self-correlation Aware Network (CSANet) for SAR target extraction based on image-level weak labels is proposed to address such challenges. Firstly, the Cross-domain Representation Alignment (CRA) strategy is proposed to learn transferrable knowledge from heterogeneous SAR datasets. Through cross-domain alignment, invariant feature space is constructed to bridge the heterogeneous SAR data and improve the generalization performance of the model. Then, we propose a self-correlation aware extraction module with image-level weak labels, which only indicate whether the images contain the targets or not. Self-Correlation Module (SCM) is designed to preserve the context dependency of SAR pixels and compensate for the gap between weak labels and dense prediction. Finally, Affinity-Guided Optimization (AGO) is introduced to learn the inner-pixel affinity and refine the coarse extraction maps with clear boundaries. Comparison with state-of-the-arts and the ablation experiments demonstrate the efficiency of our method. Sijia Ma, Libao Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | UAV Image Haze Removal Based on Saliency- Guided Parallel Learning MechanismabstractCurrent haze removal methods for unmanned aerial vehicle (UAV) images are mostly based on natural image dehazing methods, which ignore the particular imaging mechanism of UAVs. They often fail to restore regions with rich spectral and textural information. In this letter, we propose a saliency-guided parallel learning mechanism for UAV image haze removal. First, we design a saliency-guided parallel dehazing module with two parallel paths. The residual feature extraction path obtains deep-level features to realize global dehazing effectively. The key feature enhancement path, which comprises saliency dense blocks, realizes local textural preservation and spectral restoration. Second, a sporadic foreground saliency detection method is proposed for UAV images with sporadic objects. The saliency map guides the learning of significant spectral and textural information in hazy images. Finally, a multiscale reconstruction module is introduced to more accurately estimate small-scale textural details and large-scale spectral information. Experimental results show that the proposed method has better detail performance and visual effects than state-of-the-art methods. Ruohui Zheng, Libao Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Weakly Annotated Residential Area Segmentation Based on Attention Redistribution and Co-LearningabstractResidential area (RA) segmentation is of great significance in the remote-sensing (RS) field. Training the segmentation network with image-level weakly annotated data (WAD) has become a research hot spot due to the easy access to classification labels. The quality of class activation maps (CAMs) is crucial for obtaining accurate segmentation results. Limited to irregular shapes, tortuous boundaries, and greatly varied scales of targets in RS images, generating high-quality CAMs is still a great challenge. To solve these problems, a novel weakly annotated RA segmentation model based on attention redistribution and co-learning (ARC) is proposed in this letter. We develop aggregate-and-distribute-based feature coupling (ADFC) to achieve the redistribution of attention on channel and spatial dimensions, which deals with multilevel features at the same time and makes them fully embedded together. Such an arrangement can effectively capture the shape characteristic of targets and filter out complicated backgrounds. To mitigate the impact of ambiguous regions like surroundings of boundaries and potential scattered houses, a confusion co-learning (CCL) strategy is designed to jointly explore the class-specific features and refine the cross-class features through a two-stream classifier with sharing weights, which helps generate sharper edges and discover ignored targets. Experimental results on GeoEye-1, SPOT5, and Landsat8 datasets reveal that our proposal outperforms the competing methods by a large margin in both subjective and objective assessments. Wanning Zhu, Libao Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Conditional Stochastic Normalizing Flows for Blind Super-Resolution of Remote Sensing ImagesabstractRemote sensing images (RSIs) in real scenes may be disturbed by multiple factors, such as optical blur, undersampling, and additional noise, resulting in complex and diverse degradation models. At present, mainstream super-resolution (SR) algorithms only consider a single and fixed degradation (such as bicubic interpolation) and cannot flexibly handle complex degradations in real scenes. Therefore, designing an SR model that can deal with various degradations has gradually attracted researchers’ attention. Some early studies estimate degradation kernels and then perform degradation-adaptive SR but face the problems of estimation error amplification and insufficient high-frequency details in the results. Although blind SR algorithms based on generative adversarial networks (GANs) have greatly improved visual quality, they still suffer from pseudo-texture, mode collapse, and poor training stability. This article proposes a novel blind SR framework based on the stochastic normalizing flow (BlindSRSNF) to address the above problems. BlindSRSNF learns the conditional probability distribution over the high-resolution image space given a low-resolution (LR) image by explicitly optimizing the variational bound on the likelihood. BlindSRSNF is easy to train and can generate photorealistic SR results that outperform GAN-based models. In addition, we introduce a degradation representation strategy based on contrastive learning to avoid the error amplification problem caused by explicit degradation estimation. Comprehensive experiments show that the proposed algorithm can obtain SR results with excellent visual perception quality on both simulated LR and real-world RSIs. The code is available at https://github.com/hanlinwu/BlindSRSNF. Ning Ni 0002, Shan Wang 0009, Libao Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Lightweight Stepless Super-Resolution of Remote Sensing Images via Saliency-Aware Dynamic Routing StrategyabstractDeep learning-based algorithms have greatly improved the performance of remote sensing image (RSI) super-resolution (SR). However, increasing network depth and parameters cause a huge burden of computing and storage. Directly reducing the depth or width of existing models results in a large performance drop. We observe that the SR difficulty of different regions in an RSI varies greatly, and existing methods use the same deep network to process all regions in an image, resulting in a waste of computing resources. In addition, existing SR methods generally predefine integer scale factors and cannot perform stepless SR, i.e., a single model can deal with any potential scale factor. Retraining the model on each scale factor wastes considerable computing resources and model storage space. To address the above problems, we propose a saliency-aware dynamic routing network (SalDRN) for lightweight and stepless SR of RSIs. First, we introduce visual saliency as an indicator of region-level SR difficulty and integrate a lightweight saliency detector into the SalDRN to capture pixel-level visual characteristics. Then, we devise a saliency-aware dynamic routing strategy that employs path selection switches to adaptively select feature extraction paths of appropriate depth according to the SR difficulty of subimage patches. Finally, we propose a novel lightweight stepless upsampling module whose core is an implicit feature function for realizing mapping from low-resolution feature space to high-resolution feature space. Comprehensive experiments verify that the SalDRN can achieve a good tradeoff between performance and complexity. The code is available athttps://github.com/hanlinwu/SalDRN. Ning Ni 0002, Libao Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Learning Dynamic Scale Awareness and Global Implicit Functions for Continuous-Scale Super-Resolution of Remote Sensing ImagesabstractThe mainstream remote sensing image (RSI) super-resolution (SR) algorithms treat tasks with different scale factors independently, and a single model can only process a fixed integer scale factor. However, in practical applications, it is important to continuously super-resolve RSIs to multiple resolutions, as different resolutions present various levels of detail. Retraining the model for each scale factor consumes huge computational resources and storage space. Existing continuous-scale SR models employ static convolutions, and most are designed for natural scenes, ignoring dynamic feature extraction needs for different scale factors and the inherent properties of RSIs. In addition, efficiently obtaining the continuous representation of RSIs and avoiding the artifacts of RSI SR results is still a challenging problem. To address the above problems, we propose a scale-aware dynamic network (SADN) for RSI continuous-scale SR. First, we devise a scale-aware dynamic convolutional (SAD-Conv) layer to handle the strong randomness of the RSI textural distribution and achieve dynamic feature extraction according to scale factors. Second, we devise a continuous-scale upsampling module (CSUM) with the multi-bilinear global implicit function (MBGIF) for any-scale upsampling. The CSUM constructs multiple feature spaces with asymptotic resolutions to approximate the continuous representation of an image, and then, the MBGIF makes full use of multiresolution features to map arbitrary coordinates to spectral values. We evaluate our SADN using various benchmarks, and the experimental results show that the CSUM can efficiently achieve continuous-scale upsampling while maintaining excellent objective and visual performance. Moreover, our SADN uses fewer parameters and even outperforms the state-of-the-art fixed-scale SR methods. The source code is available athttps://github.com/hanlinwu/SADN. Ning Ni 0002, Libao Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Deformable Alignment And Scale-Adaptive Feature Extraction Network For Continuous-Scale Satellite Video Super-ResolutionabstractVideo super-resolution (VSR), especially continuous-scale VSR, plays a crucial role in improving the quality of satellite video. Continuous-scale VSR aims to use a single model to process arbitrary (integer or non-integer) scale factors, which is conducive to meeting the needs of video images transmission with different compression ratios and arbitrarily zooming by rolling the mouse wheel. In this article, we propose a novel network to achieve continuous-scale satellite VSR (CAVSR). Specifically, first, we propose a time-series-aware dynamic routing deformable alignment module (TDAM) for feature alignment. Second, we develop a scale-adaptive feature extraction module (SFEM), which uses the proposed scale-adaptive convolution (SA-Conv) to dynamically generate different filters based on the input scale information. Finally, we design a global implicit function feature-adaptive walk continuous-scale upsampling module (GFCUM), which can perform feature-adaptive walks according to the input features with different scale information and finally complete the continuous-scale mapping from coordinates to pixel values. Experimental results have demonstrated the CAVSR has superior reconstruction performance. Ning Ni 0002, Libao Zhang |
ICIP | 3 |
| 2022 | Dynamic Mutual Enhancement Network for Single Remote Sensing Image DehazingabstractIn this paper, we propose a dynamic mutual enhancement network (DMENet) for haze removal in remote sensing images. It has three major advantages compared with other dehazing algorithms: 1) The proposed DMENet is based on the U-Net architecture to extract features effectively, which is composed of three components, i.e., a multi-scale encoder, a middle transmission layer (MTL), and a dynamic mutual decoder. 2) The dynamic mutual enhancement (DME) module is designed to dynamically integrate multi-level feature maps in a mutual way, which contains the low-level detail information and high-level semantic information respectively. 3) To improve the robustness and generalization performance of the DMENet, the hybrid supervision is built for network training between the restored results and their ground-truth labels, which consists of the pixel-level supervision, patch-level supervision and image-level supervision. Experimental results on both synthetic datasets and real remote sensing hazy images demonstrate that the proposed DMENet can gain significant progresses over the competing methods. Shan Wang 0009, Libao Zhang |
ICIP | 2 |
| 2022 | SD-DCRN: Saliency Driven Double-Channel Residual Network for Super-Resolution of Remote Sensing ImagesabstractTraditional methods of super-resolution of remote sensing images generally ignore the fact that significant areas usually have a higher demand for super-resolution compared to nonsignificant areas. According to this feature of remote sensing images, we propose a new model of super-resolutio-n based on double-channel residual dense network driven by saliency analysis. Firstly, we use a cascaded partial decoder model to obtain the saliency image of remote sensing images which contributes to distinguishing significant areas and background areas. Secondly, we adopt different super-resolution strategies for regions with different salient values and texture complexity. For the non-salient regions, we adopt the smaller number of RDBs and their internal convolution layers to save computer resources. For the salient regions, we increase the number of RDBs and layers to extract more complex features for super-resolution of the salient regions, which is conducive to the reconstruction of complex texture. Finally, the reconstructed salient regions and non-salient regions are superimposed to obtain the complete super-resolution results. Our experimental results show that the comprehensive perform of our method outperforms other super-resolution models in terms of metrics based on the peak signal-to-noise ratio and structural similarity. Libao Zhang |
IGARSS | 4 |
| 2022 | A Spatial Attention Guided Scene Classification Method for Multiscale Remote Sensing DatasetabstractRemote sensing image contains kinds of land cover. These surface coverings form complex and diverse scenes, which brings difficulties to scene classification. In recent years, many methods based on deep learning have achieved remarkable results. The existing researches mainly focus on training convolutional neural networks. However, these methods do not clearly distinguish the key information and redundant information of the images. Inspired by the attention mechanism, we propose a scene classification network, which combines the residual unit and spatial attention mechanism. It automatically allocates large weights to the key regions of the image, so it can ignore the redundant information adaptively. In addition, we designed a multi -scale classification result voting strategy for dataset with images of different scales to improve classification accuracy. We evaluated the proposed approach with four state-of-the-art methods on the two datasets. Experimental results show that the proposed model has achieved the best classification performance. Xinran Lyu, Libao Zhang |
IGARSS | 2 |
| 2022 | SD-DSAN: Saliency-Driven Dense Spatial Attention Network for Pan-SharpeningabstractThe demands for spectral and spatial quality in remote sensing (RS) images vary from region to region. Saliency detection is an effective tool to distinguish different regions with different demands. In this paper, we introduce saliency detection to satisfy these demands and propose a novel saliency-driven pan-sharpening network to further improve the fusion quality. Firstly, we combine foreground distribution with background prior to generate the initial saliency map, and implement least-square optimization to improve the detection accuracy. Then, we construct a dense spatial attention network trained through a new spatial-spectral-based loss function designed by saliency to meet diverse spectral and spatial needs of different regions. Thus, accurate fused images can be predicted. Experiments on SPOT-5 dataset indicate that our proposal has excellent properties with respect to the unified spatial-spectral quality against state-of-the-art methods. Wanning Zhu, Yang Sun 0007, Shan Wang 0009, Libao Zhang |
IGARSS | 4 |
| 2022 | Region of Interest Extraction Based on Bayesian Joint Saliency Detection for Remote Sensing ImagesabstractSaliency detection is an essential tool to extract regions of interest (ROIs) in remote sensing (RS) images. However, many methods are applied to single image and cannot detect ROIs accurately due to the ignorance of high correlation among different RS images. Thus, we propose the Bayesian joint saliency detection method to extract ROIs. Firstly, we generate the prior saliency based on global color contrast according to co-clustering, which ensures that regions with similar features have the same saliency. Secondly, we produce the likelihood saliency by constructing intensity co-occurrence histogram, which can explore the intensity distribution of multiple images. Finally, due to the complex scenes in RS images, Bayesian enhancement strategy is applied to combine the prior saliency with the likelihood saliency, and obtain ROI with less background inference. Quantitative and qualitative experiments results indicate that our method outperforms competing methods and shows good performance in ROI extraction. Wanning Zhu, Libao Zhang, Yinggang Zhani |
IGARSS | 2 |
| 2022 | Hierarchical Feature Aggregation and Self-Learning Network for Remote Sensing Image Continuous-Scale Super-ResolutionabstractConducting research on remote sensing image (RSI) super-resolution (SR) is important, especially in terms of the continuous scale, which is beneficial to the application of RSI, such as RSI object detection and data fusion. Continuous-scale SR aims to use a single model to achieve SR at arbitrary (integer and noninteger) scale factors. Therefore, in this letter, we propose a hierarchical feature aggregation and self-learning network for RSI continuous-scale SR (RSI-HFAS). Our network can magnify the RSI continuously, which is beneficial for extracting the RSI multiscale features. First, we design a hierarchical feature aggregation module (HFAM) that is used for hierarchical feature extraction by placing convolutional layers on different floors and completing global feature fusion, which is crucial for achieving RSI continuous-scale SR with a single model. Second, the proposed network introduces a feedback mechanism, which can refine the hierarchical feature through feature feedback and enrich the texture parts of the RSI step by step. Finally, we design a self-learning upscaling structure to dynamically predict the number and weights of the upsampling filters, which can achieve RSI continuous-scale SR. Compared to the meta-learning based on enhanced deep SR (META-EDSR) method, our experimental results show a nearly 0.2-dB improvement on the metrics of the peak signal-to-noise ratio (PSNR). Ning Ni 0002, Libao Zhang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Remote Sensing Image Generation Based on Attention Mechanism and VAE-MSGAN for ROI ExtractionabstractA variety of deep learning approaches have been applied to region of interest (ROI) extraction, which is a fundamental task in the field of remote sensing image (RSI) processing. However, the unbalanced distribution of positive and negative samples in most RSIs greatly restricts the performance of these deep learning-based methods. In this study, a data augmentation method based on variational autoencoder-multiscale generative adversarial network (VAE-MSGAN) with spatial and channelwise attention (SCA) is proposed to balance the sample distribution and improve the subsequent ROI extraction results. First, we combine the original multispectral information with handcrafted texture features to make full use of the low-level visual features of RSIs. We then design a VAE-MSGAN to generate realistic RSIs with high quality and diversity. Specifically, in the generator construct, SCA blocks are introduced to adaptively recalibrate the varying importance of different channels and spatial regions. We also build a multiscale discriminator architecture to improve the visual quality of the generated samples. Finally, we compare the ROI extraction results before and after the augmentation. Our experimental results demonstrate that the proposed method can not only improve the performance of ROI extraction but also be superior to other classical generative methods. Libao Zhang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | UAV Remote Sensing Image Dehazing Based on Double-Scale Transmission Optimization StrategyabstractCurrent dehazing methods for unmanned aerial vehicle (UAV) remote sensing images often have texture detail loss and color distortion problems, especially in highlighted regions. This is mainly due to the rich texture and low intensity of UAV remote sensing images being ignored, which results in incorrect transmission estimation. In this paper, we propose a UAV remote sensing image dehazing method based on double-scale transmission optimization strategy. First, we propose a double-scale optimization strategy to estimate the transmission map with more accurate texture details and color preservation, especially in highlighted regions of hazy UAV images that are most severely distorted. Second, a UAV-adaptive haze-line prior algorithm is proposed to address the large scene depth and low intensity of UAV remote sensing images. Finally, we introduce a luminance-weighted frequency domain saliency model to avoid texture detail loss and color distortions for better transmission optimization, especially in highlighted regions. Compared with state-of-the-art methods, our method shows better detail performance and visual effects, especially for UAV images with highlighted regions. Kemeng Zhang, Sijia Ma, Ruohui Zheng, Libao Zhang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Target Detection Based on Edge-Aware and Cross-Coupling Attention for SAR ImagesabstractDue to the existence of speckle noise, background clutter, backscattering points, and geometric distortion of some targets in synthetic aperture radar (SAR) images, extracting multiscale and multilocation targets accurately is still a great challenge. To tackle these problems, a novel target detection method based on edge-aware and cross-coupling attention for SAR images is proposed in this letter. By enhancing the dependencies between targets in different locations, bridging the gap between different feature maps, and assisting the targets’ detection through cross-coupling with the edge-aware network, the performance of detecting multiscale targets in complex SAR images can be improved significantly. Specifically, residual spatial pyramid pooling (RSPP) and mixed pooling module (MPM)-based convolution block attention module (MCBAM) are combined in the decoding part to promote coupling between networks. Besides, the semi-dense connection is adopted in the encoding part based on residual convolution block (RCB), which can improve the ability of multiscale feature extraction and promote the acquirement of high-resolution features with strong semantic information. Experiments are conducted on the SAR oil tank dataset (OTD) and SAR residential area dataset (RAD). We compare our model with a traditional method and CNN-based algorithms. The experimental results verify that our model outperforms the competing models in both pixel level and geometric segmentation accuracy. Libao Zhang, Wanning Zhu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Pan-Sharpening Based on Joint Visual Saliency Analysis and Parallel Bidirectional NetworkabstractIn remote sensing (RS) images, the demands for spectral and spatial quality of different regions are different, which means the unified fusion strategy on the whole image is not suitable for pan-sharpening task. Saliency, derived from visual attention mechanism, provides an effective way to satisfy these demands. Inspired by this, we propose a novel pan-sharpening method based on joint visual saliency analysis and parallel bidirectional network (JSPBN). Firstly, considering the complex scenes and uneven distribution of targets in RS images, we develop a Bayesian optimization based joint visual saliency analysis (B-JVSA) method that integrates prior saliency based on global color contrast with likelihood saliency based on joint co-occurrence histogram, which can highlight common salient regions while suppressing individual ones and irrelevant background by exploring the correlation among multiple RS images. Secondly, we construct a parallel bidirectional feature pyramid (PBFP) network to obtain coarse fusion features, fully considering individual characteristics of panchromatic images and multispectral images. Finally, we design a saliency-aware layer (SAL) according to B-JVSA to further refine the fusion effect in salient regions and non-salient regions. With the help of SAL, diverse strategies for certain regions are learned through two independent residual dense networks and thereby generating accurate fusion results. Experimental results show that our proposal performs better than the competing methods in both spatial quality enhancement and spectral fidelity preservation. Wanning Zhu, Libao Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Weakly Supervised Region of Interest Extraction Based on Uncertainty-Aware Self-Refinement Learning for Remote Sensing ImagesabstractRegion of interest (ROI) extraction plays a significant role in the field of remote sensing image (RSI) processing. Recently, weakly supervised ROI extraction methods have attracted considerable attention due to low labeling cost. Most of them follow the pipeline of first generating pseudo labels and then using the pseudo labels to train a segmentation model. However, there remain problems to be solved: 1) the unbalanced distribution of foreground and background samples in the RSI dataset influences the network performance; 2) the pseudo labels mainly cover the most discriminative part of object regions that are incomplete; and 3) training with pseudo labels inevitably causes noise issues that degrade the model performance. To solve these issues, we propose a weakly supervised uncertainty-aware self-refinement learning (UASRL) method, where the initial unbalanced image-level labels are progressively refined to high-quality pixel-level annotations. In the proposed UASRL, we first present a deep generative model combined with self-attention modules to improve the unbalanced distribution in the weakly labeled dataset. Then, we design a confidence-weighted complementary erasing-based weakly supervised method to generate pseudo labels with high integrity. Finally, for training with noisy pseudo labels, we develop an uncertainty-aware joint optimization (UAJO) training strategy to reduce the negative effect caused by noisy labels and further refine pixelwise labels in a coarse-to-accurate manner, which in turn jointly promotes the model’s performance. Extensive experiments on three types of RSI datasets reveal that our proposed method is superior to other competing methods and shows a preferable tradeoff between annotation cost and detection performance. Libao Zhang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Remote Sensing Image Super-Resolution via Saliency-Guided Feedback GANsabstractIn remote sensing images (RSIs), the visual characteristics of different regions are versatile, which poses a considerable challenge to single image super-resolution (SISR). Most existing SISR methods for RSIs ignore the diverse reconstruction needs of different regions and thus face a serious contradiction between high perception quality and less spatial distortion. The mean square error (MSE) optimization-based methods produce results of unsatisfactory visual quality, while generative adversarial networks (GANs) can produce photo-realistic but severely distorted results caused by pseudotextures. In addition, increasingly deeper networks, although providing powerful feature representations, also face problems of overfitting and occupying too much storage space. In this article, we propose a new saliency-guided feedback GAN (SG-FBGAN) to address these problems. The proposed SG-FBGAN applies different reconstruction principles for areas with varying levels of saliency and uses feedback (FB) connections to improve the expressivity of the network while reducing parameters. First, we propose a saliency-guided FB generator with our carefully designed paired-feedback block (PFBB). The PFBB uses two branches, a salient and a nonsalient branch, to handle the FB information and generate powerful high-level representations for salient and nonsalient areas, respectively. Then, we measure the visual perception quality of salient areas, nonsalient areas, and the global image with a saliency-guided multidiscriminator, which can dramatically eliminate pseudotextures. Finally, we introduce a curriculum learning strategy to enable the proposed SG-FBGAN to handle complex degradation models. Comprehensive evaluations and ablation studies validate the effectiveness of our proposal. Libao Zhang, Jie Ma 0004 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Cosaliency Detection and Region-of-Interest Extraction via Manifold Ranking and MRF in Remote Sensing ImagesabstractSaliency-based region-of-interest (ROI) extraction is significant for the interpretation of remote sensing images (RSIs). Recently, cosaliency detection has shown its superiority of better extraction of common ROIs by using both intraimage and interimage cues. However, most existing methods still suffer from the complex backgrounds of RSIs, resulting in incomplete ROI extraction, many false positives, and blurred boundaries. In this article, we propose a cosaliency detection framework via manifold ranking and the Markov random field (MRF) for RSIs to address these problems. First, we design a two-stage manifold ranking schema for converting single-image saliency maps (SISMs) to multi-image saliency maps (MISMs). This step takes full advantage of the correlation between images to improve the integrity of ROIs and reduce false positives. Second, we locally fuse saliency proposals by minimizing the energy function in an MRF. The design of the energy function comprehensively considers the global and local performance of saliency proposals to assign appropriate fusion weights. Finally, we generate the ROI masks by thresholding the cosaliency maps. Our approach is evaluated on four RSI datasets and compared to the state-of-the-art methods. Experimental results demonstrate the effectiveness of our model in both cosaliency detection and ROI extraction. Libao Zhang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Dense Haze Removal Based on Dynamic Collaborative Inference Learning for Remote Sensing ImagesabstractHaze in remote sensing images (RSIs) usually causes serious radiance distortion and image quality degeneration, resulting in difficult remote sensing inversion and interpretation. Under the condition of dense haze, existing dehazing methods still experience problems to be solved: 1) the texture details and spectral characteristics in RSIs cannot be restored well; 2) small-scale objects, such as cars and ships, which often consist of only a few pixels in RSIs, cannot be effectively highlighted in dehazed results. To solve these issues, we propose a novel dynamic collaborative inference learning (DCIL) framework that can significantly restore real surface information from dense hazy RSIs. First, we design a dynamic mutual enhancement (DME) mechanism to reinforce the low-level texture features by integrating primary information and semantic information at different levels. Second, we propose a spectrum-aware aggregation (SAA) strategy to mine the spectrum features among multiscale restored results, which can fully capture spectral characteristics. Third, we build a collaborative criterion by constructing a Siamese network structure in the training stage to improve the robustness and generalization performance of DCIL considering the diversity of the scale range and view change of RSIs. Finally, we propose a phased learning strategy to deduce the implicit haze-relevant features by gradually increasing the concentration of haze which can effectively address small-scale objects obscured by dense haze. To this end, we develop two synthetic remote sensing dehazing datasets to train our model, which can also alleviate the dilemma of hazy RSI datasets shortages. Experimental results on both synthetic datasets and real remote sensing hazy images demonstrate that the proposed DCIL can attain significant progress compared to competing methods. The two synthetic hazy datasets are available at https://github.com/Shan-rs/DCI-Net. Libao Zhang, Shan Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Image Generation Based on Texture Guided VAE-AGAN for Regions of Interest Detection in Remote Sensing ImagesabstractDeep learning has shown great strength in regions of interest (ROIs) detection for remote sensing images (RSIs). However, for most of RSIs, the unbalanced distribution of positive and negative samples greatly limits the performance of the deep learning-based methods. To cope with this issue, we propose a novel method based on texture guided variational autoencoder-attention wise generative adversarial network (VAE-AGAN) to augment the training data for ROI detection. First, to generate realistic texture details of RSIs, we propose a texture guidance block to embed texture prior information into encoder and decoder networks. Second, we introduce the channel and spatial-wise attention layers in the discriminator construct to adaptively recalibrate the varying importance of different channels and spatial regions of input RSIs. Finally, we apply the RSI dataset balanced by our proposal to the weakly supervised ROI detection method. Experimental results demonstrate that the proposal can not only improve the performance of ROI detection, but also outperform other competing augmentation methods. Libao Zhang |
ICASSP | 1 |
| 2021 | Nighttime Haze Removal Using Saliency-Oriented Ambient Light And Transmission EstimationabstractThe ambient light and transmission of light source regions and non-light source regions have different physical properties in nighttime scenes. However, existing nighttime dehazing algorithms ignored this phenomenon and got oversaturated results looked like in the daytime. In this paper, we propose a nighttime dehazing method using fused ambient light and transmission based on saliency detection. We first propose a multi-scale saliency detection algorithm to differentiate light source regions and non-light source regions. Then we estimate ambient light and transmission of the two regions based on different methods to acquire the fused ambient light and fused transmission. Finally, we get the haze-free results based on the nighttime imaging model. Experimental results show that our method outperforms other state-of-the-art methods in color recovery and our haze-free results still look like in the nighttime. Shan Wang 0009, Libao Zhang |
ICIP | 3 |
| 2021 | Afdn: Attention-Based Feedback Dehazing Network For Uav Remote Sensing Image Haze RemovalabstractTo efficiently remove haze in unmanned aerial vehicle (UAV) remote sensing images, a novel attention-based feedback dehazing network (AFDN) is proposed, which is constructed by feedback connections and attention-based feedback blocks (AFBs). It has three major advantages compared with other dehazing algorithms: 1) The feedback connections, which allow network to use previous state to improve current performance, can effectively help the proposed AFDN generate clear remote sensing scenes progressively. 2) The AFBs are specially designed to extract global residual features, in which the dual attention block can usefully reduce redundant information and improve the fitting ability of network. 3) To obtain abundant texture information from UAV remote sensing images and restore real ground surfaces, an energy loss is employed for texture features learning. Experiments on synthetic datasets and real UAV remote sensing images verify the superiority of AFDN over several state-of-the-art methods in terms of qualitative and quantitative analysis. Shan Wang 0009, Libao Zhang |
ICIP | 3 |
| 2021 | Uav Remote Sensing Image Dehazing Based On Saliency Guided Two-Scaletransmission CorrectionabstractCurrent dehazing methods for unmanned aerial vehicle (UAV) remote sensing images often hold problems of texture detail loss in highlight regions and color distortions. This is mainly due to incorrect estimation of the transmission. In this paper, we propose a UAV dehazing method based on saliency guided two-scale transmission correction. Firstly, we propose a dehaze-driven frequency domain saliency model to detect highlight regions of hazy UAV images for better transmission correction. Secondly, we introduce a two-scale correction method to estimate the transmission map with more accurate texture details. We also introduce a suppression parameter to further suppress color distortions and energy over-reduction. Finally, the saliency map is taken as a weight of transmission correction to avoid texture detail loss and color distortions, especially in highlights. Compared with state-of-the-art methods, our method shows better visual effect and detail visibility, especially for UAV images with highlight regions. Kemeng Zhang, Ruohui Zheng, Sijia Ma, Libao Zhang |
ICIP | 4 |
| 2021 | Common Regions of Interest Extraction Based on Saliency Statistic Analysis for Multiple Remote Sensing ImagesabstractVarious landscape characteristics and irregular object boundaries often make object extraction more difficult. Automated analysis of remote sensing (RS) images is challenging and saliency detection is an effective solution. Yet, many traditional algorithms emphasize simply on a single image and would, therefore, neglect the similarity of an image set. In this paper, concerning the relationships among images, a region of interest extraction model based on common features analysis for remote sensing images is proposed. Firstly, multi-image saliency maps, showing the common salient objects, are generated by clustering in RGB and CIELab color spaces. Next, a method, highlighting the salient region, is based on global and local saliency statistics analysis. Finally, regions of interest are segmented from original images according to saliency maps which have been made boundaries holding by superpixels. Experimental evaluation shows that compared with six existing models, we get more accurate saliency maps. Xinran Lyu, Wanning Zhu, Libao Zhang |
IGARSS | 4 |
| 2021 | Region of Interest Extraction Based on Unsupervised Cross-Domain Adaptation for Remote Sensing ImagesabstractExtracting region of interest (ROI) plays an important role in many computer vision tasks. Recently, deep methods have shown excellent performance, however, when it comes to remote sensing image (RSI) domain, which lacks pixel-level annotations, training often leads to under-fitting and low-accuracy. In this paper, we propose a novel ROI extraction model based on unsupervised cross-domain adaptation for RSIs. Firstly, we pretrain the network, RS- RoINet, by large-scale natural datasets to learn general features. Through top-down propagation mechanism, we combine global and local information to generate the accurate edge of extraction maps. Then, we introduce domain adaptation module to reduce the difference between natural domain and RSI domain. Data from both domains is transferred into Reproducing Kernel Hilbert Space to measure the domain distribution distance. Finally, the model is adaptive for RSIs and extracts ROI more accurately. Compared with recent fully-supervised state-of-the-arts, our unsupervised method shows outstanding performance. Sijia Ma, Wanning Zhu, Libao Zhang |
IGARSS | 3 |
| 2021 | Single image dehazing based on bright channel prior model and saliency analysis strategyabstractAbstract Haze is a common atmospheric phenomenon that causes poor visibility in outdoor images, which greatly limits image application in later stages. Therefore, haze removal has become the first and most indispensable step when dealing with degraded images. In this paper, we propose a novel bright channel prior (BCP) model and a saliency analysis strategy for haze removal. First, we obtain a more robust and accurate atmospheric light by a superpixel‐based dark channel method. Second, we utilize the dark channel prior (DCP) to handle dark regions in hazy images. However, the DCP often mistakes white regions for opaque haze and thus causes serious colour distortion and halo effects. To solve this problem, a new BCP is proposed to accurately estimate the transmission of bright regions in hazy images. Third, we fuse the DCP and BCP using a multiscale fusion strategy with Laplacian pyramid representation to gain the correct transmission information for both bright and dark regions. Finally, a novel saliency analysis strategy for transmission refinement is proposed, so that the texture details can remain present to the greatest extent in the restored images. The experimental results illustrate that our proposed method performs well in restoring images containing bright objects. Libao Zhang, Shan Wang 0009 |
IET Image Process. | 1 |
| 2021 | Single image haze removal via attention-based transmission estimation and classification fusion network
Shan Wang 0009, Libao Zhang |
Neurocomputing | 2 |
| 2021 | Pan-Sharpening Based on Background Prior Saliency and Joint Sparse Detail ExtractionabstractFor remote sensing images, the spectral and spatial quality requirements vary across regions. Foreground objects need high spatial quality, while background regions require high spectral fidelity for the successive processing. Saliency analysis is an effective tool for distinguishing and achieving these requirements. Thus, we propose a pan-sharpening method based on background prior saliency and joint sparse detail extraction for remote sensing images. First, aimed at the characteristics of remote sensing image fusion, a saliency analysis method based on the foreground distribution and background prior is proposed to produce a regulation factor, which can reflect the different spatial and spectral information requirements of the foreground and background. Then, we propose a detail extraction and fusion method based on the guided filter and sparse representation. We extract spatial details not only from panchromatic (PAN) images but also from multispectral (MS) images and fuse them by maximizing the sparse coefficient strategy to reduce instabilities and dissimilarities. Finally, the regulation factor is used to regulate the detail injection in the pan-sharpening process. Our method can satisfy the various spatial and spectral resolution requirements for different regions more accurately. Compared to other state-of-art methods, both the visual and quantitative results reveal that our method has a better performance at improving the spatial quality and preserving the spectral fidelity. Libao Zhang, Yang Sun 0007 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2021 | Salient Object Detection Based on Progressively Supervised Learning for Remote Sensing ImagesabstractSalient object detection (SOD) is a crucial task in the field of remote sensing image (RSI) processing. Weakly supervised SOD methods, which generate saliency maps by classification convolutional neural networks (CNNs), considerably reduce labor costs. However, due to the complexity of remote sensing scenes, concerns remain about weakly supervised SOD for RSIs: 1) since the pooling operations are applied in the classification CNNs, the boundary maintenance of weakly supervised methods is unsatisfactory and 2) several sophisticated postprocessing procedures are used in previous weakly supervised methods, which are inevitably time-consuming. To solve these problems, we combine the benefits of weakly and fully supervised learning and propose a new SOD method named progressively supervised learning (PSL) for RSIs. The proposed method realizes end-to-end SOD with a lightweight model under imagewise annotations. First, to reduce the demands on large-scale pixelwise annotations, we propose a pseudo-label generation method based on a classification network and gradient-weighted class activation mapping (Grad-CAM) to compute pseudo saliency maps (PSMs) for training samples and auxiliary images in a weakly supervised manner. Then, to improve the computational efficiency, we construct a feedback saliency analysis network (FSAN), where the generated PSMs are regarded as pixelwise labels. Finally, inspired by curriculum learning, we design a new denoising loss function to further reduce the effect brought by missing judgment in PSMs and enhance the detection accuracy. Comprehensive evaluations with two remote sensing data sets and a comparison with 11 methods validate the superiority of the proposed PSL model. Libao Zhang, Jie Ma 0004 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | SC-PNN: Saliency Cascade Convolutional Neural Network for PansharpeningabstractIn many remote sensing tasks, different types of regions or targets differ in requirements for spectral and spatial quality. The discrepancy reveals that a uniform pansharpening strategy applying to the entire image may not fulfill the varying demands of different regions appropriately. From this aspect, we resort to saliency analysis to distinguish regions with different spatial and spectral requirements and then propose a new saliency cascade convolutional neural network for pansharpening (SC-PNN). SC-PNN is composed of two parts: a dilated deformable convolutional network (DDCN) for saliency analysis and a saliency cascade residual dense network (SC-RDN) for pansharpening. DDCN is a fully convolutional network based on hybrid dilated convolution and deformable convolution, aiming to separate salient regions, such as residential areas from nonsalient areas, including mountains and vegetation areas, with well-defined boundaries and integrity. In the fusion process, SC-RDN is specially designed with the help of saliency analysis. We first construct a deep regression network to estimate a primarily sharpened image and subsequently leverage the saliency map produced by DDCN to develop a saliency enhancement module. In this module, the quality of salient and nonsalient areas is further improved by two independent deep residual dense networks. Thus, a precise fused image can be predicted. Experiments on SPOT5, GeoEye-1, and WorldView-3 data sets reveal that, compared to state-of-the-art pansharpening methods, our proposal has a superior ability to improve the spatial quality and preserve spectral information. The effectiveness of the saliency enhancement module is also validated in the experiment. Libao Zhang, Jue Zhang 0001, Jie Ma 0004, Xiuping Jia |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | A Novel Saliency-Driven Oil Tank Detection Method for Synthetic Aperture Radar ImagesabstractSynthetic Aperture Radar (SAR) imaging system plays an important role in earth observation research. This leads to the significance of target detection in SAR image. In this paper, we propose a novel saliency-driven oil tank detection method (SDD) for SAR images. First, we use the enhanced directional smoothing (EDS) to remove speckle noise from SAR images; in the step of saliency analysis, the integer wavelet transforms (IWT) and the DoG filter are used to obtain orientation and intensity features, respectively. Then, the orientation feature map and the intensity feature map resulting from these two kinds of features are utilized to compute the final saliency map; after segmenting the saliency map, the obtained connected domain guide the Active Contour Model (ACM) to acquire accurate contours of tops of oil tanks, and the bottoms of the oil tanks can be detected by the strong scattering points around the tops. Experimental results show that the proposed model outperforms the classical/state-of-the-art models in maintaining complete targets and accurate boundaries. Libao Zhang, Congyang Liu |
ICASSP | 1 |
| 2020 | SDTCN: Similarity Driven Transmission Computing Network for Image DehazingabstractTransmission similarity is an important feature which can greatly increase the capability of convolutional neural network (CNN) to fit transmission map. However, it is not sufficiently utilized in existing algorithms. In this paper, we propose a novel light-weight similarity driven transmission computing network called SDTCN that is guided by the attributes of transmission similarity. First, we adopt a non-data-driven image segmentation method to acquire the transmission similarity. Compared with CNN based segmentation approaches, our method can not only greatly save computing resources, but also separate the objects and background precisely. Second, a full convolutional network is introduced to reduce blocky effects in SDTCN. Finally, unlike previous "first airlight then transmission" mode, a dependable airlight estimation approach is designed drawing on the transmission map generated by SDTCN, which can improve the accuracy of airlight effectively. Extensive experiments demonstrate that the proposed algorithm outperforms the state-of-the-art methods on synthetic and real-world images. Libao Zhang, Shan Wang 0009 |
ICASSP | 1 |
| 2020 | SD-FB-GAN: Saliency-Driven Feedback Gan for Remote Sensing Image Super-Resolution ReconstructionabstractThe visual characteristics of different regions in remote sensing images are significantly versatile, which poses a huge challenge to single image super-resolution. Although generative adversarial network (GAN) has shown great potential in generating photo-realistic results, it provides unsatisfactory performance in objective metrics owning to pseudo textures brought by adversarial learning. In this paper, we propose a new saliency-driven feedback GAN to cope with these problems. We design a saliency-driven feedback generator based on paired-feedback blocks (PFBBs) and recurrent structure to provide strong reconstruction ability. In the PFBB, the saliency map serves as an indicator to reflect the texture complexity, so different reconstruction principles can be applied to restore areas with varying levels of saliency. Besides, we propose to measure the visual quality of salient areas, non-salient areas, and the whole image with multi-discriminators, which can dramatically eliminate pseudo textures. Comprehensive evaluations and ablation studies validate the superiority of our proposal. Jie Ma 0004, Jue Zhang 0001, Libao Zhang |
ICIP | 4 |
| 2020 | Region Of Interest Extraction Based On Co-Saliency Analysis And Feedback Strategy For Remote Sensing ImagesabstractSaliency analysis has been revealed an effective method to extract the region of interest (ROI) in remote sensing images. However, most existing saliency detection methods mainly focus on extracting the ROIs from a single image, which usually are not able to generate satisfactory results because of the background interference of remote sensing images. The employment of co-saliency detection which focuses on detecting common salient objects in a set of images can provide an effective solution to this issue. In this paper, we propose a novel ROI extraction model based on co-saliency analysis and feedback strategy for remote sensing images. We combine the bottom-up measures including consistency, central prior and contrast on spectral and texture feature of multiple images together, with subsequent adjusting operation by feedback strategy to enhance the common salient objects. Experiment results reveal that our model outperforms seven state-of-the-art models. Libao Zhang |
ICIP | 1 |
| 2020 | Joint-Distribution And Gain Rate Based Saliency Model For Circular Tank Detection In Remote Sensing ImagesabstractOil tank detection plays an important role in object detection for remote sensing images. While the existence of the complex cases affects the detection accuracy, this paper proposes a joint-distribution and gain rate based saliency analysis model for circular tank detection. First, the joint-feature residual is utilized to extract common parts among the selected feature maps for intensity analysis. Besides, the local gradient specificity and local flatness descriptor are introduced to assess the texture characters. Second, the introduced feature vector is used for the clustering of the input series, and the joint-distribution is utilized to label a coarse binary mask. Third, the coarse labeled mask is used to assist to calculate the gain rate for different elements of the feature vector and generate the corresponding weights to get the saliency map. Finally, the desired targets are extracted according to the local salient parts in the saliency map. Experiments are conducted on two aspects including the qualitative evaluation and the quantitative evaluation. The assessment of pixel level and geometrical segmentation shows the superiority of the proposed method compared with the listed competing algorithms. Libao Zhang, Congyang Liu |
ICIP | 1 |
| 2020 | Pan-Sharpening Based On Joint Saliency Detection For Multiple Remote Sensing ImagesabstractRequirements of spectral and spatial quality differ from region to region in remote sensing images, which is a significant challenge for pan-sharpening. Joint saliency analysis not only fulfills these demands, but also ensures the consistency by considering the mutual information of multiple images. Thus, we propose a pan-sharpening method based on joint saliency analysis and improved intensity- hue-saturation (IHS) for multiple remote sensing images. Firstly, we introduce an improved IHS method to obtain an accurate estimation of the intensity component. Then, we design a joint saliency analysis method based on global contrast calculation and intensity feature extraction, which is subsequently compensated by texture features to generate adaptive injection gains. Finally, we use the injection gains to inject the detail into the multispectral (MS) image. Experimental results demonstrate that our method has better performance in guaranteeing consistency in multiple images, improving spatial quality and preserving spectral fidelity. Libao Zhang, Wanning Zhu, Yang Sun 0007 |
ICIP | 1 |
| 2020 | Saliency-Driven Target Detection Based on Common Visual Feature Clustering for Multiple Sar ImagesabstractSaliency detection is a newly emerging tool to extract target in image processing. However, due to the loss of color in synthetic aperture radar (SAR) images, the detection result using the traditional saliency analysis is not satisfying. Therefore, a new saliency-driven target detection model based on common visual feature clustering is introduced for multiple SAR images. Firstly, Markov Random Field is applied to extract intra-image saliency map. Secondly, intensity, texture and curve features are extracted from multiple SAR images as common visual features, which can effectively compensate for the lack of color information. And then fuzzy c-means is employed to construct inter-image saliency map. Finally, an effective fusion strategy is used to combine the intra-image saliency map with the inter-image saliency map to obtain the final common saliency map. The experimental results demonstrate that the proposed model outperforms most the state-of-the-art saliency detection models. Shan Wang 0009, Qiaoyue Sun, Sijia Ma, Libao Zhang |
IGARSS | 4 |
| 2020 | Airport Detection Based on Saliency Analysis and Geometric Feature Detection for Remote Sensing ImagesabstractOwing to the complicated background information and large data volume in remote sensing (RS) images, it's difficult to detect airport precisely and efficiently. In this paper, we propose a credible airport detection method based on saliency analysis and geometric feature detection. On the one hand, we use a novel saliency analysis model to measure both global contrast and spatial unity in RS images, by which the most salient region can be extracted accurately and the background can be suppressed preferably. On the other hand, considering the geometric features of the airport, a feature descriptor is conducted to detect proper hole structures and line segments in the saliency map. The experimental results indicate that our proposal outperforms existing saliency analysis models and shows good performance in the detection of the airport. Wanning Zhu, Qijian Zhang, Libao Zhang |
IGARSS | 3 |
| 2020 | SD-GAN: Saliency-Discriminated GAN for Remote Sensing Image SuperresolutionabstractRecently, convolutional neural networks have shown superior performance in single-image superresolution. Although existing mean-square-error-based methods achieve high peak signal-to-noise ratio (PSNR), they tend to generate oversmooth results. Generative adversarial network (GAN)-based methods can provide high-resolution (HR) images with higher perceptual quality, but produce pseudotextures in images, which generally leads to lower PSNR. Besides, different regions in remote sensing images (RSIs) reflect discrepant surface topography and visual characteristics. This means a uniform reconstruction strategy may not be suitable for all targets in RSIs. To solve these problems, we propose a novel saliency-discriminated GAN for RSI superresolution. First, hierarchical weakly supervised saliency analysis is introduced to compute a saliency map, which is subsequently employed to distinguish the diverse demands of regions in the following generator and discriminator part. Different from previous GANs, the proposed residual dense saliency generator takes saliency maps as a supplementary condition in the generator. Simultaneously, combining the characteristic of RSIs, we design a new paired discriminator to enhance the perceptual quality, which measures the distance between generated images and HR images in salient areas and nonsalient areas, respectively. Comprehensive evaluations validate the superiority of the proposed model. Jie Ma 0004, Libao Zhang, Jue Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | Oil Tank Extraction Based on Joint-Spatial Saliency Analysis for Multiple SAR ImagesabstractThe lack of true color and the presence of background clutter reduce the accuracy rate of the saliency analysis for oil tank extraction in the synthetic aperture radar (SAR) images. This letter proposes a specially designed unsupervised method to extract oil tanks using the joint-spatial saliency analysis (JSSA) for multiple SAR images. First, the intrasaliency analysis is established on a saliency driven iterative clustering. This considers the spatial intensity and texture feature within a single image and suppresses most backgrounds. Second, the cospatial residual and the local grayscale statistics are considered independently in the intersaliency analysis. The common salient parts among the input series are extracted and used to overcome the problem of the lack of true color. Third, to make the fusion of the two kinds of saliency maps, the low-rank matrix is introduced. The weights of different maps are calculated and the saliency cues are integrated efficiently. Finally, after the statistics of the highlight points within the candidates, the location of the oil tanks is refined. The experiments show the superiority of the proposed method in both the pixel level and the geometric segmentation. The result of the JSSA model appears to improve the accuracy with fewer missing objects compared with the competing algorithms. Libao Zhang, Congyang Liu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | Hierarchical Weakly Supervised Learning for Residential Area Semantic Segmentation in Remote Sensing ImagesabstractResidential-area segmentation is one of the most fundamental tasks in the field of remote sensing. Recently, fully supervised convolutional neural network (CNN)-based methods have shown superiority in the field of semantic segmentation. However, a serious problem for those CNN-based methods is that pixel-level annotations are expensive and laborious. In this study, a novel hierarchical weakly supervised learning (HWSL) method is proposed to realize pixel-level semantic segmentation in remote sensing images. First, a weakly supervised hierarchical saliency analysis is proposed to capture a sequence of class-specific hierarchical saliency maps by computing the gradient maps with respect to the middle layers of the CNN. Then, superpixels and low-rank matrix recovery are introduced to highlight the common salient areas and fuse class-specific saliency maps with adaptive weights. Finally, a subtraction operation between class-specific saliency maps is conducted to generate hierarchical residual saliency maps and fulfill residential-area segmentation. Comprehensive evaluations with two remote sensing data sets and comparison with seven methods validate the superiority of the proposed HWSL model. Libao Zhang, Jie Ma 0004, Xinran Lv |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | Saliency and Background Prior-Based Residential Area Detection for SAR ImagesabstractDue to the lack of color, the strong speckle noise, and the complex background clutter, target detection in synthetic aperture radar (SAR) images is a challengeable task. A novel saliency and background prior (SBP)-based residential area detection method for SAR images is proposed in this letter. It has three major advantages compared with other methods: 1) in saliency analysis, it deeply exploits the image feature and conceives a new texture representation using the amplitude of partitioning Fourier transform (pFT), which compensates for the lack of color and spectrum information in SAR; 2) it employs the superpixel-level background prior and monitors the average intensity level (AIL) of each superpixel for generating accurate outlines of residential areas; and 3) two regional feature-based indices are presented to select the background clutter, and the results serve as a modification to saliency analysis. Experiments using ALOS PALSAR images show that the proposed method has great priority in both quality and quantity over competing methods by extracting integrated residential areas with clear boundaries. Libao Zhang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | ROI Extraction Based on Multiview Learning and Attention Mechanism for Unbalanced Remote Sensing Data SetabstractWith an increasing number of remote sensing images (RSIs), the automatic region of interest (ROI) extraction based on convolutional neural networks (CNNs) has attracted much interest in recent years. Although fully supervised CNN-based methods have shown superiority in the field of object extraction, pixelwise annotations are expensive and time-consuming. Moreover, due to the unstable distribution of ROIs in complex landscapes, the ratios of foreground and background areas are quite different in RSIs. Training CNNs with such unbalanced data sets lead to over-fitting and low accuracy. In this article, we propose a framework that combines multiview learning and attention mechanism (MLAM) to solve the above mentioned problems. First, we develop a CNN-based weakly supervised method with a weight-balanced loss function to solve the problems caused by an unbalanced data set. It also helps to generate imagewise saliency maps by computing the gradient maps with respect to the input images. Then, we design a multiview strategy to dramatically reduce the missing inspection. Finally, we design a feedback attention mechanism based on the stage neighbor binary pattern to further modify the extraction result. In summary, the proposed framework achieves pixelwise ROI extraction under imagewise annotations through data-driven ML and a knowledge-driven visual attention mechanism. We evaluate the performance of the MLAM framework on two challenging data sets with complex backgrounds. The experimental results indicate that the proposed framework can achieve better performance than other eight ROI extraction models for unbalanced remote sensing data sets. Jie Ma 0004, Libao Zhang, Yang Sun 0007 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Remote-Sensing Image Superresolution Based on Visual Saliency Analysis and Unequal Reconstruction NetworksabstractRemote-sensing images (RSIs) generally have strong spatial characteristics for surface features. Various ground objects, such as residential areas, roads, forests, and rivers, differ substantially. According to this visual attention characteristic, regions with complicated texture features require more realistic details to reflect a better description of the topography, while regions such as farmlands should be smooth and have less noise. However, most existing single-image superresolution (SISR) methods fail to fully utilize these properties and therefore apply a uniform reconstruction strategy to the whole image. In this article, we propose a novel saliency-driven unequal single-image reconstruction network in which the demands of various regions in the superresolution (SR) process are distinguished by saliency maps. First, we design a new gradient-based saliency analysis method to produce more accurate saliency maps with imagewise annotations. The method utilizes the superiority of a multireception field to extract both high-level features and low-level features. Second, we propose a novel saliency-driven gate conditional generative adversarial network, where the saliency map is regarded as a medium during the training procedure of the whole network. The saliency map is regarded as a pixelwise condition in a generator to enhance the training capability of the network. Additionally, we design a new loss function that combines normalized content loss, saliency-driven perceptual loss, and gate-control adversarial loss to further refine details of texture-complex areas for RSIs. We evaluate the performance of our algorithm and compare it with many other state-of-the-art SR methods using a remote-sensing data set. The experimental results show that our approach achieves the optimal outcome in salient areas. Our method attains the best effect on global quality and visual performance. Libao Zhang, Jie Ma 0004, Jue Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Proper Guidance Image Generation Based on Saliency Factor for Better Transmission Refinement in Image Dehazing
Libao Zhang |
ICASSP | 1 |
| 2019 | Oil Tank Detection Using Co-Spatial Residual and Local Gradation Statistic in SAR ImagesabstractThe advance in synthetic aperture radar (SAR) technology reduces the difficulty of data acquisition, along with greatly increased computational complexity. This paper cares about this and aims to accomplish oil tank detection in SAR images from the perspective of computer vision. The whole model includes three main steps. The first is a saliency-driven clustering to accomplish the single image saliency analysis according to the intensity specificity and texture distribution. The second step introduces a common visual saliency analysis based on the co-spatial residual and local gradation statistic to extract the common visual salient parts within the input series. The final step considers the distribution of adjacent highlight points obtained from the saliency analysis to accomplish the location of tanks. Three competing methods are established in the experimental part. The evaluation in pixel level and geometric segmentation accuracy both verify the superiority of the proposed model in target detection and interferences exclusion. Libao Zhang, Congyang Liu |
ICIP | 1 |
| 2019 | Oil Tank Detection Based on Linear Clustering Saliency Analysis for Synthetic Aperture Radar ImagesabstractAs the significant position of oil tanks in oil storage, oil tank detection plays an important role in auto target detection for SAR images. This paper presents a linear clustering saliency analysis based detection model. Firstly, a linear iterative clustering of which the feature vector consists of three texture features and 2-D coordinates is used to over segment the input image. At the same time, multi intensity saliency maps constrain the shape of the superpixel using an adaptive balance weight. Secondly, feature vectors generated from the cluster centers are scattered as far as possible via Principal Component Analysis and sent to the MeanShift model to coarsely locate the candidate area. Finally, strong scattered points on the roof of tanks are utilized to locate the top of the targets. The whole method is evaluated in two aspects: the evaluation of saliency analysis and the accuracy rate of the top location. Experiment shows the efficiency and superiority of our algorithm with fewer interferences and more accurate location of oil tanks. Libao Zhang, Congyang Liu |
ICIP | 1 |
| 2019 | A New Pansharpening Method Using Objectness Based Saliency Analysis and Saliency Guided Deep Residual NetworkabstractPansharpening is a fundamental and crucial task in the remote sensing community. For remote sensing images, there is a significant difference in demands for spatial and spectral resolution in different regions. From this perspective, we propose a new pansharpening method using objectness based saliency analysis and saliency guided deep residual network to boost the fusion accuracy. We first develop an objectness based saliency analysis by incorporating texture feature and objectness measurements to estimate saliency values in images and thereby help discriminate different demands for spatial improvement and spectral preservation. Inspired by the impressive performance of deep learning, we subsequently construct a saliency guided deep residual network to implement pansharpening. In addition, in order to produce images with subtler details, we design a new loss function, the normalized mean square error, particularly for the pansharpening task. Experiments support the superiority of our proposal over six competing methods. Libao Zhang, Jue Zhang 0001, Xinran Lyu, Jie Ma 0004 |
ICIP | 1 |
| 2019 | Co-Feature and Shape Prior Based Saliency Analysis for Oil Tank Detection in Remote Sensing ImagesabstractThe great increase in the resolution of remote sensing images comes with the problem of high computational complexity. Besides, the irregular circular shape of the tanks and the difficulty to establish the dataset make the detection task a big challenge. This paper proposes a co-feature and shape prior based saliency analysis model to extract the candidates. Two kinds of saliency maps are considered including the co-feature spatial residual map which is extracted as the common parts of different feature maps and the shape saliency map which provides the complementary information with the shape prior knowledge of circular targets. Besides, the curvature driven active contour model is used to insure the precise of the extracted contour. The quantitative evaluation of the experiments is conducted considering the detection precision and the geometric error. Results show that the proposed method is superior both in the pixel level performance and the geometric shape. Congyang Liu, Libao Zhang |
IGARSS | 2 |
| 2019 | Target Detection Based on Statistical Saliency Analysis and Geodesic Active Contour Model for Sar ImageryabstractSaliency analysis is a hot topic in target detection for synthetic aperture radar (SAR) image. In this paper, we come up with a novel target detection model based on statistical saliency analysis and saliency-oriented geodesic active contour model for SAR imagery. Firstly, the contrast and homogeneity features are extracted from the gray-level co-occurrence matrix to make up the texture saliency map. Then we obtain the prior graph and likelihood map using superpixels segmentation and Otsu, respectively, which comprise the Bayesian saliency map. The two saliency maps subsequently merge into the final statistical saliency map. Finally, we present a saliency-oriented geodesic active contour model, in which a saliency orientation map is embedded into the level set-based energy functional. Based on the final saliency map, the saliency-oriented geodesic active contour model can acquire the specified and accurate target contour. In qualitative and quantitative experiments, the comprehensive performance of the proposed method outperforms the competing models in maintaining complete targets and accurate boundaries. Shan Wang 0009, Libao Zhang |
IGARSS | 3 |
| 2019 | Target heat-map network: An end-to-end deep network for target detection in remote sensing images
Huai Chen, Libao Zhang, Jie Ma 0004, Jue Zhang 0001 |
Neurocomputing | 2 |
| 2019 | Saliency-Driven Oil Tank Detection Based on Multidimensional Feature Vector Clustering for SAR ImagesabstractA novel saliency-driven oil tank detection method based on multidimensional feature vector clustering (MFVC) is proposed in this letter for synthetic aperture radar (SAR) images. There are three major contributions: 1) a specially designed MFVC method, which is suitable for SAR images without true colors, is employed to detect oil tanks roughly by saliency analysis. Five important complementary features, including intensity, texture, structure, and 2-D coordinates, form a 5-D vector, and then an unsupervised strategy is employed for clustering the 5-D feature vectors to acquire the saliency map; 2) For accurate location, the shape limit coefficient is added to the original active contour model to extract contours of top surfaces; and 3) according to the relations of top, bottom, and the brightest arch, bottoms of oil tanks are computed precisely. Experiments are conducted in two aspects: evaluation for saliency analysis, and for bottom location. Results show that the MFVC method outperforms competing methods in maintaining complete oil tanks and accurate boundaries, and removing the background clutter as much as possible. Libao Zhang, Congyang Liu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2018 | Visual Saliency Analysis for Common Region of Interest Detection in Multiple Remote Sensing ImagesabstractSaliency detection is an effective tool to extract regions of interest (ROIs) from remote sensing images. However, some existing saliency detection models focus on extracting ROI from a single image, which cannot accurately detect ROI against complex background interference. In this paper, a novel visual saliency analysis and ROI extraction model is proposed to effectively extract common ROIs from remote sensing images and exclude images without ROIs. Firstly, the single saliency maps are generated by frequency-tuned (FT) method. Secondly, the cluster method based on synthesized features is proposed to group regions with similar feature into a cluster for multiple images. Thirdly, computing the mean of saliency value as the cluster saliency suppresses the saliency value of non-common ROIs. Finally, a ROI extraction method based on the maximum saliency value is proposed to extract ROIs while eliminating the image without ROIs. Experimental results indicate our model outperforms other state-of-the-art saliency detection models, achieving highest ROC and maximal PRF values. Libao Zhang, Qiaoyue Sun, Yang Sun 0007 |
ICIP | 1 |
| 2018 | Integrating Sparse Reconstruction Saliency and Target-Aware Active Contour Model for Airport ExtractionabstractThis paper deals with automatic airport extraction in remote sensing images (RSIs). We present an innovative framework using sparse reconstruction saliency (SRS) and target-aware active contour model (TAACM). We begin with segmenting an image into superpixels and extracting the feature vectors. In feature space, we learn an airport target dictionary and a background dictionary for sparse representation of all image sub-regions. The saliency confidence can be determined by sparse reconstruction error. Based on the saliency map, we apply a novel target-aware active contour model (TAACM) for target contour tracking and provide accurate descriptions about the airport details. Extensive experiments demonstrate that the SRS algorithm outperforms nine competing saliency models in remote sensing scenes. In addition, the proposed airport extraction framework achieves higher detection accuracy compared with three competing methods. Qijian Zhang, Wenqi Shi 0002, Libao Zhang |
ICIP | 3 |
| 2018 | Salient Target Detection Based on the Combination of Super-Pixel and Statistical Saliency Feature Analysis for Remote Sensing ImagesabstractThe saliency analysis has become the important tool to detect the salient targets. However, due to complex target features and abundant background information interference, the traditional models are weak in salient target detection of remote sensing images. In this paper, a novel model based on the combination of super-pixel and statistical saliency feature analysis is proposed. The proposed model consists of three main steps. First, the statistical saliency feature map based on histogram statistical saliency analysis in the Lab color space is introduced. Then, information saliency feature map is obtained based on the combination of super-pixel segmentation and information entropy, and the statistical saliency feature map and the information saliency feature map are fused and enhanced to generate the final saliency map. Finally, the complete and accurate salient targets and regions of interest (ROIs) are obtained based on the improved Otsu segmentation method. Experimental evaluations show that the proposed model outperforms the state-of-the-art salient detection models. Libao Zhang, Yang Sun 0007 |
ICIP | 1 |
| 2018 | Target Detection Based on Saliency Analysis and Contour Extraction for Synthetic Aperture Radar ImagesabstractRecently, saliency analysis is becoming a hotspot in target detection for remote sensing images. In this paper, a target detection model based on saliency analysis and contour extraction for synthetic aperture radar (SAR) images is proposed. First, we use the coherence-enhancing diffusion model to analyze the saliency of the input SAR image. The structure tensor and the diffusion tensor are employed to calculate the saliency map in this process. Then we combine the tensor voting algorithm and the active contour model, where the negative value of the curve saliency value calculated by the tensor vote is used as the external energy of the active contour model. Therefore, the contour of the target can be extracted accurately. Finally, the accurate and complete target region is obtained by the threshold segmentation algorithm. Experimental results show that the proposed model outperforms the relevant classical models. Congyang Liu, Libao Zhang |
IGARSS | 4 |
| 2018 | Airport Detection Based on Superpixel Segmentation and Saliency Analysis for Remote Sensing ImagesabstractTraditional target detection methods are usually based on prior knowledge by template matching and classification. Nowadays, remote sensing images contain richer and richer information. It will cause high computation complexity if we still apply traditional target detection methods to remote sensing images. This paper proposes an airport detection model based on superpixel segmentation and saliency analysis. First, the input image is segmented into superpixels. Then saliency analysis is performed by calculating differences between superpixels and corresponding weights in R, G and B color channels to get the saliency map. Finally we utilize the limitation in the ratio of perimeter and area and morphology operation to eliminate the interference. Experiments compare the proposed model with three saliency analysis models qualitatively and quantitatively. Results show that the proposed model is better than the three comparative models in keeping clear boundaries, eliminating interference and maintaining intact targets. Libao Zhang |
IGARSS | 2 |
| 2018 | Saliency-based dark channel prior model for single image haze removalabstractImages degraded by haze usually have low contrast and fide colours, and thus have bad effects on applications such as object tracking, face recognition, and intelligent surveillance. So the purpose of dehazing is to recover the image contrast without colour distortion. The dark channel prior (DCP) is widely used in the field of haze removal because of its simplicity and effectiveness. However, when faced with bright white objects, DCP overestimates the haze from its true value and thus causes colour distortion. In this study, the authors propose a dehazing model combining saliency detection with DCP to obtain recovered images with little colour distortion. There are three main contributions. First, they introduce a novel saliency detection method, focusing on superpixel intensity contrast, to extract bright white objects in the hazy image. Those objects are not used to estimate the atmospheric light and transmission in the dark channel image. Second, a self‐adaptive upper bound is set for the scene radiance to prevent some regions being too bright. Third, they propose a quantitative indicator, colour variance distance, to evaluate the colour restoration. Experimental results show that their proposed model generates less colour distortion and has better comprehensive performance than competing models. Libao Zhang |
IET Image Process. | 1 |
| 2018 | Saliency detection and region of interest extraction based on multi-image common saliency analysis in satellite images
Libao Zhang, Qiaoyue Sun |
Neurocomputing | 1 |
| 2018 | A Novel Saliency-Oriented Superresolution Method for Optical Remote Sensing ImagesabstractImage superresolution (SR) techniques have widely been used to satisfy the increasing resolution demands of many advanced applications in the remote sensing domain. However, most of the existing SR models are directly operated on the whole image without considering the deviations and different objective requirements of different regions in optical remote sensing images, which are not sensible enough. To fill this gap, we propose a novel saliency-oriented adaptive SR strategy motived by the visual attention mechanism. The key idea of this letter is employing diverse treatments on different regions according to their unique requirements. For instance, the reconstruction quality of regions of interest (ROIs) should be as fine as possible. First, we employ a saliency detection strategy based on the edge-enhancement discrete wavelet transform to generate a saliency map, which clearly demonstrates the distribution of ROIs. Then, with regard to these areas, a new SR strategy is applied to get a better performance, where the training process is upgraded with the feature optimization process. In addition, the rest regions also receive the standard A+ to improve the quality there. Finally, all restored high-resolution (HR) patches are fused together as the desired reconstructed HR image. The comparative experiments validate the effectiveness of our scheme. Libao Zhang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2018 | Airport Extraction via Complementary Saliency Analysis and Saliency-Oriented Active Contour ModelabstractAutomatic airport extraction in remote sensing images (RSIs) has been widely applied in military and civil applications. An efficient airport extraction framework for RSIs is constructed in this letter. In the first step, we put forward a two-way complementary saliency analysis (CSA) scheme that combines vision-oriented saliency and knowledge-oriented saliency for the airport position estimation. In the second step, we construct a saliency-oriented active contour model (SOACM) for airport contour tracking, where a saliency orientation term is incorporated into the level-set-based energy functions. Under the guidance of saliency feature representations obtained by CSA, the SOACM can acquire well-defined and highly precise object contours. Experimental results demonstrate that the proposed extraction framework shows good adaptability in remote sensing scenes, and uniformly achieves high detection rate and low false alarm rate. Compared with three state-of-the-art algorithms, our proposal can not only estimate the location of airport targets, but also extract detailed information of the airport contours. Qijian Zhang, Libao Zhang, Wenqi Shi 0002 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | Single image haze removal based on saliency detection and dark channel priorabstractSince more and more outdoor images are often degraded by haze and suffer from bad visibility, which includes low contrast, low resolution, and high luminance, haze removal has become an important task of image restoration in recent years. In this paper, we propose a saliency prior, which introduces human visual attention mechanism into haze removal. The prior reveals the relationship between the saliency analysis and the depth of hazy scenes. In our dehazing method, firstly, saliency map and salient regions are obtained. Then, an accurate airlight and a refined transmission map can be acquired based on saliency prior and dark channel prior. Finally, we can restore the haze-free image successfully using the airlight and the transmission map. Experimental results show that our method performs much better than others in recovering large white areas that are inherently similar to the airlight. Libao Zhang, Chen She |
ICIP | 1 |
| 2017 | A new fusion method for remote sensing images based on salient region extractionabstractThe goal of the remote sensing image fusion is to inject the detail information extracted from panchromatic (PAN) images to multispectral (MS) images with minimized spectral distortion. However, different regions in the image may practically have different demands on the spatial and spectral resolution. In this paper, a new fusion method for remote sensing images based on salient region extraction is proposed. By introducing the hybrid visual saliency analysis, information in the PAN and MS image are automatically partitioned into two categories: salient and non-salient regions. Then, a sub-region fusion strategy is applied to fuse the non-salient and salient regions respectively. For non-salient regions, such as farmland and mountains, the wavelet transform is used in the process of spatial infusion to suppress spectral distortion. As for salient regions like residential areas, the windowed IHS transform is carried out for its merits of effective integration of spatial and spectrum information. Experimental results demonstrate that our proposal achieves a better balance between spatial injection and spectral maintenance in different regions. Libao Zhang, Jue Zhang 0001 |
ICIP | 1 |
| 2017 | Region-of-Interest Coding Based on Saliency Detection and Directional Wavelet for Remote Sensing ImagesabstractWith growing contradiction between the high-speed acquisition of remote sensing data and the low-speed data storage and transmission, the advantages of giving higher priority to a region of interest (ROI) in compression have become prominent. Previous research focused on ROI coding, rather than automatic ROI extraction. However, accurate ROI extraction can significantly improve coding efficiency. In this letter, we propose an automatic ROI extraction based on the improved normal directional lifting wavelet transform (LWT). Then, the compression efficiency is enhanced by a novel tangent directional LWT to reduce the signal energy of high-frequency subbands. Finally, the autogenerated ROIs are encoded by a new multibitplane alternating shift method, which supports not only arbitrarily shaped ROI coding, but also flexible adjustment of compression quality in the ROI and the background. The experimental results demonstrate that our method can effectively highlight the ROIs with well-defined boundaries, meanwhile improving the ROI coding with better visual quality. Libao Zhang, Jie Chen 0059, Bingchang Qiu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | A New Saliency-Driven Fusion Method Based on Complex Wavelet Transform for Remote Sensing ImagesabstractIn remote sensing images, demands for spectral and spatial resolution vary from region to region. Regions with abundant texture and well-defined boundaries (like residential areas and roads) need more spatial details to provide better descriptions of various ground objects while regions such as farmland and mountains are mainly discriminated by spectral characteristic. However, most existing fusion algorithms for remote sensing images execute a unified processing in the whole image, leaving those important needs out of consideration. The employment of diverse fusion strategy for regions with different needs can provide an effective solution to this problem. In this letter, we propose a new saliency-driven fusion method based on complex wavelet transform. First, an adaptive saliency detection method based on clustering and spectral dissimilarity is presented to generate saliency factor for indicating diverse needs of the two kinds of resolutions in regions. Then, we combine nonlinear intensity-hue-saturation transform with multiresolution analysis based on dual-tree complex wavelet transform in order to complement each other's advantages. Finally, saliency factor is employed to control the detail injection in the fusion, helping to satisfy different needs of different regions. Experiments reveal the validity and advantages of our proposal. Libao Zhang, Jue Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2016 | Saliency analysis and region-of-interest extraction for satellite images by biological sparse modelingabstractTraditional models for saliency analysis in satellite images cannot genuinely mimic the selection mechanism of human vision system. Furthermore, feature selection needs variant considering the complexity of data distribution of different satellite images thereby not being one-size-fits-all. Aiming at these problems, we propose a novel model based on sparse representation for saliency analysis with biological plausibility and preferably, our model only needs to decide the number of feature without considering feature complexity and massive parameters tuning in other feature learning algorithms. First, sparse filtering is adopted to learn a sparse dictionary for satellite images. Then, we use Incremental Coding Length (ICL) to measure the saliency contribution of every feature for the final saliency map. The region-of-interest (ROI) can be extracted based on saliency maps by thresholding segmentation. Experimental results show that our model achieves better performance compared with several traditional models for saliency analysis and ROIs extraction in satellite images. Libao Zhang, Jie Chen 0059 |
ICIP | 1 |
| 2016 | Multi-image saliency analysis via histogram and spectral feature clustering for satellite imagesabstractSaliency analysis is an effective method to extract interesting target regions from satellite images. However, when the satellite image contains salient background information, it is difficult to eliminate this information accurately only using single image saliency analysis. In this paper, a novel multiimage saliency analysis (MSA) model based on multiple multispectral images clustering saliency analysis (MMCS) and panchromatic image co-occurrence histogram saliency analysis (PCHS) is proposed. We obtain the MMCS maps based on contrast principle, while effectively depressing the background information that is salient only within its own image. PCHS maps are obtained by co-occurrence histogram of panchromatic images that aims to enhance the saliency of target regions. Finally, multi-image saliency maps are computed by a novel fusion strategy, which can depress the background information and highlight the target regions. Experimental results show that the MSA model outperforms other state-of-art saliency analysis methods. Libao Zhang, Qiaoyue Sun, Jie Chen 0059 |
ICIP | 1 |
| 2016 | Region of interest detection based on salient feature clustering for remote sensing imagesabstractThe region of interests (ROI) detection plays an important role in the remote sensing data processing and analysis. In this paper, a new region of interest detection method based on salient feature clustering for remote sensing images is proposed. Four steps are included in the proposed method. First, the information salient feature maps are constructed by computing the spectrum information and histograms of multispectral images. Second, a clustering strategy based on k-means is presented to generate the common salient feature maps in the CIE Lab color space. Third, the final saliency maps are generated by fusing the information salient feature maps with the common salient feature maps. Finally, we can get the ROIs by segmenting the final saliency map. Experimental results show that compared with five existing models, our model gets more accurate saliency maps without the basis of prior knowledge. Libao Zhang, Xinran Lv, Jie Chen 0059 |
IGARSS | 1 |
| 2016 | Remote sensing image segmentation based on Wilcoxon rank sum test and mean absolute deviationabstractIn this paper, a novel threshold segmentation method for remote sensing images is proposed. The proposed method is based on Wilcoxon rank sum test and mean absolute deviation (MAD) model with color feature and can segment roads and residential areas from vegetation more accurately. Three steps are used to realize the new method. First, we use blue and green color components as paired sample on Wilcoxon rank sum test to partition the vegetation. Second, a road-residential area map is constructed by mean absolute deviation on an improved two dimensional histogram to get the optimal threshold for segmentation. Finally, we fuse vegetation and residential map to get the final segmentation result. Compared with several existing algorithms, the proposed method presents a more accurate segmentation. Libao Zhang, Qiaoyue Sun, Aoxue Li |
IGARSS | 1 |
| 2016 | Saliency analysis and region of interest detection via orientation information and contrast feature in remote sensing imagesabstractSaliency analysis is an important implement for remote sensing image processing. It can effectively solve the contradiction between accuracy and computation complexity when it is applied in the region of interest (ROI) detection and extraction for remote sensing images. In this paper, we propose an efficient ROI detection model for remote sensing images based on low-level contrast feature saliency analysis. For the proposed model, we first perform fast directional integer wavelet transform (FD-IWT) to obtain multi-scale approximate and detail coefficients. Then these multi-scale orientation, local contrast, and global contrast features are exploited to generate the saliency map. Qualitative and quantitative evaluation shows that the proposed model outperforms the other nine state-of-art ROI detection models for that the proposed model can obtain highlighted and integrated ROI with well-defined boundaries, as well as eliminate the shadow interference in remote sensing image. Wen Lv, Libao Zhang |
VCIP | 3 |
| 2016 | Region of interest extraction in remote sensing images by saliency analysis with the normal directional lifting wavelet transform
Libao Zhang, Jie Chen 0059, Bingchang Qiu |
Neurocomputing | 1 |
| 2016 | Region-of-interest extraction based on spectrum saliency analysis and coherence-enhancing diffusion model in remote sensing images
Libao Zhang |
Neurocomputing | 1 |
| 2016 | Global and Local Saliency Analysis for the Extraction of Residential Areas in High-Spatial-Resolution Remote Sensing ImageabstractExtraction of residential areas plays an important role in remote sensing image processing. Extracted results can be applied to various scenarios, including disaster assessment, urban expansion, and environmental change research. Quality residential areas extracted from a remote sensing image must meet three requirements: well-defined boundaries, uniformly highlighted residential area, and no background redundancy in residential areas. Driven by these requirements, this study proposes a global and local saliency analysis model (GLSA) for the extraction of residential areas in high-spatial-resolution remote sensing images. In the proposed model, a global saliency map based on quaternion Fourier transform (QFT) and a global saliency map based on adaptive directional enhancement lifting wavelet transform (ADE-LWT) are generated along with a local saliency map, all of which are fused into a main saliency map based on complementarities. In order to analyze the correlation among spectrums in the remote sensing image, the phase spectrum information of QFT is used on the multispectral images for producing a global saliency map. To acquire the texture and edge features of different scales and orientations, the coefficients acquired by ADE-LWT are used to construct another global saliency map. To discard redundant backgrounds, the amplitude spectrum of the Fourier transform and the spatial relations among patches are introduced into the panchromatic image to generate the local saliency map. Experimental results indicate that the GLSA model can better define the boundaries of residential areas and achieve complete residential areas than current methods. Furthermore, the GLSA model can prevent redundant backgrounds in residential areas and thus acquire more accurate residential areas. Libao Zhang, Aoxue Li, Kaina Yang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Salient region detection in remote sensing images based on color information contentabstractAccurate and anti-noise detection of salient regions is a hotspot of remote sensing image analysis. In this paper, we introduce a new salient region detection model for residential areas in high-spatial-resolution remote sensing images, which is called Color Information Content model (CIC), applying color information content and outputting full resolution saliency maps. First, one-dimensional (1D) histograms of different color channels are constructed based on intensities. Second, the information content of intensities is computed on the 1D histograms and an information mapping is used to construct information maps which reflect information content of each colour channel. Finally, to establish saliency map, intensities of different color channels are fused by saliency scores based on information maps. Experimental results show that compared with existing models, our model not only gets accurate results effectively, but also has good noise immunity. Libao Zhang |
IGARSS | 1 |
| 2015 | Image denoising based on iterative generalized cross-validation and fast translation invariant
Libao Zhang, Jie Chen 0059 |
J. Vis. Commun. Image Represent. | 1 |
| 2015 | Residential area extraction based on saliency analysis for high spatial resolution remote sensing images
Libao Zhang, Jue Zhang 0001, Jie Chen 0059 |
J. Vis. Commun. Image Represent. | 1 |
| 2015 | Remote Sensing Image Segmentation Based on an Improved 2-D Gradient Histogram and MMAD ModelabstractA novel remote sensing image segmentation algorithm based on an improved 2-D gradient histogram and minimum mean absolute deviation (MMAD) model is proposed in this letter. We extract the global features as a 1-D histogram from an improved 2-D gradient histogram by diagonal projection and subsequently use the MMAD model on the 1-D histogram to implement the optimal threshold. Experiments on remote sensing images indicate that the new algorithm provides accurate segmentation results, particularly for images characterized by Laplace distribution histograms. Furthermore, the new algorithm has low time consumption. Libao Zhang, Aoxue Li, Shuaijing Xu, Xuye Yang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2014 | Region-of-Interest Extraction Based on Frequency Domain Analysis and Salient Region Detection for Remote Sensing ImageabstractTraditional approaches for detecting visually salient regions or targets in remote sensing images are inaccurate and prohibitively computationally complex. In this letter, a fast, efficient region-of-interest extraction method based on frequency domain analysis and salient region detection (FDA-SRD) is proposed. First, the HSI transform is used to preprocess the remote sensing image from RGB space to HSI space. Second, a frequency domain analysis strategy based on quaternion Fourier transform was employed to rapidly generate the saliency map. Finally, the salient regions are described by an adaptive threshold segmentation algorithm based on Gaussian Pyramids. Compared with existing models, the new algorithm is computationally more efficient and provides more visually accurate detection results. Libao Zhang, Kaina Yang |
IEEE Geosci. Remote. Sens. Lett. | 1 |