VLDB 2026 Research / reviewers in the wild / expert
Fang Liu 0001
dblp:67/5807-1
· DBLP profile ↗
303ranked-venue papers
9as first author
209since 2021 · last 2026
0000-0002-5669-9354ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 129 · 5 first-author · 80 since 2021Applied, interdisciplinary, general and emerging computing · 104 · 3 first-author · 74 since 2021Graphics, computer vision, multimedia, augmented reality and games · 74 · 65 since 2021Databases, data management, data science and information retrieval · 7 · 1 since 2021Computer networks · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evolving Semantic Propagation for Aerial Semantic 3D Gaussian SplattingabstractSemantic understanding of large-scale aerial scenes represents a critical challenge in 3D computer vision, hindered by the prohibitive cost of dense annotation. This paper introduces EvoPropGS, a novel approach for the semantic segmentation of 3D Gaussian Splatting models that requires only minimal supervision. Our core insight is to leverage the inherent structural repetitions within aerial environments to propagate semantic information from a sparse set of annotations across the entire 3D scene. Our approach constructs a prompt library by pairing SAM-generated mask candidates with DINOv2 feature embeddings from annotated views. For unannotated regions, we generate pseudo-labels by matching region proposals with these featured prompts via cosine similarity. We then formulate optimal prompt selection as a discrete optimization problem solved via evolutionary search, guided by our novel fitness function that evaluates both 3D consistency and 2D semantic coherence. Extensive experiments demonstrate that EvoPropGS achieves accurate segmentation with only 2 percent annotated pixels. Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Licheng Jiao, Puhua Chen, Wenping Ma 0001, Shuyuan Yang 0001 |
AAAI | 4 |
| 2026 | HTTrack: Learning to Perceive Targets via Historical Trajectories in Satellite Video TrackingabstractIn recent years, the rapid progress of deep learning has driven notable advancements in satellite video tracking, a critical task for applications such as environmental monitoring, disaster management, and defense. Despite these strides, existing approaches remain constrained by their inability to handle dynamic challenges, such as target appearance variations, complex motion patterns, and occlusions. Traditional methods often suffer from static template matching or overly complex update mechanisms, compromising their robustness and practicality in real-world scenarios. To address these limitations, we propose a paradigm shift in satellite video tracking by integrating historical trajectory knowledge with visual features. This fusion enhances the tracker's perceptual understanding of targets over time, enabling more adaptive and resilient tracking. By aligning spatial, temporal, and cross-modal information, our approach effectively bridges the gap between fragmented observations and coherent tracking performance, even under challenging conditions like small target detection and cluttered backgrounds. Extensive experiments conducted on multiple satellite video tracking benchmarks demonstrate the superiority of our method, with HTTrack achieving success rates of 51.5% on SV248S, 52.9% on SatSOT, and 32.6% on VISO, significantly outperforming state-of-the-art trackers and marking a step forward in achieving robust, accurate, and scalable satellite video tracking. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
AAAI | 2 |
| 2026 | Semantic Feature Purification for Adversarially-Aware RGB-T TrackingabstractRGB-T tracking is increasingly deployed in safety-critical applications such as autonomous driving, surveillance, and rescue robotics, where tracking reliability is essential under adverse conditions. Although the fusion of RGB and thermal infrared (TIR) modalities offers improved robustness in low-light and occluded scenes, recent findings show that RGB-T trackers remain highly susceptible to subtle input perturbations, human-imperceptible modifications that exploit cross-modal inconsistencies to mislead tracking outputs. In real-world scenarios, such perturbations can arise from sensor spoofing, infrared camouflage, or physical-world attacks, posing serious risks to operational safety. To address this, we propose SFPT, a Semantic Feature Purification framework that enhances RGB-T tracking at the representation level. Rather than filtering corrupted inputs at the pixel level, SFPT introduces task-specific semantic anchors into the feature space to reinforce perturbation-invariant cues. These anchors are derived from descriptive language, interact with visual features to purify representations. To further suppress modality-specific interference, we design an Adaptive Perturbation-Guided Cross-Modal Fusion (APG-CMF) module, which leverages language and visual signals to estimate reliability and dynamically reweight cross-modal features, ensuring robust fusion under perturbation conditions. Extensive experiments under diverse perturbation conditions validate the effectiveness of our approach. Notably, SFPT maintains performance comparable to clean settings even when subjected to perturbations of strength 1/255 and 4/255, demonstrating strong resilience to real-world interference. Jiahao Wang 0002, Fang Liu 0001, Hao Wang 0211, Shuo Li 0010, Puhua Chen |
AAAI | 2 |
| 2026 | WiSLAT: A Simultaneous Device Localization and Target Tracking Method for Wi-Fi Systems
Chunxi Chen, Chao Yu 0007, Fang Liu 0001, Rui Wang 0007 |
SECON | 4 |
| 2026 | Text augmentation for vision: Modality-preference aware few-shot learning
Zehua Hao, Fang Liu 0001, Shuo Li 0010, Yaoyang Du, Jiahao Wang 0002, Hao Wang 0211, Licheng Jiao |
Knowl. Based Syst. | 2 |
| 2026 | Privacy-preserving video anomaly detection via federated learning
Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Yanbiao Ma, Qianyue Bao, Lingling Li 0002, Puhua Chen |
Knowl. Based Syst. | 2 |
| 2026 | Contrastive perception representation learning for image inpainting
Maoguo Gong, Jianzhao Li, Licheng Jiao, Xu Liu 0006, Fang Liu 0001 |
Knowl. Based Syst. | 6 |
| 2026 | Like Human Rethinking: Contour Transformer AutoRegression for Referring Remote Sensing InterpretationabstractReferring remote sensing interpretation holds significant application value in various scenarios such as ecological protection, resource exploration, and emergency management. However, referring remote sensing expression comprehension and segmentation (RRSECS) faces critical challenges, including micro-target localization drift problem caused by insufficient extraction of boundary features in existing paradigms. Moreover, when transferred to remote sensing domains, polygon-based methods encounter issues such as contour-boundary misalignment and multi-task co-optimization conflicts problems. In this paper, we propose SeeFormer, a novel contour autoregressive paradigm specifically designed for RRSECS, which accurately locates and segments micro, irregular targets in remote sensing imagery. We first introduce a brain-inspired feature refocus learning (BIFRL) module that progressively attends to effective object features via a coarse-to-fine scheme, significantly boosting small-object localization and segmentation. Next, we present a language-contour enhancer (LCE) that injects shape-aware contour priors, and a corner-based contour sampler (CBCS) to improve mask-polygon reconstruction fidelity. Finally, we develop an autoregressive dual-decoder paradigm (ARDDP) that preserves sequence consistency while alleviating multi-task optimization conflicts. Extensive experiments on RefDIOR, RRSISD, and OPTRSVG datasets under varying scenarios, scales, and task paradigms demonstrate transformative performance gains: compared to the baseline PolyFormer, our proposed SeeFormer improves oIoU and mIoU by 27.58% and 39.37% for referring image segmentation and by 18.94% and 28.90% for visual grounding on the RefDIOR dataset. Jinming Chai, Licheng Jiao, Xiaoqiang Lu, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Wenping Ma 0001, Weibin Li 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Learning Evolution via Optimization Knowledge AdaptationabstractThe iterative search process of evolutionary algorithms (EAs) encapsulates optimization knowledge within historical populations and fitness evaluations. Effective utilization of this knowledge is crucial for facilitating knowledge transfer and online adaptation. However, current research typically addresses these goals in isolation and faces distinct limitations: evolutionary sequential transfer optimization often suffers from incomplete utilization of prior knowledge, while adaptive strategies, utilizing real-time knowledge, are limited to tailoring specific evolutionary operators. To simultaneously achieve these two capabilities, we introduce the Optimization Knowledge Adaptation Evolutionary Model (OKAEM), a unified learnable evolutionary framework capable of adaptively updating parameters based on available optimization knowledge. By parameterizing evolutionary operators via attention mechanisms, OKAEM enables learnable update rules that facilitate the utilization of optimization knowledge via two phases: pre-training to integrate extensive prior knowledge for efficient transfer, and adaptive optimization to dynamically update parameters based on real-time knowledge. Experimental results confirm that OKAEM significantly outperforms state-of-the-art sequential transfer methods across 12 transfer scenarios via pre-training, and surpasses advanced learnable EAs solely through its self-tuning mechanism in prior-free settings. Beyond demonstrating practical utility in prompt tuning for vision-language models, ablation studies validate the necessity of the learnable components, while visualization analyses reveal the model's capacity to autonomously discover interpretable evolutionary principles. Chao Wang 0099, Lingling Li 0002, Licheng Jiao, Jiaxuan Zhao, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Physics-Informed Matrix Factorization OperatorabstractMatrix factorization is a fundamental characterization model in machine learning and is usually solved using mathematical decomposition reconstruction loss. However, matrix factorization is a data-driven model whose results depend on data quality, making it susceptible to noise. Inspired by physics, the law of conservation of energy is used to introduce physical laws into matrix factorization, which is called Physics-informed Matrix Factorization operator (PiMF). The PiMF operator uses the heat conduction equation to construct the energy objective function for matrix factorization, thereby retaining the mathematical model's decomposition meaning and satisfying the interpretability of physics. The PiMF follows the physical laws, thereby suppressing irregular or sudden noise signals that violate these physical principles. The solutions of the PiMF operator include more comprehensive knowledge of mathematics and physics, which improves the ability to generalize complex data, especially for noisy data. We demonstrate the consistency of the energy objective function and the mathematical model, which verifies the feasibility of matrix factorization using physical energy laws. In addition, the physical interpretability of the PiMF operator is proved from the perspective of energy decline. This study proposes two practical algorithms for PiMF in classification and clustering tasks, enhancing the practicability of matrix factorization by incorporating task-specific prior information constraints. The experimental results of PiMF for classification and clustering demonstrate the advantages of the proposed operator. The importance of physics-informed matrix factorization is verified, especially for noisy data. Chenxi Tian, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Causality-inspired learning semantic segmentation in unseen domain
Pei He, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Ronghua Shang, Yuwei Guo 0001, Puhua Chen, Shuyuan Yang 0001 |
Pattern Recognit. | 5 |
| 2026 | VCGPrompt: Visual Concept Graph-Aware Prompt Learning for Vision-Language Models
Mengjia Wang, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
Pattern Recognit. | 2 |
| 2026 | Vision-by-prompt: Context-aware dual prompts for composed video retrieval
Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
Pattern Recognit. | 2 |
| 2026 | TFBTrack: Target-Aware Foreground-Background Modeling for vision-language tracking
Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
Pattern Recognit. | 2 |
| 2026 | Language-guided modulation-update for semi-supervised semantic segmentation
Libo Yan, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Jiahao Wang 0002, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Xuejian Gou |
Pattern Recognit. | 2 |
| 2026 | ERFC: Energy-Aware Reinforcement Feedback Calibration for Zero-Shot CaptioningabstractZero-shot captioning aims to generate descriptive captions for unseen image and video data by leveraging the potential of visual language models (VLMs) and language models (LMs) without requiring task-specific training. It has emerged as a critical task, but its performance is often hindered by the inherent gap between the training distribution and unseen test data. The fundamental challenge lies in the model’s strong dependence on the marginal distribution of the training data, which leads to biased predictions when handling test samples. To address this issue, we propose an Energy-aware Reinforcement Feedback Calibration (ERFC) framework to calibrate the distribution and predictions of caption models from a novel energy perspective. The calibration process of ERFC is divided into two key components: 1) We first construct an Energy Stabilizer (ES) based on the caption model, where energy is considered a measure of the affinity between the input sample and the model’s learned distribution. ES iteratively adjusts the embedding features of the input sample using Langevin Dynamics, reducing its energy to implicitly align the model’s distribution with the unseen target domain. 2) We deploy a Reinforcement Calibrator (RC) to refine and calibrate the generated captions through a reward-feedback mechanism. RC leverages the expert CLIP model as a reward signal to assess the quality of the generated captions and employs the policy gradient algorithm to reward or penalize the model, thereby improving its performance. By iteratively combining energy-based optimization and reward-driven calibration, ERFC achieves superior zero-shot generalization capabilities, as demonstrated on image benchmarks such as MSCOCO, Flickr30K, and NoCaps, as well as video benchmarks such as MSR-VTT and MSVD. Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | KCI-Net: Knowledge-Based Contourlet Inference Network for Super-ResolutionabstractTextural details are useful for image super-resolution, but massive CNN methods ignored the high-frequency components and generated over-smoothed outputs. The knowledge-based contourlet inference network is proposed in this paper. Different from other CNN-based methods that are directly infer high-resolution (HR) images, our model learns to reconstruct the HR image through the series of corresponding contourlet coefficients. Specifically, first, we consider the low-pass subbands of the contourlet as the corresponding low-resolution (LR) image. Then, feed it to the embedding net with residual blocks to provide adequate information for the contourlet coefficients prediction. Finally, we innovatively convert the estimation of contourlet coefficients into the estimation of the generalized gaussian distribution (GGD) parameters, and design the corresponding loss function to ensure training stability, which explores the smoothness of the contour effectively and guarantees the general structure and details of images. Experiments on four remote sensing datasets, four natural scenes and human-made content datasets, and the outdoor dataset demonstrate the superiority of the proposed model quantitatively and qualitatively. Mengkun Liu, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Shuang Wang 0001, Biao Hou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Knowledge-Aware Evolutionary TransformerabstractWith the Transformer architecture achieving impressive results in the vision domain. It has become a current popular research to explore more potentials of Transformer mixed architectures and explore more suitable mixed combinations. In this paper, Transformer classification network is designed and explored by multi-task architecture search algorithm. A new paradigm for multi-task architecture is designed by combining convolution and Transformer. The designed search architecture can combine the respective advantages of convolution and Transformer and can obtain better performance. At the same time, corresponding knowledge-aware multi-task genetic operators are designed to generate offspring individuals. Inter-task and inter-experience knowledge-aware is utilised to facilitate evolutionary convergence. During the search process, the reference evaluation method is utilised to reduce the redundant computation and time during the search process. In the experimental section, the search results are compared with state-of-the-art architectures and search algorithms. The experimental results confirm the effectiveness and high generalisation of the searched architectures. The ablation experimental part proves the effectiveness of the proposed architectural paradigm, genetic operators and reference evaluation. Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Evol. Comput. | 4 |
| 2026 | RefZVC: Refinable Zero-Shot Video Captioning by Test-Time Reinforcement PolishingabstractRecently, the zero-shot image captioning (zero-shot IC) method based on pre-trained visual language models (VLMs) and large language models (LLMs) has made significant progress. However, how to adapt it to the zero-shot video captioning (zero-shot VC) scenario (without video-text paired supervision) has not been well explored. Inspired by various recent test-time strategies (sacrificing additional test time to improve performance), we try to introduce a new paradigm of Test-time Reinforcement Polishing in zero-shot VC scenario. We take temporal dependency modeling as the starting point and propose a novel framework for Refinable Zero-shot VC, called RefZVC. RefZVC can greatly cover the long-term context of the video and continuously polish and refine the generated captions in a reward-feedback manner. We first design an Adaptive Frame Skipping module (AdaSkip) to skip redundant frames and select diverse keyframe sequences. Subsequently, we propose a Multi-granularity Reinforcement Polishing (MRP) mechanism, which iteratively polishes captions by leveraging Gaussian Kernel Cache (GKC) to capture temporal dynamics, store and reuse relevant historical context. In addition, MRP calculates rewards for generated captions at both the sentence-level and entity-level to achieve test-time polishing. With the MRP mechanism, RefZVC achieves superior zero-shot generalization performance, outperforming previous zero-shot VC methods on benchmarks such as MSVD, MSR-VTT, and VATEX. Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Lingling Li 0002, Xu Liu 0006 |
IEEE Trans. Image Process. | 2 |
| 2026 | Image Singularity Scattering Representation Learning ClassificationabstractThe multi-scale geometric analysis is a great representation tool. It can be used to improve the feature representation and learning process of deep networks. In addition to extracting features, the multi-scale geometric prior knowledge can also be used for the structure improvement of deep networks. In this paper, we propose a multi-scale scattering representation learning network, abbreviated as MSRLN, for image classification tasks. The exploration of structure improvement can be made with multi-scale scattering operations. In this way, the better singularity representation learning process for networks can be achieved. Firstly, the filter banks and multi-scale scattering operator are introduced for non-linear and singularity representation. Secondly, the novel multi-scale scattering representation learning network structure is designed. The scaling- wise scattering process is deployed in the shallow layer as a non-linear layer. This structure essentially supplements deep networks with geometric prior knowledge. It can further improve the non-linear activation and singularity representation process. Thirdly, we put forward the multi-stage scattering representation strategy and the prior knowledge weakening mechanism. With flexible scaling factors and learning rates, the stepwise approximation and learning process of networks can be achieved. In sum, MSRLN is a kind of structural innovative, and the scattering singularity representation structure can be extended to other backbones or tasks. Extensive experimental results show that MSRLN can achieve better image classification accuracy. Finally, necessary convergence, insight, and adaptability analyses are provided in evaluation experiments. Jie Gao 0013, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Puhua Chen, Yuwei Guo 0001, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 7 |
| 2026 | Learning to Prompt With Refining Text Knowledge for Zero-Shot Video Action RecognitionabstractFoundational vision-language models (VLMs) like CLIP are redefining the vision domain with their exceptional generalization capabilities. Prompt-based learning methods adapt pre-trained VLMs to video action recognition tasks using task-specific learnable text tokens. However, these tokens often struggle to generalize to unseen categories, as they tend to forget general textual knowledge. To address this, we construct knowledge prompts composed of handcrafted and descriptive prompts and introduce a novel knowledge-guided context mapping to enhance the generalization of learnable prompts to unseen categories. This approach mitigates the forgetting of fundamental knowledge by reducing the discrepancy between learnable prompts and knowledge prompts while simultaneously allowing the prompts to extract rich contextual knowledge from LLM data. Then, incorporating the knowledge-guided context mapping into the contrastive loss enables zero-shot transfer of prompts to new categories and data, providing discriminative prompts for both seen and unseen tasks. In addition, we propose an advanced temporal aggregation method that refines uniform mean pooling by incorporating frame-level textual relevance scoring. Extensive evaluations on multiple benchmarks demonstrate that learning to prompt with refining text knowledge is an effective quick-tuning method, achieving superior sample generalization performance without increasing training parameters. Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
IEEE Trans. Multim. | 2 |
| 2026 | Adaptive Multi-Modal Visual Tracking With Dynamic Semantic PromptsabstractRGB-based object tracking is a fundamental task in computer vision, aiming to identify, locate, and continuously track objects of interest across sequential video frames. Despite the significant advancements in the performance of traditional RGB trackers, they still face challenges in maintaining accuracy and robustness in the presence of complex backgrounds, occlusions, and rapid movements. To tackle these challenges, combining visual auxiliary modalities has gained significant attention. Beyond this, integrating natural language information offers additional advantages by providing high-level semantic context, enhancing robustness, and clarifying target priorities, further elevating tracker performance. This work proposes theAdaptiveMulti-modalVisual Tracking with Dynamic Semantic Prompts (AMVTrack) tracker, which efficiently incorporates image descriptions and avoids text dependency during tracking to improve flexibility and adaptability. AMVTrack significantly reduces computational resource consumption by freezing the parameters of the image encoder, text encoder, and Box Head and only optimizing a few learnable prompt parameters. Additionally, we introduce the Adaptive Dynamic Semantic Prompt Generator (ADSPG), which dynamically generates semantic prompts based on visual features, and theVisual-LanguageFusionAdaptation (V-L FA) method, which integrates multi-modal features to ensure consistency and complementarity of information. Additionally, we partition the Image Encoder to conduct an in-depth investigation into the relationship between the importance of features across different depth and width regions. Experimental results demonstrate that AMVTrack achieves significant performance improvements on multiple benchmark datasets, proving its effectiveness and robustness in complex scenarios. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
IEEE Trans. Multim. | 2 |
| 2026 | Adaptive Visual Prompting for Effective Satellite Video TrackingabstractSatellite video tracking presents significant challenges due to unpredictable target variations, environmental disturbances, and occlusions. Existing approaches either rely on auxiliary modalities or require full fine-tuning of foundation models, resulting in excessive parameter sensitivity and poor generalization. Meanwhile, conventional prompt-based tuning only updates parameters at a single location, limiting its ability to adapt to complex appearance changes. To address these limitations, we propose Adaptive Visual Prompting for Effective Satellite Video Tracking (AVPTrack). Unlike conventional prompts, introduced Super Prompts dynamically refine the original template at multiple distinct positions. This multi-location adaptation allows for fine-grained representation learning, enabling the tracker to better capture target variations and resist environmental disturbances. Additionally, Dynamic Templates are introduced to mitigate tracking failures in highly challenging scenarios, such as occlusions and background clutter, ensuring robust target localization. Furthermore, the Template Selection Adapter (TSA) selects the most relevant templates in real-time, enhancing tracking efficiency. These components are optimized during training while keeping other parameters frozen, ensuring parameter efficiency. We also investigate the relationship between fine-tuning proportions and learning rates to optimize model performance. Extensive evaluations on the SV248S, SatSOT, and VISO datasets demonstrate the superior adaptability and robustness of AVPTrack compared to existing methods. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Yanbiao Ma, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Mengjia Wang |
IEEE Trans. Multim. | 2 |
| 2026 | Regularized-Aware Discriminative Transformer Tracker for Satellite Videos
Xin Zhang 0167, Licheng Jiao, Lingling Li 0002, Zhongjian Huang, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 7 |
| 2026 | Multiscale Spatial-Frequency Learning for Degradation Decoupling in RS Image RestorationabstractRemote sensing (RS) images are prone to various degradations, which poses challenges to downstream tasks. Although existing single-task remote sensing image restoration methods are effective, they lack generalizability across tasks. All-in-one methods can handle multiple degradation tasks, but they usually focus on spatial information, ignoring the physical properties of the degradation information. To address the above limitations, we propose a Multiscale Spatial-Frequency Degradation Decoupling framework for All-in-One remote sensing image restoration (SFD$^{2}$IR), which decouples degradation features across different tasks to guide the model in performing task-specific image restoration. Specifically, a task-specific instruction generator (TIG) is proposed first to transform degradation features into task-specific prompts. Then, a multi-scale multi-frequency enhancement (MME) module is designed to decouple degradation effects from both spatial and frequency perspectives, thus enhancing the model's adaptability to various degradation types. Finally, a prompt feature refinement (PFR) module is developed to further refine the model's response to degraded tasks. Extensive experiments demonstrate that the proposed method achieves excellent performance on different RSIR tasks, including cloud removal, deblurring, dehazing, and super-resolution. The source code will be publicly available at SFD$^{2}$IR. Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2026 | Edit-Your-Motion: Space-Time Diffusion Decoupling Learning for Video Motion EditingabstractExisting diffusion-based methods have achieved impressive results in human motion editing. However, these methods often exhibit significant ghosting and body distortion in unseen in-the-wild cases. In this paper, we introduce Edit-Your-Motion, a video motion editing method that tackles these challenges through one-shot fine-tuning on unseen cases. Specifically, firstly, we utilized DDIM inversion to initialize the noise, preserving the appearance of the source video and designed a lightweight motion attention adapter module to enhance motion fidelity. DDIM inversion aims to obtain the implicit representations by estimating the prediction noise from the source video, which serves as a starting point for the sampling process, ensuring the appearance consistency between the source and edited videos. The Motion Attention Module (MA) enhances the model's motion editing ability by resolving the conflict between the skeleton features and the appearance features. Secondly, to effectively decouple motion and appearance of source video, we design a spatio-temporal two-stage learning strategy (STL). In the first stage, we focus on learning temporal features of human motion and propose recurrent causal attention (RCA) to ensure consistency between video frames. In the second stage, we shift focus on learning the appearance features of the source video. With Edit-Your-Motion, users can edit the motion of humans in the source video, creating more engaging and diverse content. Extensive qualitative and quantitative experiments, along with user preference studies, show that Edit-Your-Motion outperforms other methods. Yi Zuo 0003, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Wenping Ma 0001, Shuyuan Yang 0001, Yuwei Guo 0001 |
IEEE Trans. Multim. | 4 |
| 2026 | PromptVAD: Abnormal Prompt via Vision-Language ModelabstractWeakly supervised video anomaly detection (WSVAD) aims at predicting frame-level anomaly scores by modeling training videos with video-level annotations. The category names of abnormal events contain high-level knowledge abstracted by humans about abnormalities, which is of great help in identifying abnormal events. To utilize the knowledge implicit in category names, based on the visual-language pretraining model, we introduce a learnable abnormal prompt from three aspects: learnable domain prompt, learnable category prompt, and nonlearnable category definition prompt. Based on the learnable abnormal prompt, we propose a novel fine-grained WSVAD method: PromptVAD, which exploits a learnable abnormal prompt to reduce the semantic gap between visual images and anomaly categories. Through a similarity measure and our proposed coarse-grained two-class prompt module, our PromptVAD jointly learns coarse-grained and fine-grained VAD. Extensive experimental results on the ShanghaiTech, University of Central Florida (UCF)-Crime, and XD-Violence datasets show that our method achieves state-of-the-art performance. Specifically, our method achieves an area under the curve (AUC) of 88.62% on the UCF-Crime dataset. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Zehua Hao, Jiahao Wang 0002, Lingling Li 0002, Xu Liu 0006, Puhua Chen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2026 | DGNMF: Dynamic Diffusion Graph Nonnegative Matrix FactorizationabstractIn feature learning (FL), structural information shows advantages in retaining information and maintaining stability. Graph diffusion, a graph learning method that can focus on neighborhood structure and transmit information, has great research potential. In this study, a novel dynamic diffusion graph nonnegative matrix factorization (DGNMF) method is proposed, which uses a diffusion graph to improve the performance of FL and further enhances the effectiveness and stability of downstream classification tasks. DGNMF aims to mine and retain structural information more deeply in FL to build a more powerful and stable FL method. First, the model embeds graph learning into FL to obtain features containing structural information. Second, dynamic diffusion graph learning is used to mine deeper and more global structural information. Finally, we construct an updateable indicator matrix to enhance the discriminability of features. The classification experimental results of DGNMF on six databases demonstrate its advantages, verify its effectiveness and stability, and prove the importance of diffusion graph in improving FL. Chenxi Tian, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2026 | Spatial-Temporal Diffusion Model for Matrix FactorizationabstractMatrix factorization (MF) is a fundamental problem in machine learning, which is usually used as a feature learning method in various fields. For complex data involving spatiotemporal interactions, MF that only handles 2-D data will disrupt spatial dependence or temporal dynamics, failing to effectively couple spatial information with temporal factors. According to Markov chain principle, the spatial information of the present time is related to the spatial state of the previous time. We propose a spatial-temporal diffusion model for MF (STDMF), which uses graph diffusion to couple spatial-temporal information. Then, MF is used to learn the joint feature of data and spatial-temporal diffusion graph. Specifically, STDMF utilizes the graph diffusion with physical laws to generate spatial-temporal structure information. It obtains the underlying core structure of complex systems from a global perspective, which enhances the generalization ability of MF in noisy time-series data. To learn the lowest rank subspace of MF in time-series data, STDMF uses structural learning to constrain the rank of the learned features. Finally, STDMF is applied to clustering and anomaly detection of dynamic graph. The effectiveness of this method is verified by sufficient experiments, especially for noisy data. Chenxi Tian, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Licheng Jiao, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Logits DeConfusion with CLIP for Few-Shot LearningabstractWith its powerful visual-language alignment capability, CLIP performs well in zero-shot and few-shot learning tasks. However, we found in experiments that CLIP’s logits suffer from serious inter-class confusion problems in down-stream tasks, and the ambiguity between categories seriously affects the accuracy. To address this challenge, we propose a novel method called Logits DeConfusion, which effectively learns and eliminates inter-class confusion in logits by combining our Multi-level Adapter Fusion (MAF) module with our Inter-Class Deconfusion (ICD) module. Our MAF extracts features from different levels and fuses them uniformly to enhance feature representation. Our ICD learnably eliminates inter-class confusion in logits with a residual structure. Experimental results show that our method can significantly improve the classification performance and alleviate the inter-class confusion problem. The code is available at https://github.com/LiShuo1001/LDC. Shuo Li 0010, Fang Liu 0001, Zehua Hao, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
CVPR | 2 |
| 2025 | Knowledge-Guided Part Segmentation
Xuejian Gou, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, Hao Wang 0211, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
ICCV | 2 |
| 2025 | Domain-Aware Category-Level Geometry Learning Segmentation for 3D Point Clouds
Pei He, Lingling Li 0002, Licheng Jiao, Ronghua Shang, Fang Liu 0001, Shuang Wang 0001, Xu Liu 0006, Wenping Ma 0001 |
ICCV | 5 |
| 2025 | Hierarchical Variational Test-Time Prompt Generation for Zero-Shot Generalization
Zhaoyang Wu, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, LiXu Liu, Puhua Chen, Wenping Ma 0001 |
ICCV | 2 |
| 2025 | Language-Guided Hybrid Representation Learning for Visual Grounding on Remote Sensing ImagesabstractVisual grounding (VG) refers to detecting the specific objects in images based on linguistic expressions, and it has profound significance in the advanced interpretation of natural images. In remote sensing image interpretation, visual grounding is limited by characteristics such as the complex scenes and diverse object sizes. To solve this problem, we propose a novel remote sensing visual grounding (RSVG) framework, named language-guided hybrid representation learning Transformer (LGFormer). Specifically, we designed a multimodal dual-encoder Transformer structure called the adaptive multimodal feature fusion module. This structure innovatively integrates text and visual features as hybrid queries, enabling early-stage decoding queries to perceive the target position accurately. Then, the different modal information from the dual encoders is aggregated by hybrid queries to obtain the final object embedding for coordinate regression. Besides, a multi-scale cross-modal feature enhancement module (MSCM) is designed to enhance the self-representation of the extracted text and visual features and align them semantically. As for the hybrid queries, we use linguistic guidance to select visual features as the visual part and sentence-level features as the textual part. Finally, the LGFormer model we designed achieved the best results compared to existing models on the DIOR-RSVG and OPT-RSVG datasets. Xu Liu 0006, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Youlin Huang |
IJCAI | 5 |
| 2025 | FA3T: Feature-Aware Adversarial Attacks for Multi-modal TrackingabstractMulti-modal visual tracking leverages complementary sensor information to enhance robustness under challenging conditions. However, the security of multi-modal tracking systems remains largely unexplored. Existing attacks primarily target single-modal trackers or independently disrupt each modality, failing to exploit the inherent feature interactions and fusion mechanisms that define multi-modal tracking. As a result, these methods exhibit limited attack effectiveness and fail to assess multi-modal tracking systems' vulnerabilities accurately. Understanding these security risks is crucial, as adversarial threats could lead to severe failures in safety-critical applications. To address these challenges, a feature-aware adversarial attack, termed FA3T is proposed. It is designed to explicitly disrupt feature extraction and cross-modal alignment, thereby weakening the fusion process that multi-modal trackers rely on. To achieve this, a Frequency-Spatial Feature Separation (FSFS) module is constructed to perturb feature representations at multiple levels, weakening the modality-complementary advantages of multi-modal tracking. Furthermore, a Target Confusion Attack (TCA) module is devised to manipulate the target-background-template relationships, making it increasingly difficult for the tracker to distinguish the true target, significantly impairing tracking performance. Extensive experiments on five benchmark datasets (i.e., LasHeR, RGBT234, DepthTrack, VOT-RGBD2022, VisEvent) across three different modalities (RGB-T, RGB-D, and RGB-E) demonstrate that our attack substantially degrades state-of-the-art multi-modal trackers, exposing their susceptibility to adversarial threats. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
ACM Multimedia | 2 |
| 2025 | Imagining Vision From Language for Few-Shot Class-Incremental Learning
Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Yanbiao Ma, Puhua Chen, Lingling Li 0002, Xu Liu 0006, Xuejian Gou |
ACM Multimedia | 3 |
| 2025 | Preserving text space integrity for robust compositional zero-shot learning via mixture of pretrained experts
Zehua Hao, Fang Liu 0001, Licheng Jiao, Yaoyang Du, Shuo Li 0010, Hao Wang 0211, Pengfang Li, Xu Liu 0006, Puhua Chen |
Neurocomputing | 2 |
| 2025 | A Physics-Aware Collaborative Framework With Prototype Consistency for Noisy Label Signal Modulation ClassificationabstractSignal Modulation Classification (SMC) is a fundamental technique in wireless communications. However, the prevalence of label noise in practical scenarios severely constrains the advancement of SMC technology. Existing SMC methods heavily rely on high-quality labeled data and often underutilize the inherent physical prior knowledge of signals. To address these issues, this article proposes a Physics-Aware Collaborative Framework with Prototype Consistency (PhyCo-PC), designed for noisy label environments and operating without requiring reliable labels. Firstly, the framework leverages co-teaching for noise identification and incorporates a collaborative consensus-guided module for prototype learning and pseudo-label generation. Secondly, it constructs physics-guided downstream decision module that fuses deep learning features with instantaneous physical signal characteristics to enhance decision robustness. Thirdly, a Domain Knowledge-guided Adaptive Sample Selection (DKASS) strategy is introduced. DKASS parameterizes the selection rate scheduling function, incorporates domain knowledge to constrain the search space, and utilizes automated search for optimization. This enables the model to adaptively determine the optimal training strategy for varying noise environments. Finally, experimental results demonstrate that PhyCo-PC significantly improves SMC classification performance under complex label noise scenarios on the RML2016.10a/04c datasets, exhibiting excellent robustness and significant advantages. Lingling Li 0002, Jiadong Lin, Huaji Zhou, Xu Liu 0006, Fang Liu 0001, Licheng Jiao |
IEEE Internet Things J. | 6 |
| 2025 | LLM Knowledge-Driven Target Prototype Learning for Few-Shot Segmentation
Pengfang Li, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Xu Liu 0006, Puhua Chen, Lingling Li 0002, Zehua Hao |
Knowl. Based Syst. | 2 |
| 2025 | Knowledge-aware evolutionary graph neural architecture search
Chao Wang 0099, Jiaxuan Zhao, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
Knowl. Based Syst. | 5 |
| 2025 | Visual-Language Scene-Relation-Aware Zero-Shot CaptionerabstractZero-shot image captioning can harness the knowledge of pre-trained visual language models (VLMs) and language models (LMs) to generate captions for target domain images without paired sample training. Existing methods attempt to establish high-quality connections between visual and textual modalities in text-only pre-training tasks. These methods can be divided into two perspectives: sentence-level and entity-level. Although they achieve effective performance on some metrics, they suffer from hallucinations due to biased associations during training. In this paper, we propose a scene-relation-level pre-training task by considering relations as more valuable modal connection bridges. Based on this, we construct a novel Visual-Language Scene Relation Aware Captioner (SRACap), which expands the ability to predict scene relations while generating captions for images. In addition, SRACap possesses excellent cross-domain zero-shot generalization capability, which is driven by a well-designed scene reinforcement switching pipeline. We introduce a scene policy network to dynamically crop salient regions from images and feed them into a language model to generate captions. We integrate multiple expert CLIP models to form a mixture-of-rewards module (MoR) as a reward source, and deeply optimized SRACap through the policy gradient algorithm in the zero-shot inference stage. With the iteration of scene reinforcement switching, SRACap can gradually refine the generated caption details while maintaining high semantic consistency across visual-linguistic modalities. We conduct extensive experiments on multiple standard image captioning benchmarks, showing that SRACap can accurately understand scene structures and generate high-quality text, significantly outperforming other zero-shot inference methods. Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Unveiling and Mitigating Generalized Biases of DNNs Through the Intrinsic Dimensions of Perceptual ManifoldsabstractBuilding fair deep neural networks (DNNs) is a crucial step towards achieving trustworthy artificial intelligence. Delving into deeper factors that affect the fairness of DNNs is paramount and serves as the foundation for mitigating model biases. However, current methods are limited in accurately predicting DNN biases, relying solely on the number of training samples and lacking more precise measurement tools. Here, we establish a geometric perspective for analyzing the fairness of DNNs, comprehensively exploring how DNNs internally shape the intrinsic geometric characteristics of datasets-the intrinsic dimensions (IDs) of perceptual manifolds, and the impact of IDs on the fairness of DNNs. Based on multiple findings, we propose Intrinsic Dimension Regularization (IDR), which enhances the fairness and performance of models by promoting the learning of concise and ID-balanced class perceptual manifolds. In various image recognition benchmark tests, IDR significantly mitigates model bias while improving its performance. Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Lingling Li 0002, Wenping Ma 0001, Shuyuan Yang 0001, Xu Liu 0006, Puhua Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Predicting and Enhancing the Fairness of DNNs With the Curvature of Perceptual ManifoldsabstractTo address the challenges of long-tailed classification, researchers have proposed several approaches to reduce model bias, most of which assume that classes with few samples are weak classes. However, recent studies have shown that tail classes are not always hard to learn, and model bias has been observed on sample-balanced datasets, suggesting the existence of other factors that affect model bias. In this work, we first establish a geometric perspective for analyzing model fairness and then systematically propose a series of geometric measurements for perceptual manifolds in deep neural networks. Subsequently, we comprehensively explore the effect of the geometric characteristics of perceptual manifolds on classification difficulty and how learning shapes the geometric characteristics of perceptual manifolds. An unanticipated finding is that the correlation between the class accuracy and the separation degree of perceptual manifolds gradually decreases during training, while the negative correlation with the curvature gradually increases, implying that curvature imbalance leads to model bias. We thoroughly validate this finding across multiple networks and datasets, providing a solid experimental foundation for future research. We also investigate the convergence consistency between the loss function and curvature imbalance, demonstrating the lack of curvature constraints in existing optimization objectives. Building upon these observations, we propose curvature regularization to facilitate the model to learn curvature-balanced and flatter perceptual manifolds. Evaluations on multiple long-tailed and non-long-tailed datasets show the excellent performance and exciting generality of our approach, especially in achieving significant performance improvements based on current state-of-the-art techniques. Our work opens up a geometric analysis perspective on model bias and reminds researchers to pay attention to model bias on non-long-tailed and even sample-balanced datasets. Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Maoji Wen, Lingling Li 0002, Wenping Ma 0001, Shuyuan Yang 0001, Xu Liu 0006, Puhua Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Text generation and multi-modal knowledge transfer for few-shot object detection
Yaoyang Du, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Zehua Hao, Pengfang Li, Jiahao Wang 0002, Hao Wang 0211, Xu Liu 0006 |
Pattern Recognit. | 2 |
| 2025 | Knowledge-Driven Compositional Action Recognition
Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006 |
Pattern Recognit. | 2 |
| 2025 | VLPA-CLIP: Video Language Prompting and Adapting CLIP for efficient video action recognition
Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
Pattern Recognit. | 2 |
| 2025 | Knowledge-Aware Geometric Contourlet Semantic Learning for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) provides detailed spectral and spatial information, essential for precise earth observation and various applications. Deep learning has advanced HSI classification, but the scarcity of labeled data and large model parameters necessitate semi-supervised methods to enhance performance and generalization. In this paper, we propose a novel semi-supervised framework dubbed Knowledge-Aware Geometric Contourlet Semantic Learning (KGCSL), aiming to achieve high-precision HSI classification with limited samples leveraging geometric and semantic knowledge. Specifically, to fully leverage geometric knowledge, KGCSL incorporates multi-scale and multi-directional representations of the contourlet transform within the neural network, enhancing the robustness of feature extraction and interpretability. Furthermore, to fully utilize semantic knowledge, an entropy-weighted prototype loss function is designed that exploits the attribute relationships between labeled and unlabeled samples to guide the optimization of unlabeled samples, promoting comprehensive semantic learning. Comprehensive evaluations of the proposed KGCSL framework on three public HSI datasets show that it outperforms existing state-of-the-art HSI classification methods and exhibits excellent generalization capabilities in limited-sample scenarios. The source code is available athttps://github.com/ShirlySmile/KGCSL. Xueli Geng, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Prompt-Based Concept Learning for Few-Shot Class-Incremental LearningabstractFew-Shot Class-Incremental Learning (FSCIL) faces a huge stability-plasticity challenge due to continuously learning knowledge from new classes with a small number of training samples without forgetting the knowledge of previously seen old classes. To alleviate this challenge, we propose a novel method called Prompt-based Concept Learning (PCL) for FSCIL, which generalizes conceptual knowledge learned from old classes to new classes by simulating human learning capabilities. In our PCL, in the base session, we simultaneously learn common basic concepts from the training data and the class-concept weight of each class in a prompt learning manner, and in each incremental session, class-concept weights between new classes and previously learned basic concepts are learned to achieve incremental learning. Furthermore, in order to avoid catastrophic forgetting, we propose a distribution estimation module to retain feature distributions of previously seen classes and a data replay module to randomly sample features of previously seen classes in incremental sessions. We verify the effectiveness of our PCL on widely used benchmarks, such as miniImageNet, CIFAR-100, and CUB-200. Experimental results show that our PCL achieves competitive results compared with other state-of-the-art methods, especially we achieve an average accuracy of 94.02% across all sessions on the miniImageNet benchmark. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Contour Knowledge-Aware Perception Learning for Semantic SegmentationabstractThe diversity of contextual information is of great importance for accurate semantic segmentation. However, most methods focus on single spatial contextual information, which results in an overlap of the semantic content of categories and a loss of contour information of objects. In this article, we propose a novel contour knowledge-aware perception learning network (CKPL-Net) to capture diverse contextual information by space-category aggregation module (SCAM) and contour-aware calibration module (CACM). First, SCAM is introduced to enhance intraclass consistency and interclass differentiation of features. By integrating space-aware and category-aware attention, SCAM reduces the redundancy of features from a categorical perspective while maintaining spatial correlation of pixels, substantially avoiding the overlap of the semantic content in categories. Second, CACM is designed to maintain the integrity of objects by perceiving contour contextual information. It develops a novel contour-aware knowledge and adaptively transforms the grid structure of convolutions for boundary pixels, which effectively calibrates the representation of features near boundaries. Finally, the quantitative and qualitative analyses on the three public datasets: ISPRS Potsdam dataset, ISPRS Vaihingen dataset, and WHDLD dataset, demonstrate that the proposed CKPL-Net achieves superior performance compared with prevalent methods, which indicates diverse contextual information is beneficial for accurate segmentation. Chao You, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Multidistribution Time-Series Prototype Learning for Crop Mapping With Sentinel-1 SAR ImageryabstractTime-series Synthetic Aperture Radar (SAR) offers significant potential for crop mapping due to its all-weather, weather-independent imaging capabilities. Existing crop mapping methods with time-series SAR data have achieved good performance. However, these methods often ignore the phenological diversity of the same crop in time-series data, and their limited performance significantly constrains their potential for application in large-scale crop mapping. To address these issues, this paper proposes a prototype-based time-series remote sensing crop mapping framework called Multi-Distribution time-series Prototype Learning (MDPL). The framework aims to learn multiple time-series prototypes for the same crop for various phenological distributions of crops, effectively capturing the complex and varied phenological characteristics of crops. Secondly, a phenological-invariant feature learning module is proposed to enhance the model’s generalization capability for large-scale crop mapping. Additionally, a new temporal metric is proposed to capture phenological differences. Experimental results on three benchmark datasets have demonstrated the effectiveness and superiority of MDPL compared with state-of-the-art time-series SAR crop mapping methods. Yuwei Guo 0001, Licheng Jiao, Kairen Chen, Yujing Jia, Fang Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Fine-Grained Visual-Language Alignment for Remote Sensing Image-Text RetrievalabstractRemote sensing image-text retrieval (RSITR) is critical for applications, including environmental monitoring and disaster management. The main challenge in this field is that the multi-scale feature of remote sensing images and the semantic differences of professional texts make it difficult to achieve accurate alignment. Existing coarse-grained methods struggle to address the inherent difference between images and text. In light of this, we propose the Fine-Grained Visual-Language Alignment (FGVLA) method. Our FGVLA employs a hybrid loss function that combines coarse-grained contrastive and triplet loss with novel fine-grained loss. Fine-grained loss includes spatial mask loss and fine-grained contrastive loss to enhance semantic alignment. The method also introduces an inference process that works cooperatively with fine-grained loss to explicitly align image patches with textual nouns. Extensive experiments on RSICD, RSITMD, and UCM-Caption datasets demonstrate that FGVLA outperforms existing methods, achieving superior retrieval performance. The code of our FGVLA has been released at https://github.com/Ji-Haoyang/FGVLA. Shuo Li 0010, Haoyang Ji, Fang Liu 0001, Licheng Jiao, Xutong Min, Jiahao Wang 0002, Lingling Li 0002, Xu Liu 0006 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Multiscale Self-Supervised Constraints and Change-Masks-Guided Network for Weakly Supervised Change DetectionabstractRemote sensing change detection (CD) is a highly significant subtask within the field of Earth observation. Recently, weakly supervised CD (WSCD) methods based on image-level annotations have attracted interest, it is challenging to generate a clear margin between changed and unchanged regions with a lack of detailed annotation. In this article, based on class activation maps (CAMs), we propose a novel WSCD network based on self-supervised learning and change mask guidance (SSCMNet). First, we design a multiscale self-supervised constraint (MSC) module to narrow the gap between weak supervision and full supervision and compensate for the inherent shortcomings of CAMs. Second, a change mask guidance (CMG) module is proposed to further guide the network to keep the integrity of changed objects according to the consistency within unchanged regions and inconsistency within changed regions. Finally, to address the challenge of transferring commonly used post-processing methods in semantic segmentation to CD, an adaptive post-processing (APP) module is designed to adaptively select one of the input images for post-processing. We conduct experiments on three publicly available remote sensing CD datasets. Quantitative metrics and visualized results demonstrate the outstanding performance of the proposed method. Jia Liu 0020, Hejun Luo, Fang Liu 0001, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | LSCF: Long-Term Semantic-Guidance ConvFormer for Referring Remote Sensing Image SegmentationabstractReferring Remote Sensing Image Segmentation (RRSIS) task aims to generate segmentation masks for target objects based on language descriptions. It requires precise localization while distinguishing between visually similar yet semantically distinct objects. Fusing vision-language features only during extraction causes information loss and semantic forgetting in the decoder, harming similar target distinction. Additionally, high-resolution remote sensing images present challenges, including complex backgrounds, diverse object scales, and intricate boundaries, limiting the effectiveness of previous methods. To address these issues, we propose the Long-term Semantic-guidance ConvFormer (LSCF) Network. First, we fuse multi-receptive-field local features extracted by the Multi-scale CoordConv (MCC) module with language-aware global features from the Cross-modal Attention (CA) module to obtain multi-modal representations. Second, the Sampling Attention (SA) module enables fine-grained vision context alignment under semantic guidance. Finally, the Global Language Fusion (GLF) module is incorporated in the decoder to maintain long-term vision-language alignment and mitigate semantic degradation. Experimental validation on the RefSegRS, RRSIS-D, and RISBench datasets demonstrates that LSCF achieves oIoU scores of 83.27%, 77.42%, and 74.88%, and mIoU scores of 77.44%, 64.25%, and 68.53%, respectively. On RefSegRS, LSCF surpasses the SOTA method FIANet by 5.53% (oIoU) and 9.58% (mIoU), while delivering competitive performance on RRSIS-D and RISBench. Code and experimental configurations will be released. Lingling Li 0002, Xiaoqiang Lu, Licheng Jiao, Fang Liu 0001, Wenping Ma 0001, Xu Liu 0006 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | A Mamba-Aware Spatial-Spectral Cross-Modal Network for Remote Sensing ClassificationabstractThis study introduces a novel cross-modal spatial-spectral interaction Mamba (CMS2I-Mamba) for remote sensing image fusion classification. Unlike convolution-based models focusing on local details and Transformer-based models with high computational complexity, CMS2I-Mamba efficiently models global long-range dependencies in a linear complexity manner. First, multispectral (MS) and panchromatic (PAN) images each have unique advantages in the spectral and spatial attributes. Given this, this paper innovatively designs the multi-path selective-scan mechanism (MPS2M), which applies different path scanning strategies to deeply capture the global features from both spectral and spatial dimensions, aiming to enhance the robustness and complementarity of spatial-spectral features. Secondly, to overcome the characterization differences between images acquired by different sensors, this paper further introduces the channel interaction alignment module (CIAM). This module employs efficient former-last and oddeven channel interaction strategies to achieve precise semantic alignment of deep features between modalities. Finally, to leverage the shared fusion features to guide the unique singular features, this paper proposes a semantic-aware calibration module (SACM), which accurately constraints and calibrates the same semantic information in deep features. This not only enhances the model’s ability to understand scene semantics, but also promotes the deep fusion and utilization of information between different modalities. Through experimental verification on multiple datasets, the CMS2I-Mamba proposed in this paper shows excellent recognition performance and computational efficiency (parameter quantity and running speed) in fusion classification tasks. The code for CMS2I-Mamba is available at: https://github.com/ru-willow/CMSI-Mamba. Mengru Ma, Jiaxuan Zhao, Wenping Ma 0001, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | One Token for Detecting Various Changes
Licheng Jiao, Jie Chen 0098, Shuyuan Yang 0001, Fang Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Change Knowledge-Guided Vision-Language Remote Sensing Change DetectionabstractRemote sensing image change detection plays a critical role in applications like video surveillance and geographic information systems. However, existing binary and semantic change detection methods often rely solely on visual information, neglecting language information, which limits interpretability and the ability to provide specific change details. This work proposes the Change Knowledge-Guided Vision-Language Remote Sensing Change Detection (CKCD) method to address these limitations. By introducing change knowledge as language information, CKCD enhances semantic understanding and change detail representation. A Cross-Modal Affinity (CMA) module is designed to effectively fuse visual and textual features, improving information complementarity and fusion coherence. CKCD further enhances data utilization efficiency by merging change area detection and change category information into a single output through endto- end learning. This design reduces redundant data representations and simplifies the detection process, leading to a more compact and efficient use of the input data without requiring additional branches or multiple output heads. Experimental results demonstrate consistent performance improvements over traditional methods across multiple change detection datasets. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Dual Causal-Aware Detection Transformer for Remote Sensing ImagesabstractDeep neural networks often inherit biases from training data, compromising generalization. In visual recognition, distinguishing foreground from background is insufficient, as models tend to rely on spurious correlations rather than learning essential causal patterns. To address this issue, this paper proposes a novel transformer architecture, termed dual causal-aware detection transformer (DCDT), specifically designed for object detection in optical remote sensing images from a causal perspective. Specifically, we begin by constructing a structural causal model to intuitively analyze the causal effects inherent in the overall visual patterns. Building on this foundation, DCDT introduces causal constraints at the attention level by embedding dynamic multi-scale causal prototypes into the attention mechanism. The derived causal priors are subsequently used to enhance features at the representation level, thereby enforcing feature-level causal modulation. This dual causal-aware strategy enables the precise extraction and reinforcement of causally relevant features, improving both robustness and discriminative capability in complex detection scenarios. In addition, a sparse kernel-region mask is incorporated to decouple local information from global representations, effectively strengthening the modeling of fine-grained structures. Extensive experiments conducted on two challenging public datasets, DIOR and HRRSD, demonstrate that DCDT consistently outperforms existing methods and baselines. These results validate the effectiveness of DCDT in capturing both global causal semantics and local fine-grained features, highlighting its practicality in complex remote sensing scenarios. Yuhan Wang 0007, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Zhongjian Huang, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Local-Global Spectral Feature-Aware Learning for Hyperspectral Imagery ClassificationabstractEffective modeling of the relationship between local spectral details and global contextual information remains a core challenge in hyperspectral image (HSI) classification. In this paper, a local-global spectral feature-aware network (LGSFA-Net) is proposed, which achieves local and global spectral feature learning through synergistic integration of local convolutional inductive biases and global state-space models (SSMs). The architecture of LGSFA-Net comprises three sequential components, including an embedding stage using convolutions for fundamental feature extraction, an encoding stage that cascades standard Mamba blocks with specialized interactive Mamba (IMamba) blocks and an enhanced spatial-spectral feature fusion (ESSFF) module. The proposed IMamba blocks employ separable convolutions and feature interaction learning for explicitly modeling the cross-channel spectral correlations learning, which can be effective in awareness of the spectral feature. And then, the ESSFF module utilizes self-attention mechanisms to dynamically balance local and global spatial-spectral feature contributions. The final prediction stage incorporates a lightweight classification head for efficient inference. Experimental results validate the effectiveness of the proposed methods for HSI classification on four benchmark datasets, including the PaviaU, Houston, Honghu, and Hanchuan datasets. The proposed LGSFA-Net achieves approximately 1.48%-2.51% increased overall accuracy (OA), 1.34%-2.06% increased average accuracy(AA), and 1.37%-3.75% increased Kappa on the aforementioned four datasets, respectively, outperforming the contrasting methods. The code implementation will be available at https://github.com/yutinyang/LGSFA-Net. Yuting Yang 0008, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Fang Liu 0001, Shuo Li 0010, Wenping Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | CSCT: Channel-Spatial Coherent Transformer for Remote Sensing Image Super-ResolutionabstractRemote sensing image super-resolution (RSISR) techniques are crucial in practice as an economical approach to enhancing the resolution of remote sensing images (RSIs). The scale of structural information and the richness of texture details in RSIs far exceed those in natural images. Therefore, accurately restoring and preserving edge and detail information are a critical challenge in the super-resolution (SR) process. Currently, convolutional neural network (CNN)-based methods primarily rely on local feature extraction, which fails to effectively capture and integrate global contextual information. Generative adversarial network (GAN)-based methods, while improving the visual quality, often suffer from artifacts and training instability, adversely affecting image quality. Moreover, these approaches struggle to accurately represent high-frequency features, leading to blurriness or distortion when reconstructing fine details and edges. To address these limitations, we introduce the channel–spatial coherent transformer (CSCT). The core of CSCT includes the channel–spatial coherent attention (CSCA) and the frequency-gated feed-forward network (FGFN), which work synergistically to enhance edge and detail preservation while significantly improving overall image clarity. CSCA efficiently aggregates channel and spatial information, while FGFN adaptively adjusts frequency information to enhance high-frequency details and suppress low-frequency noise. Moreover, this article leverages advanced data augmentation methods that markedly boost RSISR performance, offering new avenues for further exploration. The empirical analysis across several remote sensing SR benchmark datasets reveals that our approach excels in detail restoration, effectively reduces artifacts and noise, and significantly enhances the quality of SR images. Kexin Zhang 0003, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Wenping Ma 0001, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Anomaly-Led Prompting Learning Caption Generating Model and BenchmarkabstractVideo anomaly detection (VAD) is an important intelligent system application, but most current research views it as a coarse binary classification task that lacks a fine-grained understanding of abnormal video sequences. We explore a new task for video anomaly analysis called Comprehensive Video Anomaly Caption (CVAC), which aims to generate comprehensive textual captions (containing scene information such as time, location, anomalous subject, anomalous behavior, etc.) for surveillance videos. CVAC is more consistent with human understanding than VAD, but it has not been well explored. We constructed a large-scale benchmark CVACBench to lead this research. For each video clip, we provide 6 fine-grained annotations, including scene information and abnormal keywords. A new evaluation metric Abnormal-F1 (A-F1) is also proposed to more accurately evaluate the caption generation performance of the model. We also designed a method called Anomaly-Led Generating Prompting Transformer (AGPFormer) as a baseline. In AGPFormer, we introduce an anomaly-led language modeling mechanism (Anomaly-Led MLM, AMLM) to focus on anomalous events in videos. To achieve more efficient cross-modal semantic understanding, we design the Interactive Generating Prompting (IGP) module and Scene Alignment Prompting (SAP) module to explore the divide between video and text modalities from multiple perspectives, and to improve the model's performance in understanding and reasoning about the complex semantics of videos. We conducted experiments on CVACBench by using traditional caption metrics and the proposed metrics, and the experimental results demonstrate the effectiveness of AGPFormer in the field of anomaly caption. Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Baoliang Chen |
IEEE Trans. Multim. | 2 |
| 2025 | Uncertainty Guided Progressive Few-Shot Learning Perception for Aerial View SynthesisabstractView synthesis of aerial scenes has gained attention in the recent development of applications such as urban planning, navigation, and disaster assessment. This development is closely connected to the recent advancement of the Neural Radiance Field (NeRF). However, when autonomousaerial vehicles(AAVs) encounter constraints such as limited perspectives or energy limitations, NeRF degrades with sparsely sampled views in complex aerial scenes. On this basis, we aim to solve this problem in a few-shot manner. In this paper, we propose Uncertainty Guided Perception NeRF (UPNeRF), an uncertainty-guided perceptual learning framework that focuses on applying and improving NeRF in few-shot aerial view synthesis (FSAVS). First, simply optimizing NeRF in complex aerial scenes with sparse input can lead to overfitting in training views, resulting in a collapsed model. To address this, we propose a progressive learning strategy that utilizes the uncertainty present in sparsely sampled views, enabling a gradual transition from easy to hard learning. Second, to take advantage of the inherent inductive bias in the data, we introduce an uncertainty-aware discriminator. This discriminator leverages convolutional capabilities to capture intricate patterns in the rendered patches associated with uncertainty. Third, direct optimization of NeRF lacks prior knowledge of the scene. This, coupled with a reduction in training views, can result in unrealistic rendering. To overcome this, we present a perceptual regularizer that incorporates prior knowledge through prompt tuning of a self-supervised pre-trained vision transformer. In addition, we adopt a sampled scene annealing strategy to enhance training stability. Finally, we conducted experiments with two public datasets, and the positive results indicate our method is effective. Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | LGSNet: Local-Global Semantics Learning Object DetectionabstractSelf-attention learns capturing the long-range dependencies between embeddings (e.g., image pixels). However, the memory overhead and computation cost are prohibitive due to being quadratic in term of the spatial resolution. The structure analysis reveals two crucial roles in the attention: the correlation-based dependency structure and feature normalization. In this work, an efficacious Local-Global Semantics (LGS) module is proposed to alleviate the above issues by modeling the local semantic aggregation and global semantic interaction. Our LGS module contains a group convolution and an Efficient Global Semantic Attention (EGSA). Firstly, the group convolution aggregates local semantics. Secondly, considering a feature map as a sequence of 2-D channel representations, EGSA formulates a general model for the global semantic interaction. The linear correlation is computed between global semantics. LGS has the linear memory overhead and computation cost in term of the spatial resolution. The LGS module can be smoothly incorporated into object detection frameworks. The experiment results verify its effectiveness on two popular detection datasets: the MS COCO and PASCAL VOC. Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Puhua Chen |
IEEE Trans. Multim. | 4 |
| 2025 | Semantic-Aware Wavelet Transformer for Pyramid Learning Object DetectionabstractTransformer displays the impressive capabilities on vision tasks. The built-in self-attention retains the quadratic computation burden in respect of the spatial resolution of image features. The traditional downsampling (e.g., average pooling) can reduce the resolution. Nonetheless, it may suffer from the dropping of detailed information. In this work, we propose an Efficient Wavelet Attention (EWA), which injects the wavelet transform and a Mean GELU (MGELU) function. Firstly, the wavelet transform enables the detailed information to participate in the efficient interaction modeling. Secondly, MGELU regards the statistical mean as reference and loosely passes the high relative responses. Building upon EWA, we present an effective Semantic-aware Wavelet Transformer (SWFormer), which is then employed for pyramid learning, including CNN feature hierarchy or Region of Interest (RoI) features. For the feature hierarchy, a Pyramid SWFormer (PSWFormer) incorporates SWFormer at each level to fit the bidirectional features. For RoIs, a Recognition-Localization SWFormer (RLSWFormer) is inserted into the head to fit their features from all levels. The effectiveness of our SWFormer is displayed experimentally on the MS COCO detection dataset and the Pascal VOC dataset. When exploiting Swin-small backbone, our SWFormer-based method acquires AP of 52.1 in the single-scale evaluation on the COCO test-dev set. This work will have the codes athttps://github.com/TimeIsFuture/Dt2_SWFormer. Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Puhua Chen |
IEEE Trans. Multim. | 4 |
| 2025 | Adaptive Complex Wavelet Informed Transformer OperatorabstractVisual transformers have achieved great success in representation learning. This is mainly due to efficient token dependency modeling via self-attention. However, the computational burden increases sharply as the input pixels increase. Although recent Fourier-based global frequency-domain mixing methods attempt to improve the efficiency of transformers for high-resolution image inputs, the Fourier operator has limited ability to capture the local geometric structure. Complex wavelets can perform local attention in both the spatial domain and the frequency domain. Therefore, we propose the complex wavelet informed transformer operator that uses the real and imaginary wavelets of the dual-tree complex wavelet transform to simulate the interaction in the attention kernel. In order to further reduce the computational burden of operators, we introduce an adaptive local block shared attention mechanism in the channel domain for our wavelet informed operators. Further, we construct the deep multi-head operator network consisting of a hybrid stack of complex wavelet informed transformer operators and self-attention layers. This enables the Transformer to more sparsely capture multi-scale and multi-directional structured features in the process of learning dependencies. Extensive experimental results show that our adaptive complex wavelet informed transformer operator under the Transformer architecture achieves highly competitive accuracy performance on multiple image classification benchmark datasets. And the proposed operators can be flexibly and effectively migrated to vision tasks in dynamic video scenarios. Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Hao Zhu 0009, Xu Liu 0006, Lingling Li 0002, Wenping Ma 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Uncertainty-Aware Semi-Supervised Learning Segmentation for Remote Sensing ImagesabstractDeep learning based remote sensing (RS) image segmentation significantly impacts several real application scenarios. Behind its success, massive labeled data plays an important role. However, annotating high-resolution RS images requires time-consuming and relevant expertise efforts. To address it, many works dive into semi-supervised learning which utilizes raw information embedded in unlabeled data to improve the segmentation model. Nevertheless, previous studies ignore the integrity and effectiveness of the potential context information hidden in RS data. In this work, we propose an uncertainty-aware masked consistency learning (U-MCL) framework that contains an uncertainty-aware masked denoising (U-MD) module and an uncertainty-aware masked image consistency (U-MIC) module. U-MCL initially generates a patch-wise uncertainty map for each unlabeled image during each training iteration, which is then used to derive an adaptive mask ratio for pseudo-label denoising in U-MD. Simultaneously, the uncertainty map is adopted to model a masked unlabeled image for reasoning unseen areas in U-MIC. Consequently, U-MCL is capable of enhancing model performance by engaging in accurate and stable consistency learning while preserving the integrity of the context and employing the context to infer the predictions of the masked regions safely. Extensive experiments on six RS datasets, i.e., ISPRS Vaihingen, FloodNet, MiniFrance, LoveDA, MER, and MSL, demonstrate the superiority of our U-MCL over recent most advanced methods, achieving new state-of-the-art performance under all benchmarks. Xiaoqiang Lu, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | A 3D Self-Awareness Diffusion Network for Multimodal ClassificationabstractAs imaging sensor technology in remote sensing has advanced quickly, multimodal fusion classification has become an important research direction in land cover and urban planning classification tasks. While generative models and image classification have greatly benefited from diffusion models, the present ones primarily concentrate on single-modality-driven diffusion processes. Therefore, this paper presents a 3D self-awareness diffusion network (3DSA-DiffNet) for multispectral (MS) and panchromatic (PAN) image fusion classification, which would make it easier to classify heterogeneous data from various sensors. First, in order to model the relationship between multi-channel spectra and multi-pixel spatial distributions as well as samples, respectively, a spatial-spectral joint denoising network (S$^{2}$JD-Net) is proposed. It can incorporate the diffusion process into the neural network to enhance the quality of diffusion features. Secondly, to imitate the brain's spatial-spectral coexistence learning mechanism, this work offers a 3D self-awareness module (3DSA-Module) that can learn the weight of each pixel in 3D space, resulting in extraordinarily high feature representation capabilities. Finally, experimental verification demonstrates that the 3D self-awareness diffusion fusion network driven by brain inspiration outperforms more sophisticated approaches on the Xi'an, Huhhot, and Muufl datasets. Mengru Ma, Wenping Ma 0001, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Yuwei Guo 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Tracking Like Human: Dynamic Scene Learning Reasoning Tracker in Satellite VideosabstractIn satellite video object tracking, the individual frame analysis method is usually used for target localization, ignoring informative cues of the dynamic scene. Temporal information could contribute to identifying the target from distractors. In this work, a novel dynamic scene learning reasoning tracker is proposed for satellite videos, which reasons over temporal dynamic information to derive the target location. It is inspired by the tracking pattern through human perception and reasoning. First, static-dynamic united analysis is designed to construct dynamic scenes by concatenating the static searching results along the temporal dimension. Second, the information of each response object is aggregated by wavelet transforms. Meanwhile, these scenes are projected into low-frequency and high-frequency subspaces, which could imitate different levels of perceptions of humans for scenes. Third, an object-aware reasoning transformer is proposed to utilize the temporal dynamics of input response objects. In each subspace, it models the mutual interactions between dynamic objects and further learns the intrinsic property of each object for target reasoning. Finally, to obtain the current reasoning result, inverse wavelet transforms are utilized to integrate the results of low-frequency and high-frequency subspaces. The effectiveness of the proposed method is validated on three public satellite video datasets, including SV248S, SkySat, and VISO. Qualitative and quantitative experimental results show that the proposed tracker outperforms 22 popular approaches in seven challenging tracking satellite scenarios. Licheng Jiao, Yangyang Li 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | Heterogeneous Riemannian Few-Shot Learning NetworkabstractHow to learn and accurately distinguish new concepts from few samples, as humans do, is a long-standing concern in artificial intelligence (AI). Studies in brain science and neuroscience have shown that human brain perception is based on nonlinear manifolds, and high-dimensional manifolds can facilitate concept learning in neural circuits. Based on this inspiration, in this paper, we propose a heterogeneous Riemannian few-shot learning network (HRFL-Net), which is the first few-shot learning method to perform end-to-end deep learning on heterogeneous Riemannian manifolds. Specifically, to enhance the geometric invariance of the image representation, the image features are projected into three heterogeneous Riemannian manifold spaces. Then, the implicit Riemannian kernel function maps the manifolds to the separable high-dimensional reproducing Hilbert space. It is assumed that the embedded kernel features of the complementary manifolds are mapped to the same common subspace. Thus, a novel neural network-based Riemannian metric learning method is designed to solve the subspace feature vectors by imposing orthogonal normalized projection, which overcomes the data extension limitation of the Riemannian metric. Finally, with the optimization objective of increasing the interclass distance and decreasing the intraclass distance in Hilbert space, the HRFL-Net is trained with end-to-end stochastic optimization, and the optimal aggregation subspace is learned during the gradient descent process. Thus, the proposed HRFL-Net can be easily generalized to challenging nonconvex data. The evaluation of four public datasets shows that the proposed HRFL-Net has significant superiority and also achieves competitive results compared with the state-of-the-art methods. Jie Chen 0098, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Yuwei Guo 0001, Puhua Chen, Wenping Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | A Spatial-Spectral Relation-Guided Fusion Network for Multisource Optical RS Image ClassificationabstractMultisource optical remote sensing (RS) image classification has obtained extensive research interest with demonstrated superiority. Existing approaches mainly improve classification performance by exploiting complementary information from multisource data. However, these approaches are insufficient in effectively extracting data features and utilizing correlations of multisource optical RS images. For this purpose, this article proposes a generalized spatial-spectral relation-guided fusion network (S2RGF-Net) for multisource optical RS image classification. First, we elaborate on spatial- and spectral-domain-specific feature encoders based on data characteristics to explore the rich feature information of optical RS data deeply. Subsequently, two relation-guided fusion strategies are proposed at the dual-level (intradomain and interdomain) to integrate multisource image information effectively. In the intradomain feature fusion, an adaptive de-redundancy fusion module (ADRF) is introduced to eliminate redundancy so that the spatial and spectral features are complete and compact, respectively. In interdomain feature fusion, we construct a spatial-spectral joint attention module (SSJA) based on interdomain relationships to sufficiently enhance the complementary features, so as to facilitate later fusion. Experiments on various multisource optical RS datasets demonstrate that S2RGF-Net outperforms other state-of-the-art (SOTA) methods. Xueli Geng, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Brain-Inspired Learning, Perception, and Cognition: A Comprehensive ReviewabstractThe progress of brain cognition and learning mechanisms has provided new inspiration for the next generation of artificial intelligence (AI) and provided the biological basis for the establishment of new models and methods. Brain science can effectively improve the intelligence of existing models and systems. Compared with other reviews, this article provides a comprehensive review of brain-inspired deep learning algorithms for learning, perception, and cognition from microscopic, mesoscopic, macroscopic, and super-macroscopic perspectives. First, this article introduces the brain cognition mechanism. Then, it summarizes the existing studies on brain-inspired learning and modeling from the perspectives of neural structure, cognitive module, learning mechanism, and behavioral characteristics. Next, this article introduces the potential learning directions of brain-inspired learning from four aspects: perception, cognition, understanding, and decision-making. Finally, the top-ten open problems that brain-inspired learning, perception, and cognition currently face are summarized, and the next generation of AI technology has been prospected. This work intends to provide a quick overview of the research on brain-inspired AI algorithms and to motivate future research by illuminating the latest developments in brain science. Licheng Jiao, Mengru Ma, Pei He, Xueli Geng, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001, Biao Hou, Xu Tang 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Multiscale Deep Learning for Detection and Recognition: A Comprehensive SurveyabstractRecently, the multiscale problem in computer vision has gradually attracted people's attention. This article focuses on multiscale representation for object detection and recognition, comprehensively introduces the development of multiscale deep learning, and constructs an easy-to-understand, but powerful knowledge structure. First, we give the definition of scale, explain the multiscale mechanism of human vision, and then lead to the multiscale problem discussed in computer vision. Second, advanced multiscale representation methods are introduced, including pyramid representation, scale-space representation, and multiscale geometric representation. Third, the theory of multiscale deep learning is presented, which mainly discusses the multiscale modeling in convolutional neural networks (CNNs) and Vision Transformers (ViTs). Fourth, we compare the performance of multiple multiscale methods on different tasks, illustrating the effectiveness of different multiscale structural designs. Finally, based on the in-depth understanding of the existing methods, we point out several open issues and future directions for multiscale deep learning. Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Zhixi Feng, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Complex Dual-Tree Pyramid Scattering TransformerabstractAttention-based transformer networks have recently played an increasingly important role in computer vision tasks. However, since pixel-by-pixel attention multiplication does not involve constraint assumptions such as spatial invariance, the computational complexity grows quadratically with the increase of input pixels. Therefore, this article proposes a complex pyramid scattering Transformer in dense scale space, which introduces sparse scattering constraints with a small number of wavelet basis parameters. It enhances the Transformer's flexibility and sparsity in multiscale space and, to a certain extent, slows down the increase in computational complexity caused by multiresolution input. In addition, compared with the general single-tree real wavelet transform, the dual-tree complex scattering method improves the aliasing of the scattering attention layer and helps obtain a more robust feature representation. At the same time, the multihead stepwise pyramid scattering coupling mechanism helps increase the abundance of directional priors. We conduct experiments in image classification and video tracking scenarios and verify the reliability and superiority of our dual-tree complex pyramid scattering Transformer for visual tasks with different scale requirements. The performance is better than that of the baseline Transformer and other advanced wavelet scattering networks at the same parameter scale. The code is available at https://github.com/Dawn5786/CPSTFormer. Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Hao Zhu 0009, Xin Zhang 0167, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Chain-of-Situation Aware Progressive Inference LearningabstractThe grounded situation recognition (GSR) task aims to recognize the structured semantics of an image to achieve "human-like" event understanding. Most previous studies primarily focus on the visual features of the situation, overlooking the step-by-step cognitive reasoning process that humans employ in complex task settings. Recently, the emergence of multimodal large language models (MLLMs) has provided novel directions for addressing complex problems. However, directly deploying MLLMs on the GSR task is suboptimal due to their tendency to exhibit "hallucination" issues. Additionally, fine-tuning MLLMs for the GSR task incurs high training costs. To address these challenges, inspired by human cognitive theory and the chain-of-thought (CoT) strategy, we propose the chain-of-situation progressive inference learning (CoS-PIL) framework, a lightweight approach that progressively completes verb prediction, noun prediction, and role grounding. The prediction of each step depends on the historical information of the previous step. Specifically, we first design situation prompts tailored to the GSR task and utilize MLLMs to analyze the input image and language prompts, generating heuristic response text for the current situation in the image. Instead of fine-tuning the MLLM, we activate the reasoning capabilities of the frozen MLLM and adapt its generated responses into three lightweight modules: CoS-Verb, CoS-Noun, and CoS-Ground. Considering that MLLMs may generate redundant content, we carefully design the chain-of-interest predictor (CoI-Predictor) to extract key information from the extensive response text and inject it into the model as prompts to enhance the performance. Extensive experiments on the challenging SWiG benchmark demonstrate that CoS-PIL outperforms other state-of-the-art methods. The code is publically available at https://github.com/XDLiuyyy/CoS-PIL. Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | CoT: Contourlet Transformer for Hierarchical Semantic SegmentationabstractThe Transformer-convolutional neural network (CNN) hybrid learning approach is gaining traction for balancing deep and shallow image features for hierarchical semantic segmentation. However, they are still confronted with a contradiction between comprehensive semantic understanding and meticulous detail extraction. To solve this problem, this article proposes a novel Transformer-CNN hybrid hierarchical network, dubbed contourlet transformer (CoT). In the CoT framework, the semantic representation process of the Transformer is unavoidably peppered with sparsely distributed points that, while not desired, demand finer detail. Therefore, we design a deep detail representation (DDR) structure to investigate their fine-grained features. First, through contourlet transform (CT), we distill the high-frequency directional components from the raw image, yielding localized features that accommodate the inductive bias of CNN. Second, a CNN deep sparse learning (DSL) module takes them as input to represent the underlying detailed features. This memory- and energy-efficient learning method can keep the same sparse pattern between input and output. Finally, the decoder hierarchically fuses the detailed features with the semantic features via an image reconstruction-like fashion. Experiments demonstrate that CoT achieves competitive performance on three benchmark datasets: PASCAL Context [57.21% mean intersection over union (mIoU)], ADE20K (54.16% mIoU), and Cityscapes (84.23% mIoU). Furthermore, we conducted robustness studies to validate its resistance against various sorts of corruption. Our code is available at: https://github.com/yilinshao/CoT-Contourlet-Transformer. Yilin Shao, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Automatic Graph Topology-Aware TransformerabstractExisting efforts are dedicated to designing many topologies and graph-aware strategies for the graph Transformer, which greatly improve the model's representation capabilities. However, manually determining the suitable Transformer architecture for a specific graph dataset or task requires extensive expert knowledge and laborious trials. This article proposes an evolutionary graph Transformer architecture search (EGTAS) framework to automate the construction of strong graph Transformers. We build a comprehensive graph Transformer search space with the micro-level and macro-level designs. EGTAS evolves graph Transformer topologies at the macro level and graph-aware strategies at the micro level. Furthermore, a surrogate model based on generic architectural coding is proposed to directly predict the performance of graph Transformers, substantially reducing the evaluation cost of evolutionary search. We demonstrate the efficacy of EGTAS across a range of graph-level and node-level tasks, encompassing both small-scale and large-scale graph datasets. Experimental results and ablation studies show that EGTAS can construct high-performance architectures that rival state-of-the-art manual and automated baselines. Chao Wang 0099, Jiaxuan Zhao, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationabstractPre-trained vision-language(V-L) models such as CLIP have demonstrated impressive Zero-Shot performance in many downstream tasks. Since adopting contrastive video-text pairs methods like CLIP to video tasks is limited by its high cost and scale, recent approaches focus on efficiently transferring the image-based CLIP to the video domain. A major finding is that fine-tuning the pre-trained model to achieve strong fully supervised performance leads to low zero shot, few shot, and base to novel generalization. Instead, freezing the backbone network to maintain generalization ability weakens fully supervised performance. Otherwise, no single prompt tuning branch consistently performs optimally. In this work, we proposed a multimodal prompt learning scheme that balances supervised and generalized performance. Our prompting approach contains three sections: 1) Independent prompt on both the vision and text branches to learn the language and visual contexts. 2) Inter-modal prompt mapping to ensure mutual synergy. 3) Reducing the discrepancy between the hand-crafted prompt (a video of a person doing [CLS]) and the learnable prompt, to alleviate the forgetting about essential video scenarios. Extensive validation of fully supervised, zero-shot, few-shot, base-to-novel generalization settings for video recognition indicates that the proposed approach achieves competitive performance with less commute cost. Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Zehua Hao, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
AAAI | 2 |
| 2024 | Multiplane Prior Guided Few-Shot Aerial Scene RenderingabstractNeural Radiance Fields (NeRF) have been successfully applied in various aerial scenes, yet they face challenges with sparse views due to limited supervision. The acquisition of dense aerial views is often prohibitive, as unmanned aerial vehicles (UAVs) may encounter constraints in perspective range and energy constraints. In this work, we introduce Multiplane Prior guided NeRF (MPNeRF), a novel approach tailored for few-shot aerial scene rendering-marking a pioneering effort in this domain. Our key insight is that the intrinsic geometric regularities specific to aerial imagery could be leveraged to enhance NeRF in sparse aerial scenes. By investigating NeRF's and Multiplane Image (MPI)'s behavior, we propose to guide the training process of NeRF with a Multiplane Prior. The proposed Multiplane Prior draws upon MPI's benefits and incorporates advanced image comprehension through a Swin V2 Transformer, pre-trained via SimMIM. Our extensive experiments demonstrate that MPN-eRF outperforms existing state-of-the-art methods applied in non-aerial contexts, by tripling the performance in SSIM and LPIPS even with three views available. We hope our work offers insights into the development of NeRF-based applications in aerial scenes with limited data. Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Puhua Chen, Yuwei Guo 0001 |
CVPR | 5 |
| 2024 | Multimodal Segformer for Flood Rapid Mapping with Sentinel-2 DataabstractFlood rapid mapping products play an important role in informing flood emergency response and management. To this end, the 2024 IEEE GRSS Data Fusion Contest Track 2 (DFC24-T2) establishes a multimodal benchmark for the segmentation of flood areas from Sentinel-2 multispectral images. However, the problems of imbalanced data distribution, data scarsity, and inter-modal differences severely inhibit the performance of deep-learning-based segmentation networks. In this work, we propose an end-to-end Multimodal Transformer-based Segmentation Network (MTSN) for accurate flood rapid mapping. MTSN first employs two Siamese encoders with shared parameters to accept multimodal inputs and output their respective hierarchical multiscale features, which are then enriched by several channel attention blocks. Subsequently, a Cross-modal Feature Fusion Module (CFFM) based on a gated mechanism is proposed to efficiently integrate the benefits of multimodal features, and generate informative representations. Finally, the fused features are decoded by a lightweight pure multilayer perception decoder to quickly generate mapping results of flood areas. Moreover, we introduce offline data augmentation, semi-supervised learning, test-time augmentation, and multimodal post-process to further boost the performance and generalization of our MTSN. Experimental results and extensive ablations show the effectiveness of our method. Code is available at https://github.com/xiaoqiang-lu/MMSegFormer. Xiaoqiang Lu, Tong Gou, Zhongjian Huang, Yuting Yang 0008, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001 |
IGARSS | 8 |
| 2024 | Pseudo-Viewpoint Regularized 3D Gaussian Splatting For Remote Sensing Few-Shot Novel View SynthesisabstractIn remote sensing (RS), Few-Shot Novel View Synthesis (FS-NVS) focuses on creating images of unobserved viewpoints using limited training images. Recently, 3D Gaussian Splatting (3DGS) has drawn scholars’ attention by its increasing rendering speeds and providing an explicit neural representation for 3D scenes. However, 3DGS tends to overfit limited training data. To tackle this challenge, we propose a Pseudoview Regularized 3DGS (PR3DGS) FSNVS method for RS scenarios. Our PR3DGS method introduces a pseudo-views regularization module to discriminate synthetic RS images generated from training- or pseudo-viewpoints. Therefore, our PR3DGS method can effectively mitigate overfitting in seen views and enhance the model’s capability to generate more realistic RS images from novel viewpoints. Besides, the excellent experimental results on the LEVIR-NVS dataset demonstrate the effectiveness of our method in RS FSNVS. Yijing Wang 0004, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0001, Licheng Jiao |
IGARSS | 6 |
| 2024 | Multi-Scale Co-Attention Learning For SAR Image Change DetectionabstractSynthetic aperture radar (SAR)iamge change detection is an important and challenging task. Existing methods mainly extract features from the difference map and original images in a parallel structure, ignoring the correlations between features. In this paper, we propose a novel SAR image change detection method based onmulti-scale and co-attention learning, called the Multi-scale Co-attention Network (MCNet). Specifically, we extract features from the difference map and original images using co-attention mechanism, which establishes a relationship between the two kinds of features in the feature space while retaining key information from the original images. With this attention mechanism, the model pays more attention to the changed areas in SAR images, and the features contain richer information. Furthermore, we propose a multi-scale feature fusion method to combine high-level and low-level semantic information, improving the generality and robustness of the features. Finally, the performance of the proposed method has been validated for its effectiveness. Xu Liu 0006, Licheng Jiao, Fang Liu 0001 |
IGARSS | 4 |
| 2024 | Spectral-Spatial Attentions and Deep Supervision for Change Detection in Remote Sensing ImagesabstractThe task of remote sensing image change detection involves identifying differences between images captured in the same geographical area but at different times. When dealing with dual-time-series images, lighting and seasonal variations often make recognition challenging. To address the challenges, based on Unet++, we innovatively introduce the Spectral-Spatial Attention Module (SSAM) to better focus on fine-grained details. SSAM uses different frequency components to allocate differential weights to channels, allowing the network to pay more attention to the features relevant to the current task. Moreover, to better capture the change details, a multi-level deep supervision strategy is introduced to enhance the discriminative ability and robustness of early features. Our proposed method is named as SSUNet and has been validated on the CDD and LEVIRE-CD datasets, demonstrating significant advantages in detail recognition. Jia Liu 0020, Fang Liu 0001, Jingxiang Yang, Liang Xiao 0001 |
IGARSS | 4 |
| 2024 | Time-Guided Network for Remote Sensing Change DetectionabstractIn recent years, as a very important part of remote sensing interpretation, change detection (CD) has developed rapidly with deep learning and remote sensing interpretation. However, most of the existing CD methods focus on how to extract the spatial features of the bi-temporal image pairs, ignore the importance of the temporal features. To solve this problem, a new time-guided network (TG-Net) using temporal information to guide feature extraction is proposed in this paper. In TG-Net, we use a newly proposed time-guided feature fusion (TGF2) block that uses temporal information to guide spatial feature fusion to extract temporal and spatial information comprehensively. We conducted experiments on two publicly available remote sensing datasets, LEVIR-CD and WHU, and compared them with three common CD methods. Xu Tang 0004, Xiangrong Zhang, Jingjing Ma 0001, Fang Liu 0001, Licheng Jiao |
IGARSS | 5 |
| 2024 | Domain Generalization-Aware Uncertainty Introspective Learning for 3D Point Clouds Segmentation
Pei He, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0002, Shuyuan Yang 0001, Ronghua Shang |
ACM Multimedia | 5 |
| 2024 | Geometric Prior Guided Feature Representation Learning for Long-Tailed Classification
Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006, Puhua Chen |
Int. J. Comput. Vis. | 3 |
| 2024 | Token singularity understanding and removal for transformersabstractThis work delves into unveiling the singularity issue latent in global attention-based Transformers. Empirical and theoretical analyses elucidate that interrelationships among token channels lead to singularities, impeding the training of attention weights. Concretely, the similar neighbor pixels within image patches can form intercorrelated channels after being flattened. Images that one color dominates can possess correlated channels . Furthermore, the fixed global connection architecture retains correlation relationships, contributing to the persistence of singularities. High singularity risks reducing Transformers’ performance and robustness. Based on the singularity analysis, we propose the Token Singularity Removal (TSR) strategy. It incorporates the Dual-Tree Complex Wavelet Transform (DTCWT) stem and Feature Decorrelation (FD) loss, aiming to encourage Transformers to learn tokens with unrelated channels and eliminate singularities. Experimental validation across various image classification datasets and corruption image data sets demonstrate improved accuracy and robustness of Transformers utilizing the TSR strategy. Our code is publicly available at https://github.com/wdanc/TSR . Licheng Jiao, Shuyuan Yang 0001, Fang Liu 0001 |
Knowl. Based Syst. | 5 |
| 2024 | Satellite Video Object Tracking Based on Location PromptsabstractObject Tracking in satellite videos is a challenging task due to the small target size, low spatial resolution, limited appearance and texture information, and the potential for background confusion. While current state-of-the-art tracking methods perform well on natural images, they often produce unsatisfactory results when applied to satellite videos. In this paper, we address these challenges by leveraging location prompts and refining the feature extractor and bounding box refinement module. Furthermore, we integrate motion features to effectively handle illumination variations that frequently arise in satellite videos, thereby enhancing the overall robustness of the tracker. Our proposed approach, abbreviated as SVLPNet, has been thoroughly evaluated through extensive experiments conducted on two authentic satellite video datasets. The obtained results unequivocally showcase the promising potential of SVLPNet in facilitating object tracking on satellite videos. The source code and raw results will be released at https://github.com/Wprofessor/SVLPNet. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Yingjia Gao, Hao Wang 0211, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Shuo Li 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | A Graph Association Motion-Aware Tracker for Tiny Object in Satellite VideosabstractSatellite video object tracking involves tracking a specified tiny object within a wide scene. The insufficient appearance features of these tiny objects pose significant challenges to appearance-based object trackers, particularly in situations involving occlusion, target blur, and similar interferences. In this paper, a novel Graph Association MOtion-aware tracker (GAMO) is proposed for tiny object in satellite videos, which integrates motion and spatial relationship information. First, a Gaussian motion estimator is proposed that decouples motion into velocity and direction, rather than using traditional x-y movement modeling. This estimator predicts the object’s position and estimates motion uncertainty with a directional motion probability map. Furthermore, the estimated motion serves as a prior to guide the proposal sampling. A probabilistic proposal sampling module is designed that samples candidate bounding boxes according to the directional motion probability map, focusing on the region where the target is most likely to appear. Additionally, we implement a graph association module to model and propagate the spatial relationships between the target and neighboring objects over time. This relationship information assists the appearance features in distinguishing the target from similar interferences. Experiments on the Skysat-1, SV248S, and VISO datasets demonstrate the superiority of the proposed tracker. GAMO leverages motion and surrounding information, resulting in significant improvements with minimal computational overhead. The code and results will be publicly available inhttps://github.com/Midkey/GAMO. Zhongjian Huang, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Xiangrong Zhang, Lingling Li 0002, Puhua Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Multi-Grained Gradual Inference Model for Multimedia Event ExtractionabstractWith the development of multimedia technology, events are usually presented in multimedia forms, thus multimedia event extraction (MEE) has become more and more important. Existing MEE works usually use simple strategies to align two modalities, making it difficult to precisely extract events and arguments in complex multimedia documents. To address this problem, we propose a novel Multi-grained Gradual Inference Model (MGIM) that focuses on inferring and interpreting events in complex multimedia structures in a coarse-to-fine manner. To efficiently integrate textual and visual modalities, we design a Coarse-grained Alignment (CA) module, which represents the two modalities in a graph structure and performs coarse-grained alignment. Based on the CA module, we further propose a Fine-grained Inference module (FI) that fine-grained aligns text and image by performing multiple rounds of gradual inference. MGIM provides a comprehensive interpretation of multimedia events at two information granularities (coarse and fine). Extensive experiments on the M2E2 dataset demonstrate the effectiveness of MGIM. Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Self Pseudo Entropy Knowledge Distillation for Semi-Supervised Semantic SegmentationabstractRecently, semi-supervised semantic segmentation methods based on weak-to-strong consistency learning have achieved the most advanced performance. The key to such a technique lies in strong perturbations and multi-objective co-training. However, CutMix, the most commonly used data augmentation in this field, limits the strength of perturbations as it only focuses on single random local context. Besides, complex optimization targets also reduce computational efficiency. In this work, we propose an efficient consistency learning based framework. Specifically, a novel unsupervised data augmentation strategy, EntropyMix, is present for semi-supervised semantic segmentation. Patches of unlabeled data from multi-view augmentations are combined into new training samples based on their prediction entropy, which provides more informative and powerful perturbations for consistency regularization and impels the model to focus on cross-view local context. On this basis, we further propose Self Pseudo Entropy Knowledge Distillation (SPEED) to learn global pixel relations from multi- and cross-view perturbations by optimizing a linear combination of feature-and logit-level distillation loss, enhancing model performance without additional auxiliary segmentation heads or a complex pre-trained teacher model. The collocation of the two ideas above is a plug-and-play technique without additional modification. Extensive experimental results on PASCAL VOC and Cityscapes datasets under various training settings demonstrate the superiority of the proposed data augmentation strategy and self-distillation loss, achieving new state-of-the-art performance. Remarkably, our method reaches mIoU of 75.16% using only 0.87% labeled data on PASCAL VOC and mIoU of 76.98% using only 6.25% labeled data on Cityscapes. The code is available at https://github.com/xiaoqiang-lu/SPEED. Xiaoqiang Lu, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | MBSI-Net: Multimodal Balanced Self-Learning Interaction Network for Image ClassificationabstractA growing number of earth observation satellites are able to simultaneously gather multimodal images of the same area due to the expanding availability and resolution of satellite remote sensing data. This paper proposes a novel multimodal balanced self-learning interaction network (MBSI-Net) for the classification task. It involves a dual-branch teacher-student network that enables knowledge interaction and transfer between the multimodalities. Firstly, in order to introduce statistical information in addition to local and global structural information, a texture feature equalization module (TFE-Module) is proposed. This can enhance the texture information of features through histogram equalization and further improve the representation ability of features. Secondly, to enable the student network to provide timely feedback questions, the paper proposes a feature fusion module (F2-Module) that models and enhances teacher features through the student network. This helps to raise the classification’s accuracy by incorporating information from multimodal images. Finally, the paper proposes a loss function based on structural similarity analysis to ensure balanced self-learning between the student and the teacher networks. Taking the multispectral (MS) and the panchromatic (PAN) images of the same scene as examples, through experimental verification, the proposed method can achieve good results on multiple datasets compared with other methods. Therefore, it offers an effective method for classifying and fusing multimodal data. Mengru Ma, Wenping Ma 0001, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Shuyuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Visual and Language Collaborative Learning for RGBT Object TrackingabstractDespite the extensive research on RGBT object tracking, there are still several challenges and issues in practical applications, such as modality differences, lighting variations and disappearance of the target, and changes in viewpoint. Existing methods mostly address these issues by fusing image features, while neglecting a significant amount of target label information. To address these challenges, this paper introduces text to drive the alignment of visible and infrared image features, transforming features from different modalities into the same feature space and fully using complementary features between different modalities. Furthermore, inspired by the success of prompt learning in various tasks, we utilize prior boxes and language as prompts to further guide the model in tracking the target. Extensive experiments demonstrate that the proposed VLCTrack tracker has excellent potential in RGBT object tracking. Compared to previous methods developed for this purpose, our approach achieves state-of-the-art performance on three benchmark datasets. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Yingjia Gao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Domain Adaptation-Aware Transformer for Hyperspectral Object TrackingabstractVisual object tracking in natural scenes is a popular but challenging task, owing to the difficulties of feature representation from various changes of the targets, such as size change, deformation, illumination change, rotations, motion blur, background clutter, etc. High-speed hyperspectral imaging systems capture hyperspectral videos (HSVs) in wide spectral ranges and provide abundant spectral and spatial information to tell targets apart from backgrounds, alleviating the model drift in appearance-based tracking methods. However, different hyperspectral imagers, such as near-infrared (NIR), red-to-near-infrared (RedNIR), and visible (VIS), obtain heterogeneous types of data that could not be handled by common object trackers. In this paper, a domain adaptive Transformer framework is proposed for hyperspectral object tracking. Considering the HSVs are from different types of sensors, their heterogeneous features are learned in an adversarial way by domain label reverse learning with a gradient reversed layer. To fully utilize the spectral information in HSV frames, a band-wise spatial attention module (BSAM) is designed to emphasize the salient area near the target of interest. We adopt a Siamese-like Transformer tracker as the main structure for tracking. Our tracker outperforms top-ranking methods on a hyperspectral object tracking benchmark dataset containing three types, 87 hyperspectral videos in total. The comparison experiments validate the effectiveness of the proposed method. The source code and trained models of this work will be publicly available soon at https://github.com/LianYi233/Trans-DAT. Yinan Wu 0001, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Lingling Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Efficient LWPooling: Rethinking the Wavelet Pooling for Scene ParsingabstractExisting wavelet pooling methods discard the high-frequency sub-bands, which can improve the noise-robustness of convolutional neural networks (CNNs) but lose the essential detailed features. Besides, most of them depend on different wavelets, which is not adaptive. In this paper, a novel efficient lifting-based wavelet pooling (LWPooling) is proposed to alleviate the problems above. Firstly, wavelet pooling is rethought based on the equivalence of 2D discrete wavelet transform (DWT) and standard average pooling (SAP), which suggests the lack of detailed information on traditional wavelet pooling. Secondly, the efficient LWPooling module is proposed to adaptively capture and preserve the critical high-frequency features via lifting-based wavelets. It can constrain the features linear independence, which efficiently makes important features salient. Thirdly, the lifting-based wavelet collaborative network (LWCNet) is constructed for classification and segmentation tasks based on the efficient LWPooling module. Experiments are validated on Cifar10, Cifar100, and ADE20K datasets. It suggests that the efficient LWPooling can enhance CNN’s representation and achieve a particular performance advantage compared to average, maximum, and original wavelet pooling. Besides, the proposed LWCNet shows the potential for scene parsing. The code implementation will be available at https://github.com/yutinyang/LWCNet. Yuting Yang 0008, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Xiangrong Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Knowledge Guided Evolutionary Transformer for Remote Sensing Scene ClassificationabstractSolving the complex challenges of sophisticated terrain and multi-scale targets in remote sensing (RS) images requires a synergistic combination of Transformer and convolutional neural network (CNN). However, crafting effective CNN architectures remains a major challenge. To address these difficulties, this study introduces the knowledge guided evolutionary Transformer for RS scene classification (Evo RSFormer). It amalgamates adaptive evolutionary CNN (Evo CNN) with Transformers in a hybrid strategy synergistically, which combines fine-grained local feature extraction of CNNs with long-range contextual dependency modeling of Transformers. Furthermore, for the development of Evo CNN blocks, this paper presents a knowledge-guided adaptive efficient multi-objective evolutionary neural architecture search (MOE2-NAS) strategy. This approach markedly diminishes the labor-intensive characteristics associated with traditional CNN design, striking a balance for both accuracy and compactness. Additionally, by leveraging domain knowledge from natural scene analysis into the RS field, MOE2-NAS facilitates the efficiency of classical NAS. It utilizes a priori knowledge to generate promising initial solutions and constructs a surrogate model for efficient search. The effectiveness of the proposed Evo RSFormer has been rigorously tested on various benchmark RS datasets, including UC Merced, NWPU45, and AID. Empirical results strongly support the superiority of Evo RSFormer over existing methods. Furthermore, experiments on MOE2-NAS have been studied to confirm the important role of knowledge guidance in improving the efficiency of NAS. Jiaxuan Zhao, Licheng Jiao, Chao Wang 0099, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Mengru Ma, Shuyuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Evolutionary Dual-Stream TransformerabstractVision transformers (ViTs) are rapidly evolving and are widely used in computer vision. However, high-performance ViTs require many computations, which limit their further development in the vision field. In this article, a novel evolutionary dual-stream transformer (E-DST) model is proposed to alleviate the computational resource demand problem. A hybrid attention mechanism structure is proposed for a DST model. The DST model uses a dual-branch structure to fuse convolutional and transformer features. Combining the features learned by the transformer and convolution effectively saves model computational resources. In addition, an evolutionary optimizer is proposed to optimize the parameters of the model. The excellent search ability of the evolutionary algorithm is utilized to optimize the transformer model parameters. The convergence of the evolutionary optimizer is proved in this article. In addition, the proposed E-DST model is experimentally compared with a variety of classic models and their deformations based on three datasets. And, the evolutionary optimizer proves its generality in convolutional and recurrent neural networks. The experimental results show that the E-DST model can effectively reduce computational resources and that the evolutionary optimizer can solve large-scale optimization problems. In conclusion, our proposed method is feasible and effective. Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Cybern. | 4 |
| 2024 | Bi-Level Multiobjective Evolutionary Learning: A Case Study on Multitask Graph Neural Topology SearchabstractThe construction of machine learning models involves many bi-level multiobjective optimization problems (BL-MOPs), where upper-level (UL) candidate solutions must be evaluated via training weights of a model in the lower level (LL). Due to the Pareto optimality of subproblems and the complex dependency across UL solutions and LL weights, a UL solution is feasible if and only if the LL weight is Pareto optimal. It is computationally expensive to determine which LL Pareto weight in the LL Pareto weight set is the most appropriate for each UL solution. This article proposes a bi-level multiobjective learning framework (BLMOL), coupling the above decision-making process with the optimization process of the upper-level MOP (UL-MOP) by introducing LL preference$\boldsymbol {r}$. Specifically, the UL variable and$\boldsymbol {r}$are simultaneously searched to minimize multiple UL objectives by evolutionary multiobjective algorithms. The LL weight with respect to$\boldsymbol {r}$is trained to minimize multiple LL objectives via gradient-based preference multiobjective algorithms. In addition, the preference surrogate model is constructed to replace the expensive evaluation process of the UL-MOP. We consider a novel case study on multitask graph neural topology search. It aims to find a set of Pareto topologies and their Pareto weights, representing different tradeoffs across tasks at UL and LL, respectively. The found graph neural network is employed to solve multiple tasks simultaneously, including graph classification, node classification, and link prediction. Experimental results demonstrate that BLMOL can outperform some state-of-the-art algorithms and generate well-representative UL solutions and LL weights. Chao Wang 0099, Licheng Jiao, Jiaxuan Zhao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Evol. Comput. | 6 |
| 2024 | A Quantum Evolutionary Learning Tracker for VideoabstractVideo object tracking has been a popular area in the field of computer vision. As video data evolves, more special perspectives and challenging video data are constantly kept up to date. This poses challenges for object tracking tasks and places higher demands on the generalization capabilities of the models. In this article, we propose a novel quantum evolutionary learning tracker (QELT) for video. The model combines quantum evolution with deep networks for tracking video objects. The model uses a QELT to generate a reliable population of candidate regions and a deep network for classification. In particular, the quantum evolutionary predictor predicts the object motion state through rotation operator and trajectory inference, and provides motion state information for the tracker. The predictor can incorporate object history contextual information and can provide stable candidate estimation populations for the model in case of failure of appearance features. Both quantum evolution and deep networks are combined to form an end-to-end online video object tracker. In addition, we propose a new video object tracking evaluation algorithm, Balanced Intersection over Union. The evaluation algorithm uses aspect ratios to balance the share of overlap and distance. Finally, we test the model on the OTB 2015 dataset for natural video and on the SV248A10-SOT dataset for satellite video. The performance of the proposed model is also analyzed and validated by comparing it with more than 20 classical tracker models. The experimental results show that our model has high generalization ability and robustness. Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Evol. Comput. | 5 |
| 2024 | Mask-Guided Correlation Learning for Few-Shot Segmentation in Remote Sensing ImageryabstractFew-shot segmentation aims to segment specific objects in a query image based on a few densely annotated images and has been extensively studied in recent years. In remote sensing, image segmentation faces challenges such as less training data, large intraclass diversity, and low foreground-background contrast. In this work, we propose a novel few-shot segmentation method in remote sensing imagery based on mask-guided correlation learning (MGCL) to alleviate the above challenges. In our MGCL, a novel mask-guided feature enhancement (MGFE) module is proposed, which makes features have intramask consistency by leveraging oversegmented masks. In order to enhance the contrast between foreground and background, a novel foreground-background correlation (FBC) module is proposed, which enhances background correlation representation by learning foreground correlation and background correlation separately. Furthermore, a novel mask-guided correlation decoder (MGCD) module is proposed to guide the decoder to focus on the consistency within the mask, thereby learning how to segment complete objects and improving segmentation accuracy. Sufficient experiments on the iSAID-$5^{i}$and DLRSD-$5^{i}$datasets show that our MGCL outperforms all comparative methods. In particular, in the one-shot setting of the iSAID-$5^{i}$dataset, we achieve an mIoU of 39.92 based on ResNet50, which is an improvement of 4.25 over the state-of-the-art (SOAT) method. The visualization of features before and after the MGFE module further concretely demonstrates the motivation and advantages of our MGCL. The code is available athttps://github.com/LiShuo1001/MGCL. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Xu Liu 0006, Puhua Chen, Lingling Li 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | MutSimNet: Mutually Reinforcing Similarity Learning for RS Image Change DetectionabstractChange detection involves analysis of discrepancies between two phases. However, when the unchanged elements are known, the changed features to be identified become straightforward. In addition, remote sensing image is constrained by limited spectral information, which leads to blurred boundaries between different semantics. Based on these two prior knowledge, in this artical, we introduce a novel change detection framework, named the mutually reinforcing similarity network (MutSimNet). This architecture aims to minimize false alarms along changing boundaries and reduce misjudgment rates among outliers. First, similarity learning is applied to change detection. The relationship between the two phases is considered when deriving the change feature maps. Second, we devise a mutually reinforcing loss function that integrates initial features with final features. Third, a self-attention module is connected in the feature pyramid network. This design mitigates information loss during the down-sampling process. Fourth, an attention feature fusion strategy is proposed for the integration of multi-layer features. This strategy takes into account the interaction between layer-by-layer features. Fifth, experimental results validate MutSimNet’s efficiency, particularly its ability to focus on edge contour learning. The MutSimNet also achieves superior performance on two benchmark datasets and predicts positive samples with higher probability. The codebase is accessible at https://github.com/ly-yu/MutSimNet. Xu Liu 0006, Yu Liu 0005, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | MGPACNet: A Multiscale Geometric Prior Aware Cross-Modal Network for Images Fusion ClassificationabstractConvolutional neural networks (CNNs) and self-attention (SA) are highly effective techniques used for the fusion of multisource remote sensing (RS) data, and they have found extensive application in Earth observation (EO) tasks. Nevertheless, CNNs are insufficient for the comprehensive extraction of contextual information and the representation of the sequential properties of spectral features. Furthermore, the loss of edge geometry information is often a consequence of information mining, which limits its application in RS. To address the abovementioned limitations, we propose a method called “multiscale geometric prior aware cross-modal network (MGPACNet)” for RS image fusion classification. First, a geometric prior feature enhanced residual module (GPFEResM) is created to extract shallow multiscale geometric edge prior features and detailed information from multimodal RS data to enhance feature boundary information. Second, a multiscale global-local spatial-spectral feature extraction module (MG-LS2FEM) uses multiscale spatial modeling and global-local spectral modeling to perceive rich semantic information in the spatial-spectral domain. Finally, a dual attention fusion module (DAFM) is designed to use pixel-level SA and cross-attention between heterogeneous data to achieve deep aggregation and cross-focusing of cross-modal information in two branches, and enhance the complementarity of heterogeneous data. A comprehensive examination of public RS data (hyperspectral-synthetic aperture radar (HS-SAR) Augsuburg/Berlin, hyperspectral-light detection and ranging (HS-LiDAR) Trento/MUUFL) from four distinct modalities (HS/SAR/LiDAR) has revealed that our method outperforms alternative models. Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Lighter and Robust: A Rotation-Invariant Transformer for VHR Image Change DetectionabstractIn recent years, change detection (CD) has emerged as an increasingly intricate research domain. However, in natural images, the orientation of objects is often aligned with the image boundaries, whereas in RS images, the imaging angles are random. As a result, existing CD methods encounter limitations when effectively representing vector features. In this article, we propose a rotation-invariant CD architecture named RFormer. It effectively utilizes direction-sensitive position embedding (DSPE) to represent features in RS images. To address the challenge of the quadratic growth in attention mechanism complexity with sequence length, we introduce low-cost cross attention (LC2A) to reduce its complexity to$1/{C^{2}}$. Furthermore, we employ the implicit timing extraction process (TEP) to represent interframe bitemporal features. TEP plays a crucial role in mitigating prediction biases caused by seasonal changes in land cover and prevents overconfident discrimination by the classifier in CD tasks. Experimental results demonstrate that RFormer achieves competitive performance on WHU, deeply supervised image fusion network (DSIFN)-CD, CDD, and LEVIR-CD datasets. Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | TrTr-CMR: Cross-Modal Reasoning Dual Transformer for Remote Sensing Image CaptioningabstractRemote sensing image captioning (RSIC) is an interesting but challenging cross-modal reasoning task for computer vision and natural language processing. Most of the recent popular approaches for RSIC utilize encoder-decoder architectures, which focus on visual features captured by convolutional neural network (CNN)-based encoder and semantic information by recurrent neural network (RNN)-based or long short-term memory (LSTM)-based decoder, but encounter difficulties with multiscale, multicategories, and direction ambiguity challenges. To make the most of semantic understanding ability of Transformers, in this article, we propose a new attention-based visual-linguistic reasoning framework with dual Transformer for RSIC. Specifically, Swin Transformer (SwinT) encoder with shifted window partitioning scheme is introduced for multiscale visual feature extraction to discover the intrinsic relationship in the objects, and then, a Transformer language model (TLM) with self-attention and cross attention is designed as the decoder to generate a well-formed sentence for the image. Extensive experiments are conducted on the public RSIC benchmark datasets, including UCM-Captions, Sydney-Captions, and RSICD. The impressive performance verifies the effectiveness and superiority of the proposed method. In addition, the source code and models of this work are publicly available athttps://github.com/LianYi233/TrTr-CMR. Yinan Wu 0001, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | High-Order Relation Learning Transformer for Satellite Video Object TrackingabstractSurrounding contexts are generally perceived as interfering with object tracking in satellite videos, leading to model drift. From another perspective, they can also be seen as reference objects of the tracked target, the dynamic interactions between them could provide essential information. In this article, a high-order relation learning transformer (HRLT) is proposed for satellite video object tracking, which not only models the high-order interactions of different target-context pairs but also reasons the associations between these high-order relations across multiple frames. First, a spatial high-order relation reasoning (SHR2) module is designed to model the high-order interactions between the target and scene contexts. Second, a temporal high-order relation reasoning (THR2) module is proposed to associate and reason these spatial high-order relations across multiple frames. Third, historical high-order relations are collected to provide more reasoning bases for the current frame prediction. Finally, qualitative and quantitative evaluations are performed on the SV248S, SkySat, and VISO datasets. The results show that HRLT outperforms 20 popular methods in different challenging scenarios. Licheng Jiao, Yangyang Li 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | LGLFormer: Local-Global Lifting Transformer for Remote Sensing Scene ParsingabstractIn deep learning, convolutional neural networks (CNNs) and transformers have gained excellent achievements in remote sensing scene parsing. Strong feature representation ability is still a challenge for them. Besides, the complex scenes are still essential challenges for deep learning in remote sensing scene parsing. In this article, an efficient local–global lifting transformer (LGLFormer) framework is proposed to ease the challenges above. It effectively combines CNNs, transformer, and wavelet transform to build a strong local–global (LG) feature representation network. Besides, global feature learning driven by LG adaptive features is proposed based on the 2-D LG adaptive feature extractor (LGAFE) and refined global feature attention module. The 2-D LG lifting feature extractor is inspired by the lifting scheme, which introduces local and global dependency. Furthermore, two LG lifting schemes are proposed, including the series and parallel modes, which can effectively learn LG relations between pixels. Finally, experiments are validated on three remote sensing benchmark datasets. The proposed LGLFormer achieves the state-of-the-art with 99.02%, 99.2%, and 99.48% overall accuracy (OA) on AID, WHU-RS19, and UCM datasets, respectively. In addition, LGLFormer shows good convergence with competitive parameters. The experimental code will be available athttps://github.com/yutinyang/LGLFormer. Yuting Yang 0008, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Relation Learning Reasoning Meets Tiny Object Tracking in Satellite VideosabstractTiny objects in satellite videos are usually not independent individuals, there exist rich semantic and temporal relations with each other. Thus, modeling and reasoning the variation of such intrinsic relationships can be beneficial for tiny object tracking. In this paper, a relation learning reasoning method is proposed for tiny object tracking in satellite videos. The core of the proposed is the relation reasoning network that consists of a key context module, a global semantic module, and a relation reasoning module sequentially. First, the key context module exploits global key contexts which explicitly or implicitly contribute to the target object, modeling the intrinsic relations with the target. Second, to reason the contribution, the global semantic module analyses the interaction between them in the same frame. Third, the relation reasoning module deduces the target based on the variation of the semantic relations among different frames. Such a relation learning reasoning approach which takes the target as the core is aligned with the satellite tiny object tracking task, significantly improves the identification performance in dense similarity scenes and the retrieval ability after completely occluded. Furthermore, the proposed method is shown to report improved qualitative and quantitative results on Jilin-1 and SkySat satellite video datasets. Licheng Jiao, Yangyang Li 0001, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Adaptive Multi-Scale Transformer Tracker for Satellite VideosabstractSatellite video tracking tasks are often characterized by blurred foreground boundaries in vast scenes, a wide range of targets varying in scale, and irregular changes in appearance. These challenges significantly impact the optimization of robust tracker performance. Therefore, it is imperative to extract diverse features with dynamic adaptive learning capabilities for the target being tracked in each sequence. In this article, we explore a novel adaptive multi-scale Transformer (MT) tracker for satellite videos to explore the potential spatiotemporal information of the target effectively. Specifically, a multi-scale spatial Transformer (MSST) is designed to leverage stage-by-stage spatial reduction and channel doubling, thereby enhancing the representation capabilities for the tracked target. In dynamic feature learning, an adaptive temporal Transformer (ATT) is then introduced based on multiple cross attentions, which analyzes the adaptive learning capacity for the dynamic target. It analyzes the weight proportion of different attentions automatically in the specific sequence through the learnable parameters. Finally, a multi-scale feature (MSF) regression module is crafted to improve the positioning accuracy of targets with low pixel counts in satellite scenes. This module accomplishes precise annotation of target boxes by effectively fusing features from diverse stages. We evaluate the proposed tracker performance on several public satellite datasets, including SatSOT, SV248S, and VISO. Experimental results show that the performance of our model can be comparable to the state-of-the-art trackers. Xin Zhang 0167, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Globally-Aware Continuous-Time Redistribution Learning for RS Image Change DetectionabstractChange detection (CD) based on deep learning has achieved excellent performance in recent years. However, these models exhibit limited capability in complete temporal modeling or face problems with fine-grained spatial features being overshadowed by the temporal context. Pure CNN-based CD pipelines also struggle to establish long-range connections. In this article, a globally aware continuous-time redistribution network (GCRNet) is proposed for RSCD. First, a boundary extraction branch is designed to preserve the semantic invariance of objects within the same boundary. This is achieved by providing boundary attention to adaptively guide the integration of temporal and spatial information. Then, a globally aware operator (GAO) is developed to obtain global interaction features. GAO utilizes the convolution theorem, which combines the Fourier transform and inverse Fourier transform, achieving it with low computational costs. Finally, an adaptive feature redistribution (AFR) module is designed to increase the distance between positive and negative samples in the latent space with change perception. It alleviates the effects of the severe class imbalance issue. Experimental results demonstrate that our proposed GCRNet surpasses 13 state-of-the-art CD methods. It achieves F1-score 0.33%, 0.62%, 0.84%, 0.17%, and 1.54% higher than the second-best model on the LEVIR-CD, LEVIR-CD+, WHU, CDD, and DSIFN datasets. The code of GCRNet is available athttps://github.com/XiaowenZhang-kuku/GCRNet. Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Effective and Robust: A Discriminative Temporal Learning Transformer for Satellite VideosabstractRobust feature learning has always been a research hotspot in dynamic temporal tasks. It makes the model almost unaffected by some challenging properties. The sequential nature of the transformer means attractive for temporal learning tasks, making it perform well in the video field. It is a current research hotspot for learning effective features by utilizing the target motion trends in satellite videos with multiple attributes, such as similar objects (SOBs) interference and occlusion. In this article, a novel discriminative temporal learning transformer tracker (DTLTracker) is introduced to characterize the dynamic target information for satellite videos. A discriminative transformer (DT) is proposed to comprehensively explore the dynamic target features with multiple attention mechanisms. It focuses on the primary information of the search area, making the target more discriminative. A fast convergence (FC) filter is designed to accelerate the weights convergence in calculating the target correlation operation, thereby ensuring the efficiency of model learning. The effectiveness and convergence have been demonstrated for the proposed optimization method. Additionally, a motion prior correction (MPC) module is constructed to utilize temporal information for target tracklet prediction, assisting the tracker in predicting the correct target. Numerous experiments are performed on three satellite videos to verify the effectiveness and feasibility of the proposed DTLTracker. It shows robustness compared to the state-of-the-art trackers on some challenging properties. Xin Zhang 0167, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Renormalized Connection for Scale-Preferred Object Detection in Satellite ImageryabstractSatellite imagery, due to its long-range imaging, brings with it a variety of scale-preferred tasks, such as the detection of tiny/small objects, making the precise localization and detection of small objects of interest a challenging task. In this article, we design a knowledge discovery network (KDN) to implement the renormalization group theory in terms of efficient feature extraction (FE). Renormalized connection (RC) on the KDN enables “synergistic focusing” of multiscale features. Based on our observations of KDN, we abstract a class of RCs with different connection strengths, called$n21$C, and generalize it to feature pyramid network (FPN)-based multibranch detectors. In a series of FPN experiments on the scale-preferred tasks, we found that the “divide-and-conquer” idea of FPN severely hampers the detector’s learning in the right direction due to the large number of large-scale negative samples and interference from background noise. Moreover, these negative samples cannot be eliminated by the focal loss function. The RCs extends the multilevel feature’s “divide-and-conquer” mechanism of the FPN-based detectors to a wide range of scale-preferred tasks, and enables synergistic effects of multilevel features on the specific learning goal. In addition, interference activations in two aspects are greatly reduced and the detector learns in a more correct direction. Extensive experiments of 17 well-designed detection architectures embedded with$n21$Cs on five different levels of scale-preferred tasks validate the effectiveness and efficiency of the RCs. Especially the simplest linear form of RC—E421C performs well in all tasks, and it satisfies the scaling property of renormalization group theory. All experiments can be trained and tested on a graphics card with 8 GB of video memory, which greatly enhances the applicability of our methodology. We hope that our approach will transfer a large number of well-designed detectors from the computer vision community to the remote sensing community. Datasets and codes will be available at:https://github.com/rabbitme/ Fan Zhang 0041, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Robust Instance-Based Semi-Supervised Learning Change Detection for Remote Sensing ImagesabstractSemi-supervised change detection (SSCD) has experienced rapid development, with numerous semi-supervised methods being proposed to reduce the reliance on labeled data in change detection. Existing approaches typically rely on manually set high-confidence thresholds to select robust pseudo-labels. However, the single-pixel threshold filtering method for pseudo-labels (STFP) lacks context correlation, cannot eliminate high-confidence false positive samples, and leads to erroneously filtering out low-confidence true positive samples. To address this issue, we propose robust instance-based semi-supervised learning change detection (RISL) for remote sensing images. RISL evaluates the reliability of each instance object by linking the semantic information of the context, thereby generating robust pseudo-labels. In RISL, firstly, a simple boundary trimming module (BT) as a preprocessing method for change prediction map is introduced. BT can effectively remove low-confidence false positive samples while avoiding confusion in the category of instance objects, thereby improving the quality of instance objects. Then, we propose a reliable instance evaluation module (RIEM) to evaluate the reliability of each instance object. RIEM combines the semantic information of the entire instance and establishes correlations between sample contexts to determine the reliability of the instance, effectively eliminating high false positive samples. In addition, the consistency regularization (CR) is integrated into RISL, and a new strategy suitable for RIEM is constructed. This strategy enhances the model’s generalization ability by mining and hiding semantic information from different views of unlabeled data. Experimental results on the challenging WHU-CD, LEVIR-CD, and CDD-CD datasets show that the proposed method achieves 89.80%, 90.01%, and 87.56% F1 scores on labeled data with 5% distribution. RISL achieves state-of-the-art performance compared to other methods. Yi Zuo 0003, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Hierarchical Dynamic Graph Clustering NetworkabstractConnections between visual components are ubiquitous. Graphs, as a highly flexible data structure, not only allow imposing relational induction bias on data, but can provide a completely distinct learning perspective for regular image data. In this paper, we propose a hierarchical dynamic graph clustering network (HDGCN) for visual feature learning. We construct hierarchical graph representations in graph domain in an adaptive, data-adaptive and task-adaptive manner. First, the initial graph is constructed in high-dimensional feature domain of images. To mine the hierarchical geometric features in latent graph space, adaptive clustering network (ClusterNet) is performed to learn discriminative clusters and generates cluster-based coarse graph. Then, graph convolutional networks (GCNs) are used to diffuse, transform and aggregate information among clusters. So, the intra-class and inter-class information is fully explored to increase the discriminativity of graph representations. Next, coarsened graph representations are mapped to grid based on its affinity with linear projection features. To further improve the task adaptation of clusters and hierarchical graph representations, ClusterNet and GCNs are fused in the same framework for end-to-end training and clusters is updated dynamically. We have conducted extensive experiments on classification and segmentation tasks. The experimental results fully validate the robustness of the proposed algorithm. Jie Chen 0098, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Puhua Chen, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Fast and Effective: Progressive Hierarchical Fusion Classification for Remote Sensing ImagesabstractMultisource remote sensing image fusion classification aims to produce accurate pixel-level classification maps by combining complementary information from different sources of remote sensing data. Existing methods based on Convolutional Neural Networks (CNN-based) utilize a patch-based learning framework, which has a high computational cost, leading to poor real-time performance. In contrast, methods based on Fully Convolutional Networks (FCN-based) can process the entire image directly, achieving fast inference. However, FCN-based methods require high computational resources and exhibit shortcomings in feature fusion, hindering practical applications. In this paper, a lightweight FCN-based Progressive Hierarchical Fusion Network (PHFNet) is tailored for multisource remote sensing image classification. PHFNet comprises a pyramid dual-path encoder and a pyramid decoder. In the encoder, cross-source features are hierarchically fused via the adaptive modulation fusion module (AMF), which leverages style calibration for cross-source alignment and promotes the complementarity of the fusion feature. In the decoder, we introduced an improved convolutional gated recurrent unit (iConvGRU) to progressively integrate the semantic and detailed information of hierarchical features, producing a context-enhanced global representation. In addition, we consider the relation between the channel number, convolutional kernel size, and parameter count to make the model as lightweight as possible. Comprehensive evaluations on three multisource remote sensing datasets demonstrate that PHFNet improves overall accuracy by 1.5% to 2.8% with a low computational overhead compared to state-of-the-art methods. The source code is avaliable athttps://github.com/ShirlySmile/PHFNet. Xueli Geng, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Cross-Domain Scene Unsupervised Learning Segmentation With Dynamic SubdomainsabstractUnsupervised cross-domain scene segmentation approach adapts the source model to the target domain, which utilizes two-stage strategies to minimize the inter-domain and intra-domain gap. However, the accumulation of errors in the previous stages affects the training of the subsequent stages. In this paper, a framework called statistical and structural domain adaptation (SSDA) is proposed to optimize inter-domain and intra-domain adaptation jointly. Firstly, the statistical inter-domain adaptation (StaIA) is proposed to model dynamic subdomains, which continuously adjust seed samples during the process of domain adaptation to mitigate error accumulation. The dynamic subdomains are modeled by exploring Bayesian uncertainty statistics and global balance statistics, which alleviate the imbalance problem in uncertainty estimation. StaIA encourages the model to transfer comprehensive and genuine knowledge through the seed loss for inter-domain adaptation. Secondly, the structural intra-domain adaptation (StrIA) is proposed to align the intra-domain gap among dynamic subdomains by the structural priors. Specifically, the StrIA models structural priors by truncated conditional random field (TruCRF) loss within the neighborhood, which constrains intra-domain semantic consistency to reduce the intra-domain gap. Experimental results demonstrate the effectiveness of the proposed cross-domain scene segmentation approaches on two commonly-used unsupervised domain adaptation benchmarks. The code is available at https://github.com/ChicalH/SSDA. Pei He, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Ronghua Shang, Shuang Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | A Category-Aware Curriculum Learning for Data-Free Knowledge DistillationabstractConstructing effective proxy data is one of the core challenges in data-free knowledge distillation. The existing models ignore the influence of the category entanglement of the generated data on the distillation. To alleviate this issue, imitating the human learning process, a new category-aware curriculum learning mechanism is proposed in this paper to perform data-free knowledge distillation, called CCL-D. The main ideology of this category-aware curriculum learning mechanism is to provide a new learning mode for data generation and network training, which enables the model to realize the knowledge distillation process from easy to difficult through automated curriculum learning. In this novel learning mechanism, a category-aware monitoring module is proposed to constrain the category attribute of generated data. Based on this monitoring module, the curriculum learning process for data generation and network training is designed and applied. Initially, the generator is guided to obtain new data with clear category features. The utilization of data with apparent category features is easy for student network training, and it enables the student network to learn clear and significant category features at the early training stage. Subsequently, the generator is guided to generate data with category entanglement. Utilizing these new data with category entanglement problems can improve the recognition ability of the student network to interclass interference and enhance network robustness. The effectiveness of the CCL-D is verified on the six benchmark experimental datasets (MNIST, CIFAR-10, CIFAR-100, SVHN, Caltech-101, Tiny-Imagenet). Xiufang Li, Licheng Jiao, Qigong Sun, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Multi-Scale Contourlet Knowledge Guide Learning SegmentationabstractFor accurate segmentation, effective feature extraction has always been a challenging problem, since the variability of appearance and the fuzziness of object boundaries. Convolutional neural networks have recently gained recognition in feature representation learning. However, it is only conducted in the spatial domain, and lacks effective representation of directionality, singularity and regularity in the spectral domain for anomaly detection of images. This is the key to feature learning representation of high-order singularity. To solve this problem, a multi-scale contourlet knowledge guide learning network is proposed in this paper. It is novel in this sense that, different from the CNNs in the spatial domain, the proposed method learns the multi-scale contourlet sparse representation to obtain more effective and sparse features in multi-scales and multi-directions. Furthermore, the contourlet knowledge guide learning can enhance the representation of spectral domain features. It is shown that the proposed network can learn the multi-level discriminative features and capture the more accurate object boundaries. The segmentation ability in theoretical analysis and experiments on five polyp segmentation datasets (CVC-ColonDB, CVC-ClinicDB, Kvasir-SEG, ETIS-LaribPolypDB, EndoSceneStill) and two building datasets (Massachusetts, WHU) are compared with developed methods. It must be emphasized that there is potential in effective feature learning representation and the generalization capability of the proposed method in deep learning, recognition and interpretation. Mengkun Liu, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Shuang Wang 0001, Biao Hou |
IEEE Trans. Multim. | 5 |
| 2024 | Bio-Inspired Multi-Scale Contourlet Attention NetworksabstractInspired by the sparse and hierarchical features representation in the ventral stream of the human visual system, the biologically inspired multi-scale contourlet attention network (BMCAnet) is proposed to extract robust discriminative features. First, we constructed the multi-scale contourlet filter banks as a population of neurons in the primary visual cortex (V1), and extracted sparse features in a multi-scale and multi-direction way. It simulated a simple cell in V1 that responds to stimuli in a specific direction. Second, in order to refine contourlet features adaptively, the Shannon block attention module (SBAM) is introduced by integrating Shannon entropy as the third branch of the channel attention module (CAM), thus the weights of contourlet coefficients can be learned adaptively. Third, the responses of the spatial and spectral features are pooled by the proposed contourlet pooling layer to obtain the invariant structure features with the specified rules, which roughly stimulate the pooling process of complex cells in the V1 area. Last, the combination of global average pooling (GAP) and full connection (FC) is used for classification. The competitive results on eight databases demonstrate that the BMCAnet can effectively extract sparse and effective features for the classification tasks. Mengkun Liu, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Xiangrong Zhang |
IEEE Trans. Multim. | 5 |
| 2024 | A Knowledge-Based Hierarchical Causal Inference Network for Video Action RecognitionabstractCurrently, existing action recognition methods mainly use a data-driven method to extract spatio-temporal representations of actions for recognition. However, this method may face performance bottlenecks. At the same time, existing action recognition methods are easily affected by the bias of scene information and object information in videos. In order to explore the essential causal relationship between factors and remove bias in action recognition, we introduce the theory of causal inference into the field of action recognition and propose a Knowledge-based Hierarchical Causal Inference Network (KHCIN) to help us step toward a new direction of inference in action recognition. First, we construct a Knowledge-based Hierarchical Causal Graph (KHCG) to structurally represent the scene, object and motion knowledge of a video. Then, in the model inference stage, we perform factual causal inference on a video on the constructed KHCG, and then deploy counterfactual inference on the Direct Content Hierarchy (DCH) and Indirect Interaction Hierarchy (IIH) in the KHCG. For DCH, we intervene in the model at the decision level to highlight bias errors in the model predictions. For the IIH, we focus on intervening in the feature modelling process. The biased interactions are revealed by interrupting the information communication in the feature space. By comparing the results of factual and counterfactual inference, we can easily expose the biased information in the original representations and eliminate them. Driven by counterfactual causal inference, our approach can significantly improve the performance of action recognition while improving model explainability. Extensive experiments demonstrate the effectiveness of this method. We hope that KHCIN can provide some new ideas for better introduction of causal inference theory in the action recognition community in the future. Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Lingling Li 0002, Yuwei Guo 0001, Puhua Chen |
IEEE Trans. Multim. | 2 |
| 2024 | Feature Distribution Representation Learning Based on Knowledge Transfer for Long-Tailed ClassificationabstractReal-world data typically follows a long-tailed distribution. When a small sample of tail classes does not cover the underlying distribution well, methods such as class re-balancing strategies and decoupled training are difficult to work, and additional knowledge needs to be introduced to recover the underlying distribution of the tail classes. In this work, we observe that the similarity between the variances of the feature distributions increases with the class similarity. Then, we also find that well-represented feature distributions typically contain multiple subcenters, which allows for denser samples at the edges of the distribution and promotes model learning to more robust decision bounds. Based on these observations, we propose to calibrate the feature distribution of the tail class by transferring the variance of the feature distribution of the head class, and then sample from the calibrated tail class distribution to generate augmented samples. To coordinate with the tail class calibration method, we also propose label-aware noise suppression (LANS) for reducing the generation of noisy samples and a three-stage training scheme for reshaping decision boundaries and compacting feature learning. Experimental results on iNaturalist2018, ImageNet-LT, CIFAR-10-LT, and CIFAR-100-LT show that our method achieves state-of-the-art performance in most metrics compared to similar approaches. Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006, Puhua Chen |
IEEE Trans. Multim. | 3 |
| 2024 | Multiresolution Interpretable Contourlet Graph Network for Image ClassificationabstractModeling contextual relationships in images as graph inference is an interesting and promising research topic. However, existing approaches only perform graph modeling of entities, ignoring the intrinsic geometric features of images. To overcome this problem, a novel multiresolution interpretable contourlet graph network (MICGNet) is proposed in this article. MICGNet delicately balances graph representation learning with the multiscale and multidirectional features of images, where contourlet is used to capture the hyperplanar directional singularities of images and multilevel sparse contourlet coefficients are encoded into graph for further graph representation learning. This process provides interpretable theoretical support for optimizing the model structure. Specifically, first, the superpixel-based region graph is constructed. Then, the region graph is applied to code the nonsubsampled contourlet transform (NSCT) coefficients of the image, which are considered as node features. Considering the statistical properties of the NSCT coefficients, we calculate the node similarity, i.e., the adjacency matrix, using Mahalanobis distance. Next, graph convolutional networks (GCNs) are employed to further learn more abstract multilevel NSCT-enhanced graph representations. Finally, the learnable graph assignment matrix is designed to get the geometric association representations, which accomplish the assignment of graph representations to grid feature maps. We conduct comparative experiments on six publicly available datasets, and the experimental analysis shows that MICGNet is significantly more effective and efficient than other algorithms of recent years. Jie Chen 0098, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Multiscale Dynamic Curvelet Scattering NetworkabstractThe feature representation learning process greatly determines the performance of networks in classification tasks. By combining multiscale geometric tools and networks, better representation and learning can be achieved. However, relatively fixed geometric features and multiscale structures are always used. In this article, we propose a more flexible framework called the multiscale dynamic curvelet scattering network (MSDCCN). This data-driven dynamic network is based on multiscale geometric prior knowledge. First, multiresolution scattering and multiscale curvelet features are efficiently aggregated in different levels. Then, these features can be reused in networks flexibly and dynamically, depending on the multiscale intervention flag. The initial value of this flag is based on the complexity assessment, and it is updated according to feature sparsity statistics on the pretrained model. With the multiscale dynamic reuse structure, the feature representation learning process can be improved in the following training process. Also, multistage fine-tuning can be performed to further improve the classification accuracy. Furthermore, a novel multiscale dynamic curvelet scattering module, which is more flexible, is developed to be further embedded into other networks. Extensive experimental results show that better classification accuracies can be achieved by MSDCCN. In addition, necessary evaluation experiments have been performed, including convergence analysis, insight analysis, and adaptability analysis. Jie Gao 0013, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | A Patch Diversity Transformer for Domain Generalized Semantic SegmentationabstractDomain generalization (DG) is one of the critical issues for deep learning in unknown domains. How to effectively represent domain-invariant context (DIC) is a difficult problem that DG needs to solve. Transformers have shown the potential to learn generalized features, since the powerful ability to learn global context. In this article, a novel method named patch diversity Transformer (PDTrans) is proposed to improve the DG for scene segmentation by learning global multidomain semantic relations. Specifically, patch photometric perturbation (PPP) is proposed to improve the representation of multidomain in the global context information, which helps the Transformer learn the relationship between multiple domains. Besides, patch statistics perturbation (PSP) is proposed to model the feature statistics of patches under different domain shifts, which enables the model to encode domain-invariant semantic features and improve generalization. PPP and PSP can help to diversify the source domain at the patch level and feature level. PDTrans learns context across diverse patches and takes advantage of self-attention to improve DG. Extensive experiments demonstrate the tremendous performance advantages of the PDTrans over state-of-the-art DG methods. Pei He, Licheng Jiao, Ronghua Shang, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Xiangrong Zhang, Shuang Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | A Complex-Former Tracker With Dynamic Polar Spatio-Temporal EncodingabstractRecently, the excellent performance of transformer has attracted the attention of the visual community. Visual transformer models usually reshape images into sequence format and encode them sequentially. However, it is difficult to explicitly represent the relative relationship in distance and direction of visual data with typical 2-D spatial structures. Also, the temporal motion properties of consecutive frames are hardly exploited when it comes to dynamic video tasks like tracking. Therefore, we propose a novel dynamic polar spatio-temporal encoding for video scenes. We use spiral functions in polar space to fully exploit the spatial dependences of distance and direction in real scenes. We then design a dynamic relative encoding mode for continuous frames to capture the continuous spatio-temporal motion characteristics among video frames. Finally, we construct a complex-former framework with the proposed encoding applied to video-tracking tasks, where the complex fusion mode (CFM) realizes the effective fusion of scenes and positions for consecutive frames. The theoretical analysis demonstrates the feasibility and effectiveness of our proposed method. The experimental results on multiple datasets validate that our method can improve tracker performance in various video scenarios. Licheng Jiao, Hao Zhu 0009, Zhongjian Huang, Fang Liu 0001, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Self-Supervised Self-Organizing Clustering Network: A Novel Unsupervised Representation Learning MethodabstractDeep learning-based clustering methods usually regard feature extraction and feature clustering as two independent steps. In this way, the features of all images need to be extracted before feature clustering, which consumes a lot of calculation. Inspired by the self-organizing map network, a self-supervised self-organizing clustering network ( [Formula: see text]OCNet) is proposed to jointly learn feature extraction and feature clustering, thus realizing a single-stage clustering method. In order to achieve joint learning, we propose a self-organizing clustering header (SOCH), which takes the weight of the self-organizing layer as the cluster centers, and the output of the self-organizing layer as the similarities between the feature and the cluster centers. In order to optimize our network, we first convert the similarities into probabilities which represents a soft cluster assignment, and then we obtain a target for self-supervised learning by transforming the soft cluster assignment into a hard cluster assignment, and finally we jointly optimize backbone and SOCH. By setting different feature dimensions, a Multilayer SOCHs strategy is further proposed by cascading SOCHs. This strategy achieves clustering features in multiple clustering spaces. [Formula: see text]OCNet is evaluated on widely used image classification benchmarks such as Canadian Institute For Advanced Research (CIFAR)-10, CIFAR-100, Self-Taught Learning (STL)-10, and Tiny ImageNet. Experimental results show that our method significant improvement over other related methods. The visualization of features and images shows that our method can achieve good clustering results. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Puhua Chen, Lingling Li 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Contrastive Learning-Based Dual Dynamic GCN for SAR Image Scene ClassificationabstractAs a typical label-limited task, it is significant and valuable to explore networks that enable to utilize labeled and unlabeled samples simultaneously for synthetic aperture radar (SAR) image scene classification. Graph convolutional network (GCN) is a powerful semisupervised learning paradigm that helps to capture the topological relationships of scenes in SAR images. While the performance is not satisfactory when existing GCNs are directly used for SAR image scene classification with limited labels, because few methods to characterize the nodes and edges for SAR images. To tackle these issues, we propose a contrastive learning-based dual dynamic GCN (DDGCN) for SAR image scene classification. Specifically, we design a novel contrastive loss to capture the structures of views and scenes, and develop a clustering-based contrastive self-supervised learning model for mapping SAR images from pixel space to high-level embedding space, which facilitates the subsequent node representation and message passing in GCNs. Afterward, we propose a multiple features and parameter sharing dual network framework called DDGCN. One network is a dynamic GCN to keep the local consistency and nonlocal dependency of the same scene with the help of a node attention module and a dynamic correlation matrix learning algorithm. The other is a multiscale and multidirectional fully connected network (FCN) to enlarge the discrepancies between different scenes. Finally, the features obtained by the two branches are fused for classification. A series of experiments on synthetic and real SAR images demonstrate that the proposed method achieves consistently better classification performance than the existing methods. Fang Liu 0001, Xiaoxue Qian, Licheng Jiao, Xiangrong Zhang, Lingling Li 0002, Yuanhao Cui |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | An Adaptive Migration Collaborative Network for Multimodal Image ClassificationabstractThe multispectral (MS) and the panchromatic (PAN) images belong to different modalities with specific advantageous properties. Therefore, there is a large representation gap between them. Moreover, the features extracted independently by the two branches belong to different feature spaces, which is not conducive to the subsequent collaborative classification. At the same time, different layers also have different representation capabilities for objects with large size differences. In order to dynamically and adaptively transfer the dominant attributes, reduce the gap between them, find the best shared layer representation, and fuse the features of different representation capabilities, this article proposes an adaptive migration collaborative network (AMC-Net) for multimodal remote-sensing (RS) images classification. First, for the input of the network, we combine principal component analysis (PCA) and nonsubsampled contourlet transformation (NSCT) to migrate the advantageous attributes of the PAN and the MS images to each other. This not only improves the quality of images themselves, but also increases the similarity between the two images, thereby reducing the representational gap between them and the pressure on the subsequent classification network. Second, for the interaction on the feature migrate branch, we design a feature progressive migration fusion unit (FPMF-Unit) based on the adaptive cross-stitch unit of correlation coefficient analysis (CCA), which can make the network automatically learn the features that need to be shared and migrated, aiming to find the best shared-layer representation for multifeature learning. And we design an adaptive layer fusion mechanism module (ALFM-Module), which can adaptively fuse features of different layers, aiming to clearly model the dependencies among multiple layers for different sized objects. Finally, for the output of the network, we add the calculation of the correlation coefficient to the loss function, which can make the network converge to the global optimum as much as possible. The experimental results indicate that AMC-Net can achieve competitive performance. And the code for the network framework is available at: https://github.com/ru-willow/A-AFM-ResNet. Wenping Ma 0001, Mengru Ma, Licheng Jiao, Fang Liu 0001, Hao Zhu 0009, Xu Liu 0006, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Robust and Effective: A Deep Matrix Factorization Framework for ClassificationabstractFor complex data, high dimension and high noise are challenging problems, and deep matrix factorization shows great potential in data dimensionality reduction. In this article, a novel robust and effective deep matrix factorization framework is proposed. This method constructs a dual-angle feature for single-modal gene data to improve the effectiveness and robustness, which can solve the problem of high-dimensional tumor classification. The proposed framework consists of three parts, deep matrix factorization, double-angle decomposition, and feature purification. First, a robust deep matrix factorization (RDMF) model is proposed in the feature learning, to enhance the classification stability and obtain better feature when faced with noisy data. Second, a double-angle feature (RDMF-DA) is designed by cascading the RDMF features with sparse features, which contains the more comprehensive information in gene data. Third, to avoid the influence of redundant genes on the representation ability, a gene selection method is proposed to purify the features by RDMF-DA, based on the principle of sparse representation (SR) and gene coexpression. Finally, the proposed algorithm is applied to the gene expression profiling datasets, and the performance of the algorithm is fully verified. Chenxi Tian, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | A Fast Evolutionary Knowledge Transfer Search for Multiscale Deep Neural ArchitectureabstractThe emergence of neural architecture search (NAS) algorithms has removed the constraints on manually designed neural network architectures, so that neural network development no longer requires extensive professional knowledge, trial and error. However, the extremely high computational cost limits the development of NAS algorithms. In this article, in order to reduce computational costs and to improve the efficiency and effectiveness of evolutionary NAS (ENAS) is investigated. In this article, we present a fast ENAS framework for multiscale convolutional networks based on evolutionary knowledge transfer search (EKTS). This framework is novel, in that it combines global optimization methods with local optimization methods for search, and searches a multiscale network architecture. In this article, evolutionary computation is used as a global optimization algorithm with high robustness and wide applicability for searching neural architectures. At the same time, for fast search, we combine knowledge transfer and local fast learning to improve the search speed. In addition, we explore a multiscale gray-box structure. This gray box structure combines the Bandelet transform with convolution to improve network approximation, learning, and generalization. Finally, we compare the architectures with more than 40 different neural architectures, and the results confirmed its effectiveness. Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Curvature-Balanced Feature Manifold Learning for Long-Tailed ClassificationabstractTo address the challenges of long-tailed classification, researchers have proposed several approaches to reduce model bias, most of which assume that classes with few samples are weak classes. However, recent studies have shown that tail classes are not always hard to learn, and model bias has been observed on sample-balanced datasets, suggesting the existence of other factors that affect model bias. In this work, we systematically propose a series of geometric measurements for perceptual manifolds in deep neural networks, and then explore the effect of the geometric characteristics of perceptual manifolds on classification difficulty and how learning shapes the geometric characteristics of perceptual manifolds. An unanticipated finding is that the correlation between the class accuracy and the separation degree of perceptual manifolds gradually decreases during training, while the negative correlation with the curvature gradually increases, implying that curvature imbalance leads to model bias. Therefore, we propose curvature regularization to facilitate the model to learn curvature-balanced and flatter perceptual manifolds. Evaluations on multiple long-tailed and non-long-tailed datasets show the excellent performance and exciting generality of our approach, especially in achieving significant performance improvements based on current state-of-the-art techniques. Our work opens up a geometric analysis perspective on model bias and reminds researchers to pay attention to model bias on non-long-tailed and even sample-balanced datasets. The code and model will be made public. Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006, Lingling Li 0002 |
CVPR | 3 |
| 2023 | Delving into Semantic Scale Imbalance
Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006 |
ICLR | 3 |
| 2023 | Spatial-Preserving and Edge-Orienting High-Resolution Network for Remote Sensing Change DetectionabstractRemote sensing change detection (RSCD), with a view to probing surface changes between bi-temporal images, makes a spurt of progress with the continuous innovation of deep learning. However, the extraction of multi-scale features and the detection of small domain of variation as well as the detail information in RSCD task still has large development space. Besides, current existing methods mostly focus on learning regional information but pay less regard to boundary identification, which leads to inaccurate detection results. Therefore, a spatial-preserving and edge-orienting high-resolution network is proposed to address the problems. In the overall architecture, a dual-branch encoder consists of a pyramid feature extracted branch and an enhanced HR network branch is designed to extract muti-scale bi-temporal features and small change objectives, while two edge-orienting modules (EOM) are embedded in order to utilize edge prior knowledge for further improving the accuracy of change detection. Moreover, spatial-preserving module (SPM) based on the self-attention calculation in spatial dimension is applied in the pyramid part to alleviate the poor location information of the high-level features. The experimental results demonstrate that the proposed network outperforms the cited state-of-the-art methods on LEVIR change detection datasets (LEVIR-CD). Fang Liu 0001, Jia Liu 0020, Liang Xiao 0001, Xu Tang 0004 |
IGARSS | 2 |
| 2023 | Domain-Specific and Domain-Common Feature Enhancement for Cross-Domain Few-Shot Hyperspectral Image ClassificationabstractThere is a small sample problem in hyperspectral image (HSI) classification task due to the difficulty of labeling samples. It is generally solved using a combination of few-shot learning and cross-domain method. In the paper, we propose a domain-specific and domain-common feature enhancement method for cross-domain few-shot HSI classification. It consists of a domain adaptation module and a feature enhancement module. The former is used to learn domain-specific features of both domains from the beginning of the network, and the latter is used to reduce domain differences by learning domain-common features through feature enhancement. The experimental results indicate that our proposed method performs better than the advanced classification methods. Wenfei Gao, Fang Liu 0001, Jia Liu 0020, Liang Xiao 0001, Xu Tang 0004 |
IGARSS | 2 |
| 2023 | Swin Resnetswin Transformers for Change Detection in Remote Sensing ImagesabstractThe change detection task of remote sensing images is a basic scientific problem, and has been further widely used in real life. Recently, transformer model has shown strong learning and representation abilities in visual interpretation. In this article, Inspired by the success of the Vision Transformer and its variants, we propose a novel change detection model for remote sensing images, named Swin ResNet Transformers (Swin ResNet). Different from other methods, the proposed Swin ResNet architecture uses a Swin transform encoder, which extracts feature representations of multiple resolutions through a shift window mechanism to calculate self-attention. On three datasets, the proposed model showed good performance, and demonstrate that the Swin transformer has a strong ability to learn long-term dependencies of multi-scale context representation. Xu Liu 0006, Yu Liu 0005, Licheng Jiao, Lingling Li 0002, Fang Liu 0001 |
IGARSS | 5 |
| 2023 | Unsupervised Domain Adaption for Remote Sensing Semantic Segmentation with Self-Attention MechanismabstractThe domain shift between the source and target domains limits the performance of traditional convolutional neural networks (CNNs) for feature extraction in remote sensing tasks. We propose an image translation network that uses generative adversarial networks (GANs) to transfer spectral distributions from training to test data, enhancing cross-domain semantic segmentation. Our approach fine-tunes the DeepLab-V3 framework on synthetic training data generated by the proposed network. Experimental results show improved performance in cross-domain semantic segmentation tasks for remote sensing images. Keming Liu, Fang Liu 0001, Jia Liu 0020, Liang Xiao 0001, Xu Tang 0004 |
IGARSS | 2 |
| 2023 | A Strong Vision Transformer Adapter with Adaptive Thresholding for fine-Grained Building ClassificationabstractFine-grained building classification provides a solid basis for the comparison of city morphologies and the investigation of urban planning. To this aim, the DFC23 establishes a large-scale and multi-modal benchmark for the classification of building roof types. However, the problems of long-tailed distribution, data insufficient, inter-class similarity, and intra-class difference severely inhibit the performance of the detector. In this work, we build a strong vision transformer adapter fine-tuned on the cropped building instances to enhance the capacity of feature extraction and design a cross-modal fusion (CMF) module to effectively aggregate features from RGB and SAR data. When transferring to building instance segmentation, we construct a robust training pipeline and a two-stage test-time results ensemble scheme. Furthermore, we introduce self-training with two key denoising techniques, global average filtering (GAF) and intra-class adaptive thresholding (IAT), to boost the generalization of the model. Experimental results show the effectiveness of our method, ranking 2nd in the test phase of the contest. Xiaoqiang Lu, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Yuting Yang 0008 |
IGARSS | 5 |
| 2023 | Trident Cooperation Network for Building Extraction and Height EstimationabstractBuilding extraction and height estimation provide solid fundamentals for reconstructing city morphologies and investigating urban planning. To this aim, the DFC23 establishes a large-scale and multi-modal benchmark for multi-task learning of building reconstruction. However, the problems of data limitation and fore-background confusion severely inhibit the performance of the model. In this work, we propose a novel trident cooperation network (TCNet) to perform end-to-end building extraction and height estimation using RGB and SAR data. Specifically, to enrich the feature representation and generalization of the shared backbone, we introduce a vision transformer adapter to inject vision-specific inductive biases and design a cross-modal fusion (CMF) module to effectively aggregate features from multi-modal data. For downstream visual tasks, we construct trident decoders including a detector, a lightweight MLP segmentation head, and a pixel-wise regression head. Moreover, to highlight the foreground object, we use the binary mask predicted by the MLP head to cooperate with the height estimation map predicted by the estimator. And the weighted sub-task losses are gathered to optimize our TCNet. Experimental results show the effectiveness of our method, ranking 2nd in the test phase of the contest. Xiaoqiang Lu, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Yuting Yang 0008 |
IGARSS | 5 |
| 2023 | Edge-Guided Feature Dense Fusion Network for Remote Sensing Image Change DetectionabstractRemote sensing change detection (CD) is of great importance to Earth observation. Recently, Deep Learning (DL) has been increasingly used to extract useful features and make accurate decisions in a large number of remote sensing images, due to its ability to automatically learn semantic features. However, insufficient fusion of bitemporal images and the lack of prior knowledge of edge structures in current DL methods will result in inaccurate CD results, especially for building boundaries. To alleviate these problems, an edge-guided feature-densely-fused network (EGFDFN) is proposed in this paper. In contrast to conventional Siamese networks, EGFDFN extracts bitemporal features from an extra dual decoder instead of a dual encoder to obtain more accurate change features. In addition, an attention and dense fusion module (ADFM) and an edge guidance module (EGM) are used to enhance features and make full use of edge information. Experimental results demonstrate that the proposed method outperforms on LEVIR-CD dataset among other representative methods. Hejun Luo, Jia Liu 0020, Fang Liu 0001, Jingxiang Yang, Liang Xiao 0001 |
IGARSS | 3 |
| 2023 | Orthogonal Uncertainty Representation of Data Manifold for Robust Long-Tailed LearningabstractIn scenarios with long-tailed distributions, the model's ability to identify tail classes is limited due to the under-representation of tail samples. Class rebalancing, information augmentation, and other techniques have been proposed to facilitate models to learn the potential distribution of tail classes. The disadvantage is that these methods generally pursue models with balanced class accuracy on the data manifold, while ignoring the ability of the model to resist interference. By constructing noisy data manifold, we found that the robustness of models trained on unbalanced data has a long-tail phenomenon. That is, even if the class accuracy is balanced on the data domain, it still has bias on the noisy data manifold. However, existing methods cannot effectively mitigate the above phenomenon, which makes the model vulnerable in long-tailed scenarios. In this work, we propose an Orthogonal Uncertainty Representation (hOUR) of feature embedding and an end-to-end training strategy to improve the long-tail phenomenon of model robustness. As a general enhancement tool, OUR has excellent compatibility with other methods and does not require additional data generation, ensuring fast and efficient training. Comprehensive evaluations on long-tailed datasets show that our method significantly improves the long-tail phenomenon of robustness, bringing consistent performance gains to other long-tailed learning methods. Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006, Lingling Li 0002 |
ACM Multimedia | 3 |
| 2023 | Task context transformer and GCN for few-shot learning of cross-domain
Pengfang Li, Fang Liu 0001, Licheng Jiao, Lingling Li 0002, Puhua Chen, Shuo Li 0010 |
Neurocomputing | 2 |
| 2023 | MinEnt: Minimum entropy for self-supervised representation learning
Shuo Li 0010, Fang Liu 0001, Zehua Hao, Licheng Jiao, Xu Liu 0006, Yuwei Guo 0001 |
Pattern Recognit. | 2 |
| 2023 | Knowledge transduction for cross-domain few-shot learning
Pengfang Li, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006 |
Pattern Recognit. | 2 |
| 2023 | Knowledge transfer evolutionary search for lightweight neural architecture with dynamic inference
Xiaoxue Qian, Fang Liu 0001, Licheng Jiao, Xiangrong Zhang, Shuo Li 0010, Puhua Chen, Xu Liu 0006 |
Pattern Recognit. | 2 |
| 2023 | Dual Wavelet Attention Networks for Image ClassificationabstractGlobal average pooling (GAP) plays an important role in traditional channel attention. However, there is the disadvantage of insufficient information to use the result of GAP as the channel scalar. At the same time, the existing spatial attention models focus on the areas of interest using average pooling or convolutional networks, but there is a loss of feature information and neglect of the structural feature. In this paper, dual wavelet attention is proposed, which can effectively alleviate the aforementioned problems and enhance the representation ability of CNNs. Firstly, the equivalence between the sum of the low-frequency subband coefficients of 2D DWT (Haar) and GAP is proved. On this basis, the statistical characteristics of low-frequency and high-frequency subbands are effectively combined to obtain the channel scalars, which can better measure the importance of each channel. In addition, 2D DWT can effectively capture the approximate and detailed structural features. Thus, wavelet spatial attention is proposed, which can effectively focus on the key spatial structural features. Different from traditional spatial attention, it can better curve the structural and spatial attention for different channels. The experiments are verified on four natural image data sets and three remote sensing scene classification data sets, which shows the effectiveness and versatility of the proposed methods. The code of this paper will be available athttps://github.com/yutinyang/DWAN. Yuting Yang 0008, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Lingling Li 0002, Puhua Chen, Xiufang Li, Zhongjian Huang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | DFAT: Dynamic Feature-Adaptive TrackingabstractDuring target tracking process, the state of the target is usually unpredictable. In theory, it is often beneficial to automatically assign suitable features to describe the specific target in each frame. Inspired by this, in this paper, we propose a novel dynamic feature-adaptive tracking framework (DFAT) which automatically assigns appropriate features to the consecutive frames during tracking process to boost the tracking performance. To implement DFAT, a large pool consisting of trackers/experts based on correlation filtering (CF) is constructed which is called candidate pool (CandPool). The diversity of the experts lies in their feature configurations and we call them candidate experts (CandExp). In this way, different features can be assigned for continuously changed scenarios and the target. Then to assign suitable experts, for each frame, we design the dynamic tracking process as the following three steps: (1) Several experts which are called executive experts (ExeExp) are selected from the CandPool according to CandExps’ past performance. (2) The ExeExps generate the tracking results and the performance of them are evaluated via a novel evaluation mechanism. (3) The selection rate of each CandExp in the CandPool is updated according to the performance evaluation and the final tracking result is selected. To better evaluate the CandExp, we propose two novel criteria: (1) content similarity weighted intra-evaluation, and (2) response confidence based self-evaluation. Compared with traditional post-event ensemble trackers that use fixed experts, the proposed method learns to dynamically assign appropriate ExeExps selected from a large CandPool which leads to adaption to different cases. Moreover, overfitting caused by fixed experts can also be mitigated via dynamic tracking. Experiments on both public available general and satellite videos based data sets demonstrate the superiority of the proposed method. Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Jia Liu 0020 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | A Collaborative Learning Tracking Network for Remote Sensing VideosabstractWith the increasing accessibility of remote sensing videos, remote sensing tracking is gradually becoming a hot issue. However, accurately detecting and tracking in complex remote sensing scenes is still a challenge. In this article, we propose a collaborative learning tracking network for remote sensing videos, including a consistent receptive field parallel fusion module (CRFPF), dual-branch spatial-channel co-attention (DSCA) module, and geometric constraint retrack strategy (GCRT). Considering the small-size objects of remote sensing scenes are difficult for general forward networks to extract effective features, we propose a CRFPF-module to establish parallel branches with consistent receptive fields to separately extract from shallow to deep features and then fuse hierarchical features adaptively. Since the objects and their background are difficult to distinguish, the proposed DSCA-module uses the spatial-channel co-attention mechanism to collaboratively learn the relevant information, which enhances the saliency of the objects and regresses to precise bounding boxes. Considering the interference of similar objects, we designed a GCRT-strategy to judge whether there is a false detection through the estimated motion trajectory and then recover the correct object by weakening the feature response of interference. The experimental results and theoretical analysis on multiple datasets demonstrate our proposed method's feasibility and effectiveness. Code and net are available at https://github.com/Dawn5786/CoCRF-TrackNet. Licheng Jiao, Hao Zhu 0009, Fang Liu 0001, Shuyuan Yang 0001, Xiangrong Zhang, Shuang Wang 0001, Rong Qu |
IEEE Trans. Cybern. | 4 |
| 2023 | Learning Salient Feature for Salient Object Detection Without LabelsabstractSupervised salient object detection (SOD) methods achieve state-of-the-art performance by relying on human-annotated saliency maps, while unsupervised methods attempt to achieve SOD by not using any annotations. In unsupervised SOD, how to obtain saliency in a completely unsupervised manner is a huge challenge. Existing unsupervised methods usually gain saliency by introducing other handcrafted feature-based saliency methods. In general, the location information of salient objects is included in the feature maps. If the features belonging to salient objects are called salient features and the features that do not belong to salient objects, such as background, are called nonsalient features, by dividing the feature maps into salient features and nonsalient features in an unsupervised way, then the object at the location of the salient feature is the salient object. Based on the above motivation, a novel method called learning salient feature (LSF) is proposed, which achieves unsupervised SOD by LSF from the data itself. This method takes enhancing salient feature and suppressing nonsalient features as the objective. Furthermore, a salient object localization method is proposed to roughly locate objects where the salient feature is located, so as to obtain the salient activation map. Usually, the object in the salient activation map is incomplete and contains a lot of noise. To address this issue, a saliency map update strategy is introduced to gradually remove noise and strengthen boundaries. The visualization of images and their salient activation maps show that our method can effectively learn salient visual objects. Experiments show that we achieve superior unsupervised performance on a series of datasets. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Xu Liu 0006, Puhua Chen |
IEEE Trans. Cybern. | 2 |
| 2023 | Fast and Effective: A Novel Sequential Single-Path Search for Mixed-Precision-Quantized NetworksabstractModel quantization can reduce the model size and computational latency, it has been successfully applied for many applications of mobile phones, embedded devices, and smart chips. Mixed-precision quantization models can match different bit precision according to the sensitivity of different layers to achieve great performance. However, it is difficult to quickly determine the quantization bit precision of each layer in deep neural networks under some constraints (for example, hardware resources, energy consumption, model size, and computational latency). In this article, a novel sequential single-path search (SSPS) method for mixed-precision model quantization is proposed, in which some given constraints are introduced to guide the searching process. A single-path search cell is proposed to combine a fully differentiable supernet, which can be optimized by gradient-based algorithms. Moreover, we sequentially determine the candidate precisions according to the selection certainties to exponentially reduce the search space and speed up the convergence of the searching process. Experiments show that our method can efficiently search the mixed-precision models for different architectures (for example, ResNet-20, 18, 34, 50, and MobileNet-V2) and datasets (for example, CIFAR-10, ImageNet, and COCO) under given constraints, and our experimental results verify that SSPS significantly outperforms their uniform-precision counterparts. Qigong Sun, Xiufang Li, Licheng Jiao, Yan Ren 0002, Fanhua Shang, Fang Liu 0001 |
IEEE Trans. Cybern. | 6 |
| 2023 | Multisource Joint Representation Learning Fusion Classification for Remote Sensing ImagesabstractMultisource remote sensing images provide complementary multidimensional information for reliable and accurate classification. However, gaps in imaging mechanisms result in heterogeneity between multiple source images. During fusion, this heterogeneity causes the generated multisource representations may be redundant and ignore discriminative uni-source information, which significantly hampers the fusion classification performance. To address this challenge, we introduce a novel multisource joint representation learning method for remote sensing image fusion classification, termed Multisource Information Bottleneck Fusion Network (MIBF-Net). Based on the Information Bottleneck principle, MIBF-Net employs mutual information constraints to effectively integrate multisource information, generating a comprehensive and non-redundant multisource representation. Specifically, MIBF-Net first introduces an attribution-driven noise adaptation layer to dynamically balance the speed of feature learning across sources for extracting discriminative uni-source intrinsic information. Furthermore, a cross-source relationship encoding module is designed to fully explore cross-source complex dependencies for enhancing the richness of fused representations. Finally, we design an information bottleneck fusion module to fuse uni-source semantic information and cross-source information while reducing redundancy. In particular, we employ variational inference techniques to effectively address the mutual information optimization problem and provide theoretical derivations. Extensive experimental results on three heterogeneous multisource remote sensing data benchmarks show that the model significantly outperforms the state-of-the-art methods. Xueli Geng, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001, Xiangrong Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Weak-to-Strong Consistency Learning for Semisupervised Image SegmentationabstractSupervised remote sensing (RS) image segmentation has achieved remarkable success with large amounts of manually labeled data, which may be difficult to acquire in some practical application scenarios. Semisupervised RS image segmentation can efficiently utilize the knowledge embedded in unlabeled data to improve recognition performance, which is of great significance for the generalization application of segmentation models. In this work, we propose an end-to-end semisupervised RS image segmentation method based on weak-to-strong consistency learning, denoted as WSCL. Specifically, a common strong data augmentation technique for image segmentation is introduced to provide powerful input perturbation to decouple self-biased cognition. By forcing weakly augmented, and strongly augmented perspectives from the same sample to be consistent, WSCL not only enables the model to steadily learn knowledge contained in unlabeled data but also alleviates overfitting. In addition, a novel sparse dual-view cross-sample image generation method is presented to generate new training samples, which helps provide a more comprehensive diversity of perturbations. Furthermore, an adaptive re-weighting strategy based on the entropy maps of the outputs of strongly perturbed samples is proposed to suppress noise, guiding the training process in a positive direction. Extensive experiments demonstrate the significant advantage of WSCL over other advanced methods, achieving new state-of-the-art under several evaluation metrics on DFC22, iSAID, MER, MSL, Vaihingen, and GID-15 datasets. The source code is open-sourced at https://github.com/xiaoqiang-lu/WSCL. Xiaoqiang Lu, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001, Zhixi Feng, Puhua Chen |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Multipretext-Task Prototypes Guided Dynamic Contrastive Learning Network for Few-Shot Remote Sensing Scene ClassificationabstractAs a content management technique, remote sensing (RS) scene classification (RSSC) always attracts researchers’ attention. In the past decades, many successful methods have been proposed. Nevertheless, their prerequisite is that there are large labeled data sets, which is a strict demand in practice. To resolve this contradiction, developing RSSC models with the help of few-shot learning (FSL) has become popular. Due to lacking prior knowledge, most of the existing few-shot RSSC models pay attention to the learning algorithm. However, they do not attach importance to the complex contents within RS scenes and the intricate inter-/intra-class relations between RS scenes. This would influence their performance negatively. In this paper, we propose a new few-shot RSSC model named multi-pretext-task prototypes guided dynamic contrastive learning network (MPCL-Net). MPCL-Net consists of a multi-pretext tasks generation sub-module, a deep feature learning sub-module, and a joint optimization sub-module. First, two RS-oriented pretext tasks are constructed under the self-supervised learning (SSL) framework in the multi-pretext tasks generation sub-module, which aim to explore multi-scale and rotation-invariant information from RS scenes. Second, a simple convolutional neural network (CNN) is developed in the deep feature learning sub-module to transform the RS scenes into visual features. Third, three loss functions are formulated and integrated in the joint optimization sub-module. Their goals are to fully capture the diverse land covers within RS scenes and compact/separate the intra-/inter-class samples with limited supervision. Finally, our MPCL-Net can be trained in a meta way. The positive results counted on the three public RS scene data sets confirm that our MPCL-Net is helpful to RSSC tasks under the few-shot scenario. Our source codes are available at https://github.com/TangXu-Group/Remote-Sensing-Images-Classification/tree/main/MPCL. Jingjing Ma 0001, Weiquan Lin, Xu Tang 0004, Xiangrong Zhang, Fang Liu 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | A Spatial-Spectral Bilinear Representation Fusion Network for Multimodal ClassificationabstractThe complementary and heterogeneous properties fusion of multimodal data (such as hyperspectral, lidar, and synthetic aperture radar data) can significantly improve the accuracy of remote sensing (RS) images joint classification. Thus, we propose a spatial-spectral bilinear representation fusion network (S2BRFNet), which captures long-range dependencies cross-modality and within the same modality to achieve the final joint classification. Firstly, a cross-modal spatial-spectral representation module (S2RM) is designed, it utilizes spatial-spectral attention and self-attention between heterogeneous data to enhance the characterization capabilities of cross-modal complementary properties and spatial-spectral features of single-source data. Secondly, a semantic space-guided bilinear feature fusion module (S2BFM) is developed, which uses deep and shallow features to regain fine-grained features. It uses shallow location details to improve the semantic prediction of deep features. Furthermore, it uses the different representation capabilities of different layers for objects with obvious feature differences to enhance the feature advantages. Therefore, rich global context information is obtained. Finally, the semantic space re-weight strategy is used to guide the outer product fusion of heterogeneous features, which enhances the ability of the network to identify similar features. Classification experiments are carried out on four common datasets of different modality combinations (HS-SAR-DSM Augsburg, Berlin, Trento, and Muufl), and this can prove the superiority of the S2BRFNet. Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Which Target to Focus on: Class-Perception for Semantic Segmentation of Remote SensingabstractDeep Learning-based (DL) methods have dominated the task of semantic segmentation of remote sensing images. However, the sizes of different objects vary widely, and there is a great deal of label-noise due to the inevitable shadows. Therefore, there is an urgent need for a method that can precisely handle complex ground data. In this paper, we propose an Inter-Class Enhanced Network (ICEN) for representing features of varying sizes. It comprises two branches: Sparse Representation Network (SPN) and Feature Extraction Network (FEN). Then, a Class-Perception Block is inserted between the two branches to instruct the SPN’s low-level semantic features to be merged into the deeper network. Such a block can reduce label-noise in remote sensing image segmentation. In addition, the proposed EIRI provides a more precise classification process for target edges containing many misclassified points without requiring excessive computational overhead. The experimental results of our proposed Class-Perception Network (C-PNet) achieve competitive performance on the Vaihingen, Potsdam, LoveDA, and UAVid datasets. Lingling Li 0002, Yilin Shao, Licheng Jiao, Xu Liu 0006, Puhua Chen, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | WNet: W-Shaped Hierarchical Network for Remote-Sensing Image Change DetectionabstractChange detection (CD) is a hot research topic in the remote sensing (RS) community. With the increasing availability of high-resolution (HR) RS images, there is a growing demand for CD models with high detection accuracy and generalization ability. In other words, the CD models are expected to work well for various HRRS images. Convolutional neural networks (CNNs) have been dominated in HRRS image CD due to their excellent information extraction and nonlinear fitting capabilities. However, they are not skilled in modeling long-range contexts hidden in HRRS images, which limits their performance in CD tasks more or less. Recently, the Transformer, which is good at extracting global context dependencies, has become popular in the RS community. Nevertheless, detailed local knowledge receives insufficient emphasis in common Transformers. Considering the above discussion, we combine CNN and Transformer and propose a new W-shaped dual Siamese branch hierarchical network for HRRS image CD named WNet. WNet first incorporates a Siamese CNN and a Siamese Transformer into a dual-branch encoder to extract multi-level local fine-grained features and global long-range contextual dependencies. Also, we introduce deformable ideas into the Siamese CNN and Transformer to make WNet understand the critical and irregular areas within HRRS images. Second, the difference enhancement module (DEM) is developed and embedded into the encoder to produce the difference feature maps at different levels. Using simple pixel-wise subtraction and channel-wise concatenation, the changes of interest and irrelevant changes can be highlighted and suppressed in a learnable manner. Next, the multi-level difference feature maps are fused stage by stage by CNN-Transformer fusion modules (CTFMs), which are the basic units of the decoder in WNet. In CTFM, the local, global, and cross-scale clues are taken into account to ensure the integrity of information. Finally, a simple classifier is constructed and added at the top of the decoder to predict the change maps. Positive experimental results counted on four public datasets demonstrate that the proposed WNet is helpful in HRRS image CD tasks. Our source codes are available at https://github.com/TangXu-Group/Remote-Sensing-Image-Change-Detection/tree/main/WNet. Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Cross-Domain Few-Shot Hyperspectral Image Classification With Class-Wise AttentionabstractFew-shot learning (FSL) is an effective method to solve the problem of hyperspectral image (HSI) classification with few labeled samples. It learns transferable knowledge from sufficient labeled auxiliary data to classify unseen classes with limited labeled samples for training. However, the distribution difference between auxiliary data and unseen classes results in the learned transferable knowledge not being well applied to the new task. Therefore, a class-wise attentive cross-domain FSL (CA-CFSL) framework is proposed in this article, in which a feature extractor is learned to extract data features with discriminability and domain invariance. The class-wise attention metric module (CAMM) introduces class-wise attention on the FSL framework to learn more discriminative features, which improves the interclass decision boundaries. Furthermore, an asymmetric domain adversarial module (ADAM) is designed to enhance the ability of extracting domain-invariant representations, which combines asymmetric adversarial training with embedded domain-specific information. Experimental results on four public HSI datasets demonstrate that the proposed method outperforms the existing methods. Wenzhen Wang, Fang Liu 0001, Jia Liu 0020, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | SDCDNet: A Semi-Dual Change Detection Network Framework With Super-Weak Label for Remote Sensing ImageabstractMost current change detection methods require a large amount of labeled data to train huge parameters. To break this limitation, this paper proposes a novel semi-supervised learning framework for remote sensing change detection, named a semi-dual change detection network (SDCDNet). The SDCDNet consists of a dual shared network and dual branching networks. The dual shared network is designed to exploit the full potential of the data, and the dual branching network is proposed to differentiate the kinds of annotated data and eliminate the disturbance between different types of data. In addition, the adaptive weighting module (AWM) enhances the features of weak branching, and the mask constraint module (MCM) is proposed to increase the ability of the network to extract foreground features. To solve the complex problem of data labeling, a patch-based weak label construction method is proposed to build super-weak labels. Experiments show that the proposed SDCDNet achieves excellent results on two remote sensing image change detection datasets. Jiahao Wang 0002, Fang Liu 0001, Hao Wang 0211, Xu Liu 0006, Licheng Jiao, Lingling Li 0002, Puhua Chen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | CSLT: Contourlet-Based Siamese Learning Tracker for Dim and Small Targets in Satellite VideosabstractMost popular visual trackers for natural scenarios always adopt handcraft features or deep features to track the target in a video. However, they face with difficulties in discriminative feature representation and usually suffer from severe model drift for satellite videos, especially when encountering challenges of dim and small targets, low contrast or similar target interference. To overcome these difficulties, we propose a Contourlet-based Siamese Learning Tracker (CSLT), which mainly aims at tracking dim and small objects in satellite videos. In contrast to conventional methods, the contourlet transform enriches directional multi-resolution information which is crucial to discriminative feature representation for dim and small targets in satellite video frames that lack distinguishable appearance features. We jointly use multi-resolution features with deep features by spatial-attention fusion strategy and then track the targets by a Siamese structure network. To further improve the accuracy and robustness, a model drift alarm and calibration module, including translation drifting penalty and rotation drifting penalty, is employed during tracking. We conduct extensive comparisons with 16 popular state-of-the-art trackers on three satellite video datasets. The experimental results validate the effectiveness of the proposed tracker. Yinan Wu 0001, Licheng Jiao, Fang Liu 0001, Zhaoliang Pi, Xu Liu 0006, Lingling Li 0002, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Multicue Contrastive Self-Supervised Learning for Change Detection in Remote SensingabstractContrastive self-supervised learning (CSSL) is a promising method in extracting effective features from unlabeled data. It performs well in image-level tasks, such as image classification and retrieval. However, the existing CSSL methods are not suitable for pixel-level tasks, e.g., change detection (CD), since they ignore the correlation between local patches or pixels. In this paper, we firstly propose a multi-cue contrastive self-supervised learning (MC-CSSL) method to derive dense features for change detection. Besides data augmentation, the MC-CSSL takes advantage of more cues based on the semantic meaning and temporal correlation of local patches. Specially, the positive pair is built from local patches with the similar semantic meaning or temporal ones with the same geographic location. The assumption is that local patches belonging to the same kind of land-covering tend to share similar features. Secondly, the affinity matrix is truncated and introduced to extract change information between two temporal patches obtained from different types of sensors. As a result, some initial unchanged pixels are selected to serve as the supervision for mapping the dense features into a consistent space. Based on the distance between all bi-temporal pixels in the consistent space, a difference image (DI) is generated and more unchanged pixels can be available. The dense feature mapping and unchanged pixel updating proceed alternately. The proposed CD method is evaluated in both homogeneous and heterogeneous cases and the experimental results demonstrate its effectiveness and priority after comparison with some existing state-of-the-art methods. The source code will be available at https://github.com/Yang202308/ChangeDetection_CSSL. Meijuan Yang, Licheng Jiao, Fang Liu 0001, Biao Hou, Shuyuan Yang 0001, Yake Zhang, Jianlong Wang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | An Explainable Spatial-Frequency Multiscale Transformer for Remote Sensing Scene ClassificationabstractDeep convolutional neural networks (CNNs) are significant in remote sensing. Due to the strong local representation learning ability, CNNs have excellent performance in remote sensing scene classification. However, CNNs focus on location-sensitive representations in the spatial domain and lack contextual information mining capabilities. Meanwhile, remote sensing scene classification still faces challenges, such as complex scenes and significant differences in target sizes. To address the problems and challenges above, more robust feature representation learning networks are necessary. In this paper, a novel and explainable spatial-frequency multi-scale Transformer framework, SF-MSFormer, is proposed for remote sensing scene classification. It mainly comprises spatial-domain and frequency-domain multi-scale Transformer branches, which consider the spatial-frequency global multi-scale representation features. Besides, the texture-enhanced encoder is designed in the frequency-domain multi-scale Transformer branch, which is adaptive to capture the global texture features. In addition, an adaptive feature aggregation module is designed to integrate the spatial-frequency multi-scale feature for final recognition. The experimental results verify the effectiveness of SF-MSFormer and show better convergence. It achieves state-of-the-art results (98.72%, 98.6%, 99.72%, and 94.83% overall accuracies, respectively) on the AID, UCM, WHU-RS19, and NWPU-RESISC45 datasets. Besides, the feature visualizations evaluate the explainability of the texture-enhanced encoder. The code implementation of this article will be available at https://github.com/yutinyang/SF-MSFormer. Yuting Yang 0008, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Boundary-Aware Multiscale Learning Perception for Remote Sensing Image SegmentationabstractFor remote sensing image segmentation, the boundaries of objects are difficult to distinguish, which is ignored by most methods. Therefore, it is challenging how to excavate and recover the boundaries of objects accurately. In this article, we propose a boundary-aware multi-scale network (BMNet) to solve this problem. The key components of BMNet include the scale attention module (SA-module) and boundary guidance module (BG-module). Specifically, SA-module is proposed to guide the refinement of multi-scale features in a context-aware way. It enhances the discriminability of multi-scale features by establishing contextual dependencies, which enables the refinement of the prediction of objects. Then, BG-module is proposed to enable networks to distinguish the boundary of objects. It utilizes manifold information of features to generate boundary guidance maps and forces the network to focus more on the boundary of objects. The effectiveness of the proposed BMNet is demonstrated on two public remote sensing datasets: ISPRS 2-D semantic labeling Potsdam dataset and Vaihingen dataset, where BMNet achieves better segmentation than prevalent methods. Finally, the experimental results indicate that BMNet can produce sharper boundaries of objects to reconstruct more detailed segmentation results. Chao You, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | A Spatial Hierarchical Reasoning Network for Remote Sensing Visual Question AnsweringabstractFor visual question answering on remote sensing (RSVQA), current methods scarcely consider geospatial objects typically with large-scale differences and positional sensitive properties. Besides, modeling and reasoning the relationships between entities have rarely been explored, which leads to one-sided and inaccurate answer predictions. In this article, a novel method called spatial hierarchical reasoning network (SHRNet) is proposed, which endows a remote sensing (RS) visual question answering (VQA) system with enhanced visual–spatial reasoning capability. Specifically, a hash-based spatial multiscale visual representation module is first designed to encode multiscale visual features embedded with spatial positional information. Then, spatial hierarchical reasoning is conducted to learn the high-order inner group object relations across multiple scales under the guidance of linguistic cues. Finally, a visual-question (VQ) interaction module is employed to learn an effective image–text joint embedding for the final answer predicting. Experimental results on three public RS VQA datasets confirm the effectiveness and superiority of our model SHRNet. Zixiao Zhang, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Fang Liu 0001, Yuxuan Li 0004, Zhicheng Guo |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Curvelet Adversarial Augmented Neural Network for SAR Image ClassificationabstractConvolutional neural networks (CNNs) have superior feature learning capabilities with large numbers of labeled samples. The reality is that labeling these samples is costly in terms of human labor. Existing data augmentation methods alleviate the scarcity of labeled samples. However, these methods are not suitable for synthetic aperture radar (SAR) images, owing to special imaging mechanisms and observational objects. The generative SAR images by existing augmented methods show structure distortion. To address this issue, we introduce a curvelet adversarial augmented neural network (CA2NN) for SAR image classification. Specifically, an$\text{A}^{2}$NN is established, which consists of two generative streams and one discriminative stream. In the generative stream, through the mutual transformation between the whole and partial images, more new samples with structural consistency are generated to augment the limited labeled data. In the discriminative stream, these generated samples show certain appearance variations after adversarial training based on the novel joint discriminant criterion. Simultaneously, given the multiscale and multidirectional nature of SAR images, we construct discretized curvelet in 2-D space, aiming to extract the singularity features and avoid overfitting. By integrating curvelet kernels into$\text{A}^{2}$NN, CA2NN can automatically generate more representative features adapting to complex terrain, while greatly reducing the complexity of the network. Experiments are conducted on the SAR images with large-scale and complex scenes, suggesting that the proposed approach significantly improves the classification performance with few labeled samples. Yake Zhang, Fang Liu 0001, Licheng Jiao, Shuyuan Yang 0001, Lingling Li 0002, Meijuan Yang, Jianlong Wang, Xu Liu 0006 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | GeoFormer: A Geometric Representation Transformer for Change DetectionabstractDeep representation learning has improved automatic remote change detection (RSCD) in recent years. Existing methods emphasize primarily convolutional neural networks (CNNs) or Transformer-based networks. However, most of them neither effectively combine CNNs and Transformer nor use prior geometric information to refine regions. In this paper, a novel geometric representation Transformer (GeoFormer) is proposed for high-resolution RSCD. GeoFormer utilizes convolutional information to guide the Transformer by employing geometric prior knowledge. Specifically, the proposed GeoFormer consists of three carefully designed components: the geometric-based Swin Transformer (Geo-Swin Transformer) encoder, the Laplace attention fusion (LAFusion) module, and the UNet++CD decoder. Firstly, Geo-Swin Transformer is a novel designed non-local Siamese encoder that combines geometric convolution with Transformer to provide local geometric representation information for remote contextual features. Then, a LAFusion module is proposed to achieve robust bi-temporal feature fusion, which is founded on attention mechanism and edge information. Finally, UNet++CD decodes fine-grained information from the fused features by dense multiscale upsampling process. Experimental results demonstrate that the proposed GeoFormer performs better than benchmark methods on four change detection datasets (LEVIR-CD, WHU-CD, DSIFN-CD, and CDD) and is able to detect the edges of change regions more precisely. Our code is available at https://github.com/Jiaxzhao/GeoFormer. Jiaxuan Zhao, Licheng Jiao, Chao Wang 0099, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | A Universal Quaternion Hypergraph Network for Multimodal Video Question AnsweringabstractFusion and interaction of multimodal features are essential for video question answering. Structural information composed of the relationships between different objects in videos is very complex, which restricts understanding and reasoning. In this paper, we propose a quaternion hypergraph network (QHGN) for multimodal video question answering, to simultaneously involve multimodal features and structural information. Since quaternion operations are suitable for multimodal interactions, four components of the quaternion vectors are applied to represent the multimodal features. Furthermore, we construct a hypergraph based on the visual objects detected in the video. Most importantly, the quaternion hypergraph convolution operator is theoretically derived to realize multimodal and relational reasoning. Question and candidate answers are embedded in quaternion space, and a Q&A reasoning module is creatively designed for selecting the answer accurately. Moreover, the unified framework can be extended to other video-text tasks with different quaternion decoders. Experimental evaluations on the TVQA dataset and DramaQA dataset show that our method achieves state-of-the-art performance. Zhicheng Guo, Jiaxuan Zhao, Licheng Jiao, Xu Liu 0006, Fang Liu 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | D³K: Dynastic Data-Free Knowledge DistillationabstractData-free knowledge distillation further broadens the applications of the distillation model. Nevertheless, the problem of providing diverse data with rich expression patterns needs to be further explored. In this paper, a novel dynastic data-free knowledge distillation ($D^{3}K$) model is proposed to alleviate this problem. In this model, a dynastic supernet generator (D-SG) with a flexible network structure is proposed to generate diverse data. The D-SG can adaptively alter architectural configurations and activate different subnet generators in different sequential iteration spaces. The variable network structure increases the complexity and capacity of the generator, and strengthens its ability to generate diversified data. In addition, a novel additive constraint based on the differentiable dhash (D-Dhash) is designed to guide the structure parameter selection of the D-SG. This constraint forces the D-SG to constantly jump out of the fixed generation mode and generate diverse data in semantics and instance. The effectiveness of the proposed model is verified on the experimental benchmark datasets (MNIST, CIFAR-10, CIFAR-100, and SVHN). Xiufang Li, Qigong Sun, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Yi Zuo 0003 |
IEEE Trans. Multim. | 4 |
| 2023 | Transformer Based Conditional GAN for Multimodal Image FusionabstractMultimodal Image fusion is becoming urgent in multi-sensor information utilization. However, existing end-to-end image fusion frameworks ignore a priori knowledge integration and long-distance dependencies across domains, which brings challenges to the network convergence and global image perception in complex scenes. In this paper, a conditional generative adversarial network with transformer (TCGAN) is proposed for multimodal image fusion. The generator is to generate a fused image with the source images content. The discriminators are adopted to distinguish the differences between the fused image and the source images. Adversarial training makes the final fused image to maintain the structural and textural details in the cross-modal images simultaneously. In particular, a wavelet fusion module makes the inputs contain image content from different domains as much as possible. The extracted convolutional features interact in the multiscale cross-modal transformer fusion module to fully complement the associated information. It makes the generator to focus on both local and global context. TCGAN fully considers the training efficiency of the adversarial process and the integrated retention of redundant information. Various experimental results of TCGAN have highlighted targets, rich details, and fast convergence properties on public datasets. Jun Zhang 0045, Licheng Jiao, Wenping Ma 0001, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Multiscale Curvelet Scattering NetworkabstractFeature representation has received more and more attention in image classification. Existing methods always directly extract features via convolutional neural networks (CNNs). Recent studies have shown the potential of CNNs when dealing with images' edges and textures, and some methods have been explored to further improve the representation process of CNNs. In this article, we propose a novel classification framework called the multiscale curvelet scattering network (MSCCN). Using the multiscale curvelet-scattering module (CCM), image features can be effectively represented. There are two parts in MSCCN, which are the multiresolution scattering process and the multiscale curvelet module. According to multiscale geometric analysis, curvelet features are utilized to improve the scattering process with more effective multiscale directional information. Specifically, the scattering process and curvelet features are effectively formulated into a unified optimization structure, with features from different scale levels being efficiently aggregated and learned. Furthermore, a one-level CCM, which can essentially improve the quality of feature representation, is constructed to be embedded into other existing networks. Extensive experimental results illustrate that MSCCN achieves better classification accuracy when compared with state-of-the-art techniques. Eventually, the convergence, insight, and adaptability are evaluated by calculating the trend of loss function's values, visualizing some feature maps, and performing generalization analysis. Jie Gao 0013, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou, Xu Liu 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Deep Learning in Visual Tracking: A ReviewabstractDeep learning (DL) has made breakthroughs in many computer vision tasks and also in visual tracking. From the beginning of the research on the automatic acquisition of high abstract feature representation, DL has gone deep into all aspects of tracking to date, to name a few, similarity metric, data association, and bounding box estimation. Also, pure DL-based trackers have obtained the state-of-the-art performance after the community's constant research. We believe that it is time to comprehensively review the development of DL research in visual tracking. In this article, we overview the critical improvements brought to the field by DL: deep feature representations, network architecture, and four crucial issues in visual tracking (spatiotemporal information integration, target-specific classification, target information update, and bounding box estimation). The scope of the survey of DL-based tracking covers two primary subtasks for the first time, single-object tracking and multiple-object tracking. Also, we analyze the performance of DL-based approaches and give meaningful conclusions. Finally, we provide several promising directions and tasks in visual tracking and relevant fields. Licheng Jiao, Yidong Bai, Puhua Chen, Fang Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Learning Social Spatio-Temporal Relation Graph in the Wild and a Video BenchmarkabstractSocial relations are ubiquitous and form the basis of social structure in our daily life. However, existing studies mainly focus on recognizing social relations from still images and movie clips, which are different from real-world scenarios. For example, movie-based datasets define the task as the video classification, only recognizing one relation in the scene. In this article, we aim to study the problem of social relation recognition in an open environment. To close the gap, we provide the first video dataset collected from real-life scenarios, named social relation in the wild (SRIW), where the number of people can be huge and vary, and each pair of relations needs to be recognized. To overcome new challenges, we propose a spatio-temporal relation graph convolutional network (STRGCN) architecture, utilizing correlative visual features to recognize social relations intuitively. Our method decouples the task into two classification tasks: person-level and pair-level relation recognition. Specifically, we propose a person behavior and character module to encode moving and static features in two explicit ways. Then we take them as node features to build a relation graph with meaningful edges in a scene. Based on the relation graph, we introduce the graph convolutional network (GCN) and local GCN to encode social relation features which are used for both recognitions. Experimental results demonstrate the effectiveness of the proposed framework, achieving 83.1% and 40.8% mAP in person-level and pair-level classification. Moreover, the study also contributes to the practicality in this field. Haoran Wang 0008, Licheng Jiao, Fang Liu 0001, Lingling Li 0002, Xu Liu 0006, Deyi Ji, Weihao Gan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | RDLNet: A Regularized Descriptor Learning NetworkabstractLocal image descriptor learning has been instrumental in various computer vision tasks. Recent innovations lie with similarity measurement of descriptor vectors with metric learning for randomly selected Siamese or triplet patches. Local image descriptor learning focuses more on hard samples since easy samples do not contribute much to optimization. However, few studies focus on hard samples of image patches from the perspective of loss functions and design appropriate learning algorithms to obtain a more compact descriptor representation. This article proposes a regularized descriptor learning network (RDLNet) that makes the network focus on the learning of hard samples and compact descriptor with triplet networks. A novel hard sample mining strategy is designed to select the hardest negative samples in mini-batch. Then batch margin loss concerned with hard samples is adopted to optimize the distance of extreme cases. Finally, for a more stable network and preventing network collapsing, orthogonal regularization is designed to constrain convolutional kernels and obtain rich deep features. RDLNet provides a compact discriminative low-dimensional representation and can be embedded in other pipelines easily. This article gives extensive experimental results for large benchmarks in multiple scenarios and generalization in matching applications with significant improvements. Jun Zhang 0045, Licheng Jiao, Wenping Ma 0001, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Hao Zhu 0009 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Self-Training Multi-Sequence Learning with Transformer for Weakly Supervised Video Anomaly DetectionabstractWeakly supervised Video Anomaly Detection (VAD) using Multi-Instance Learning (MIL) is usually based on the fact that the anomaly score of an abnormal snippet is higher than that of a normal snippet. In the beginning of training, due to the limited accuracy of the model, it is easy to select the wrong abnormal snippet. In order to reduce the probability of selection errors, we first propose a Multi-Sequence Learning (MSL) method and a hinge-based MSL ranking loss that uses a sequence composed of multiple snippets as an optimization unit. We then design a Transformer-based MSL network to learn both video-level anomaly probability and snippet-level anomaly scores. In the inference stage, we propose to use the video-level anomaly probability to suppress the fluctuation of snippet-level anomaly scores. Finally, since VAD needs to predict the snippet-level anomaly scores, by gradually reducing the length of selected sequence, we propose a self-training strategy to gradually refine the anomaly scores. Experimental results show that our method achieves significant improvements on ShanghaiTech, UCF-Crime, and XD-Violence. Shuo Li 0010, Fang Liu 0001, Licheng Jiao |
AAAI | 2 |
| 2022 | Unsupervised Few-Shot Image Classification by Learning Features into Clustering Space
Shuo Li 0010, Fang Liu 0001, Zehua Hao, Kaibo Zhao 0001, Licheng Jiao |
ECCV (31) | 2 |
| 2022 | A Dual-Fusion Semantic Segmentation Framework with Gan for SAR ImagesabstractDeep learning based semantic segmentation is one of the popular methods in remote sensing image segmentation. In this paper, a network based on the widely used encoder-decoder architecture is proposed to accomplish the synthetic aperture radar (SAR) images segmentation. With the better representation capability of optical images, we propose to enrich SAR images with generated optical images via the generative adversative network (GAN) trained by numerous SAR and optical images. These optical images can be used as expansions of original SAR images, thus ensuring robust result of segmentation. Then the optical images generated by the GAN are stitched together with the corresponding real images. An attention module following the stitched data is used to strengthen the representation of the objects. Experiments indicate that our method is efficient compared to other commonly used methods. Jia Liu 0020, Fang Liu 0001, Andi Zhang 0003, Wenfei Gao, Jiao Shi |
IGARSS | 3 |
| 2022 | Remote Sensing Image Change Detection Based on Deep Dictionary LearningabstractAs a hot topic in the field of remote sensing (RS), change detection aims to identify the semantic change between bitemporal RS images. Due to the semantic complexity of RS images, how to accurately detect the semantic change has become a challenging problem. Recently, many deep-based methods are proposed to solve this issue. However, ignoring the representation difference of same semantics in different periods limits their performance, such as river is liquid in summer and solid in winter. Therefore, a new method is presented, named dictionary learning based change detector (DLCDet), which consists of feature pyramid network, deep dictionary learning and dual supervision modules. In DLCDet, the deep dictionary learning is proposed to reduce the representation difference so that DLCDet identifies the potential semantic change more accurately. Experiments are conducted on two public datasets change detection dataset (CDD) and building change detection dataset (BCDD), which demonstrates the effectiveness of the proposed method. Yuqun Yang, Xu Tang 0004, Fang Liu 0001, Jingjing Ma 0001, Licheng Jiao |
IGARSS | 3 |
| 2022 | Hierarchical Scene Normality-Binding Modeling for Anomaly Detection in Surveillance VideosabstractAnomaly detection in surveillance videos is an important topic in the multimedia community, which requires efficient scene context extraction and the capture of temporal information as a basis for decision. From the perspective of hierarchical modeling, we parse the surveillance scene from global to local and propose a Hierarchical Scene Normality-Binding Modeling framework (HSNBM) to handle anomaly detection. For the static background hierarchy, we design a Region Clustering-driven Multi-task Memory Autoencoder (RCM-MemAE), which can simultaneously perform region segmentation and scene reconstruction. The normal prototypes of each local region are stored, and the frame reconstruction error is subsequently amplified by global memory augmentation. For the dynamic foreground object hierarchy, we employ a Scene-Object Binding Frame Prediction module (SOB-FP) to bind all foreground objects in the frame with the prototypes stored in the background hierarchy according their positions, thus fully exploit the normality relationship between foreground and background. The bound features are then fed into the decoder to predict the future movement of the objects. With the binding mechanism between foreground and background, HSNBM effectively integrates the "reconstruction" and "prediction" tasks and builds a semantic bridge between the two hierarchies. Finally, HSNBM fuses the anomaly scores of the two hierarchies to make a comprehensive decision. Extensive empirical studies on three standard video anomaly detection datasets demonstrate the effectiveness of the proposed HSNBM framework. Qianyue Bao, Fang Liu 0001, Yang Liu 0349, Licheng Jiao, Xu Liu 0006, Lingling Li 0002 |
ACM Multimedia | 2 |
| 2022 | MRIQA: Subjective Method and Objective Model for Magnetic Resonance Image Quality AssessmentabstractMagnetic Resonance Imaging (MRI) is widely used for medical diagnosis, staging and follow-up of disease. However, MRI images may have artifacts due to various reasons such as patient movement or machine distortion, which may be unintentionally introduced during the procedure of medical image acquisition, processing, etc. These artifacts may affect the effectiveness of diagnosis or even cause false diagnosis. To solve this problem, we propose a general medical image quality assessment (MIQA) methodology, including subjective MIQA procedures and objective MIQA algorithms. We further apply this methodology to MRI images in this paper due to its widespread use in practical applications. We first establish a magnetic resonance imaging quality assessment (MRIQA) database, which contains 3809 MRI images. Then a subjective image quality assessment experiment is conducted by expert doctors according to the diagnostic value of these images, which split all MRI images into 1285 low quality images and 2524 high quality images. We then conduct a baseline deep learning experiment, and propose an attention based MIQANet model to automatically separate MRI images into high quality and low quality based on their diagnosis value. Our proposed method achieves a great quality assessment accuracy of 96.59%. The constructed MRIQA database and proposed MIQA model will be public available to further promote medical IQA research. Fang Liu 0001, Huiyu Duan, Xiongkuo Min, Guangtao Zhai |
VCIP | 2 |
| 2022 | Augmentative contrastive learning for one-shot object detection
Yaoyang Du, Fang Liu 0001, Licheng Jiao, Zehua Hao, Shuo Li 0010, Xu Liu 0006, Jing Liu 0006 |
Neurocomputing | 2 |
| 2022 | Region NMS-based deep network for gigapixel level pedestrian detection with two-step cropping
Lingling Li 0002, Xiaohui Guo, Jingjing Ma 0001, Licheng Jiao, Fang Liu 0001, Xu Liu 0006 |
Neurocomputing | 6 |
| 2022 | Entire Deformable ConvNets for semantic segmentation
Bingqi Yu, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Xu Tang 0004 |
Knowl. Based Syst. | 5 |
| 2022 | Deep Multiview Union Learning Network for Multisource Image ClassificationabstractWith the development of the imaging technology of various sensors, multisource image classification has become a key challenge in the field of image interpretation. In this article, a novel classification method, called the deep multiview union learning network (DMULN), is proposed to classify multisensor data. First, an associated feature extractor is designed to process the multisource data by canonical correlation analysis (CCA) in the head of the network. Second, an improved deep learning architecture with two branches is presented to extract high-level view features from the associated features. Third, a novel pooling, called view union pooling, is proposed to fuse the multiview feature from the deep model. Finally, the fused feature is fed into the classifier. The proposed framework is easy to optimize since it is an end-to-end network. Extensive experiments and analysis on the datasets IEEE_grss_dfc_2017 and IEEE_grss_dfc_2018 show that the proposed method achieves comparable results. Our results demonstrate that abundant multisource information can improve the classification performance. Xu Liu 0006, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Cybern. | 5 |
| 2022 | GAFnet: Group Attention Fusion Network for PAN and MS Image High-Resolution ClassificationabstractPanchromatic (PAN) and multispectral (MS) images have coordinated and paired spatial spectral information, which can complement each other and make up for their shortcomings for image interpretation. In this article, a novel classification method called the deep group spatial-spectral attention fusion network is proposed for PAN and MS images. First, the MS image is processed by unpooling to obtain the same resolution as that of the PAN image. Second, the group spatial attention and group spectral attention modules are proposed to extract image features. The PAN and the processed MS images are regarded as the input of the two modules, respectively. Third, the features from the previous step are fused by the attention fusion module, which aims to fully fuse multilevel features, take into account both the low-level features and the high-level features, and maintain the global abstract and local detailed information of the pixels. Finally, the fusion feature is fed into the classifier and the resulting map is obtained by pixel level. Extensive experiments and analysis on four datasets show that the proposed method achieves comparable results. Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Biao Hou, Shuyuan Yang 0001, Licheng Jiao |
IEEE Trans. Cybern. | 3 |
| 2022 | Automatic Graph Learning Convolutional Networks for Hyperspectral Image ClassificationabstractThe excellent performance of graph convolutional networks (GCNs) on non-Euclidean data has drawn widespread attention from the hyperspectral image classification (HSIC) community, where the predefined graph (including node modeling and adjacency matrix calculation) plays a key role. However, existing GCN-based methods rely on manual efforts in constructing and updating graphs, and the superpixel-based node features lack high-level semantics. In this article, we propose an automatic graph learning convolutional network (Auto-GCN), which unifies the graph learning and HSIC in a “network-in-network” manner. Specifically, the graph is employed to model the interaction of the high-order tensors. Considering the powerful learning and representation capabilities of convolutional neural networks (CNNs), the semisupervised Siamese network (SiamNet) is embedded into GCNs and HSIC networks to accomplish the automatic learning and dynamic updating of the graph. GCNs further encode and infer the dynamic graph, and then, the learnable graph reprojection matrix is designed to assign graph representations to pixels. The dynamic graph serves the HSIC task during forward propagation, while the HSIC task continuously corrects the graph during backward propagation. Therefore, the “automatic” of the proposed Auto-GCN is not only reflected in the fact that the graph representation is designed and updated by an end-to-end network but is also HSIC task-oriented. The experimental results show that the proposed Auto-GCN outperforms other state-of-the-art methods on four publicly available hyperspectral datasets. Jie Chen 0098, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Polarimetric Multipath Convolutional Neural Network for PolSAR Image ClassificationabstractScatter targets of complex land covers in polarimetric synthetic aperture radar (PolSAR) images are often randomly oriented and cause randomly fluctuating echoes, which brings a challenge to PolSAR image classification. Therefore, many existing methods have alleviated this problem through orientation compensation. However, there are still two obstacles that limit the improvement of classification accuracy. On the one hand, generally, these methods process PolSAR images with fixed polarization rotation angles, which is experience-dependent and inflexible. On the other hand, for the different land covers of a PolSAR image, the existing methods do not consider these rotation angles separately. For the first obstacle, we design a group of convolution kernels called polarization rotation kernels (PRKs) and utilize them to build the polarimetric convolutional neural network (CNN) (PolCNN). The PolCNN is the base network of our final model, and it can learn polarization rotation angles adaptively. For the second obstacle, we extend the PolCNN into a multipath structure, the final model polarimetric multipath CNN (PolMPCNN). The polarization rotation angles of different land covers are directly related to the networks of different paths within the PolMPCNN. Furthermore, we also put forward the two-scale sampling and the stagewise training algorithm in order that our PolMPCNN can fit different scales of PolSAR targets and pays more attention to difficult training samples. Experiments on real PolSAR images show that the proposed model achieves the best classification results with an extremely low sampling rate of 0.1%. Yuanhao Cui, Fang Liu 0001, Licheng Jiao, Yuwei Guo 0001, Xuefeng Liang, Lingling Li 0002, Shuyuan Yang 0001, Xiaoxue Qian |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Deep Shearlet Network for Change Detection in SAR ImagesabstractConvolutional neural networks (CNN) can extract shift-invariant features, and have been widely applied in change detection task. However, common CNN lacks noise robustness and needs supervised data, to alleviate these problems, in this paper, we propose a novel deep shearlet network (ShearNet) for change detection in SAR images. In the network, a shearlet denoising layer (SDL) is designed to enhance the representation ability of common CNN. In SDL, feature maps are decomposed into subband coefficients by shearlet transform (ST). Due to optimal sparse representation property and highly direction sensitivity of ST, the network can capture important geometric information. Then, hard-threshold shrinkage is applied to high frequency subbands to drop small coefficients that are most likely to be noise, so that reduce the effect of noise. Finally, ShearNet is trained by introducing a noise-robust loss with noisy labels. The noisy labels are obtained by deep clustering that shows more robustness than existing preclassification methods. This fine-tuning process novelly follows the paradigm of learning from noisy labels to aside the difficulty of precisely labeling samples. Our experimental results on multiple real SAR datasets show that ShearNet can boost accuracy, and have better applicability for change detection in SAR images. The source code is available at https://github.com/yizhilanmaodhh/ShearNet. Huihui Dong, Licheng Jiao, Wenping Ma 0001, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | A Multiscale Self-Attention Deep Clustering for Change Detection in SAR ImagesabstractSynthetic aperture radar (SAR) image change detection (CD) is an important application in the field of remote sensing. Due to the lack of labeled data especially in the pixelwise task, it is urgent to develop unsupervised techniques to effectively detect changes. In this article, we propose a novel unsupervised representation learning framework for CD in SAR images, called multiscale self-attention (SA) deep clustering based on octave convolution. The main motivation is that a convolutional neural network (CNN) has the ability to extract significant feature hidden in input images, but it relies heavily on annotated data. Clustering is typically free from supervision; however, SAR images always suffer from speckle noise, which is unfriendly for clustering. Thus, we integrate unsupervised clustering with CNN to learn clustering-friendly feature representations. In the unified framework, CNN feature learning and clustering can be optimized end-to-end without supervision. To better suppress speckle noise and boost the joint optimization for distinguishing changes and unchanges, we use the K-means++ algorithm that is robust to noise as the clustering algorithm. In the meanwhile, we introduce the octave convolution and SA mechanism into the network to fully mine important spatial structure information for enhancing noise resistance of the network. Moreover, multiscale fusion modules are proposed to fuse multiscale input into a complementary feature representation that contains more context and semantic information around each pixel so that it refines the difference feature extraction while reducing speckle noise. Experiments on challenging SAR data sets demonstrate the effectiveness and potential of the proposed model compared with the current state-of-the-art algorithms. Huihui Dong, Wenping Ma 0001, Licheng Jiao, Fang Liu 0001, Lingling Li 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Adaptive Fuzzy Learning Superpixel Representation for PolSAR Image ClassificationabstractThe increasing applications of polarimetric synthetic aperture radar (PolSAR) image classification demand for effective superpixels’ algorithms. Fuzzy superpixels’ algorithms reduce the misclassification rate by dividing pixels into superpixels, which are groups of pixels of homogenous appearance and undetermined pixels. However, two key issues remain to be addressed in designing a fuzzy superpixel algorithm for PolSAR image classification. First, the polarimetric scattering information, which is unique in PolSAR images, is not effectively used. Such information can be utilized to generate superpixels more suitable for PolSAR images. Second, the ratio of undetermined pixels is fixed for each image in the existing techniques, ignoring the fact that the difficulty of classifying different objects varies in an image. To address these two issues, we propose a polarimetric scattering information-based adaptive fuzzy superpixel (AFS) algorithm for PolSAR images classification. In AFS, the correlation between pixels’ polarimetric scattering information, for the first time, is considered through fuzzy rough set theory to generate superpixels. This correlation is further used to dynamically and adaptively update the ratio of undetermined pixels. AFS is evaluated extensively against different evaluation metrics and compared with the state-of-the-art superpixels’ algorithms on three PolSAR images. The experimental results demonstrate the superiority of AFS on PolSAR image classification problems. Yuwei Guo 0001, Licheng Jiao, Rong Qu, Zhuangzhuang Sun, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | Multitask Semantic Boundary Awareness Network for Remote Sensing Image SegmentationabstractIn remote sensing images, boundary information plays a crucial role in land-cover segmentation. However, it is a challenging problem that sufficiently extracts complete and sharp boundaries from complex very-high-resolution (VHR) remote sensing images. To tackle this problem, we propose a semantic boundary awareness network (SBANet). The SBANet captures refined boundary information of land covers in feature extraction and then supervises its learning with a designed boundary loss. The key of SBANet includes boundary attention module (BA-module) and adaptive weights of multitask learning (AWML). The BA-module is proposed to capture land-cover boundary information from hierarchical features aggregation in a bottom-up manner. It emphasizes useful boundary information and relieves noise information in low-level features with the guidance of high-level features. To directly learn the boundary information, AWML adds a boundary loss to the original semantic loss by an adaptive fusion manner. This multitask learning enables the semantic information and the boundary information to work collaboratively and promote each other. Note that the BA-module and AWML are plug-and-play. Experimental results demonstrate the effectiveness of the proposed SBANet on the available ISPRS 2-D semantic labeling Potsdam and Vaihingen data sets. The SBANet also achieves the state-of-the-art performance in terms of overall accuracy (OA) and mean$F_{1}$score (m-$F_{1}$). Aijin Li, Licheng Jiao, Hao Zhu 0009, Lingling Li 0002, Fang Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Simple and Efficient: A Semisupervised Learning Framework for Remote Sensing Image Semantic SegmentationabstractSemantic segmentation based on deep learning has achieved impressive results in recent years, but these results are supported by a large amount of labeled data which requires intensive annotation at the pixel level, particularly for high-resolution remote sensing (RS) images. In this work, we propose a simple yet efficient semisupervised learning framework based on linear sampling self-training, named LSST, to improve the performance of RS image semantic segmentation. Specifically, the classical pseudo-labeling-based self-training paradigm is enhanced by injecting strong data augmentations (SDA) applicable to RS images, based on which a powerful baseline is constructed. Nevertheless, the problem of insufficient data training to generate pseudo-labels with a high level of noise persists, and the noisy pseudo-labels will continue to accumulate and impede model improvement during the re-training phase. Previous works commonly employ a pre-defined threshold to remove noise, but it will lead to overfitting the model to easily identified classes. To address it, a method using linear sampling (LS) is presented for assigning thresholds to different classes in an adaptive manner, which provides noiseless regions for re-training. Experiments prove that the proposed pixel-wise selection is more available for segmentation than image-level selection in RS images. Finally, LSST achieves state-of-the-art on several datasets and different evaluation metrics. The source code of the this paper is available at https://github.com/xiaoqiang-lu/LSST. Xiaoqiang Lu, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006, Zhixi Feng, Lingling Li 0002, Puhua Chen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Transfer Representation Learning Meets Multimodal Fusion Classification for Remote Sensing ImagesabstractTo maximize the complementary advantages of synergistic multimodal, a transfer representation learning fusion network (TRLF-Net) is proposed for multisource remote sensing images collaborative classification in this article. First, with respect to the feature encoding, we design a dual-branch attention sparse transfer module (DAST-Module), which combines the spatial and channel attention (CA) masks to migrate the advantage attributes of the panchromatic (PAN) and the MS images mutually. This not only enhances their respective image advantages but also facilitates the sparse fusion of low-level features. Second, for the separation of multiscale information, a deep dual-scale decomposition module (DDSD-Module) is designed, which allows the decompose of high-frequency and low-frequency components. Then it uses the decomposed information to make the essential difference as small as possible, and the surrounding contour difference is as large as possible of the complementary multimodal image through the design of the loss function. Finally, to address the problem of large intraclass and small interclass differences, we develop a representation fusion of the global and local features’ module (RFGAL-Module). It mainly adopts global features to sort local features within classes, and then outputs them in a cascade. Thus, the characterization ability of features is improved, and the global and local features are used in a coordinated manner to accomplish the sample classification tasks. In particular, the experimental results demonstrate that TRLF-Net can obtain much improved accuracy and efficiency. The code is accessible in:https://github.com/ru-willow/SRLF-Net. Mengru Ma, Wenping Ma 0001, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Very Low-Resolution Moving Vehicle Detection in Satellite VideosabstractThis paper proposes a practical end-to-end neural network framework to detect tiny moving vehicles in satellite videos with low imaging quality. Some instability factors such as illumination changes, motion blurs, and low contrast to the cluttered background make it difficult to distinguish true objects from noise and other point-shaped distractors. Moving vehicle detection in satellite videos can be carried out based on background subtraction or frame differencing. However, these methods are prone to produce lots of false alarms and miss many positive targets. Appearance-based detection can be an alternative but is not well-suited since classifier models are of weak discriminative power for the vehicles in top view at such low resolution. This article addresses these issues by integrating motion information from adjacent frames to facilitate the extraction of semantic features and incorporating the Transformer to refine the features for key points estimation and scale prediction. Our proposed model can well identify the actual moving targets and suppress interference from stationary targets or background. The experiments and evaluations using satellite videos show that the proposed approach can accurately locate the targets under weak feature attributes and improve the detection performance in complex scenarios. Zhaoliang Pi, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Biao Hou, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Hybrid Network With Structural Constraints for SAR Image Scene ClassificationabstractData-based image classification methods, such as convolutional neural networks (CNNs), have achieved state-of-the-art performance. They usually leverage thousands of labeled samples to train the networks but ignore some prior knowledge. However, labeled samples are difficult to be obtained for synthetic aperture radar (SAR) images. Model-based methods are adept at utilizing the prior information of data, while they have to introduce some restrictions or assumptions during the realization of models. Consequently, to develop the advantages of both methods and improve their disadvantages, we propose a hybrid network by coupling the data-based with model-based methods for SAR image scene classification in this article. First, to fully use the prior information of SAR images and large amounts of unlabeled samples, we improve the$G^{0}$-based variational Bayesian inference model (GVBI) and construct a$G^{0}$-based convolutional variational auto-encoder (GCVAE) for unsupervised learning of the distributional characteristics of SAR images. After that, we further extend the GCVAE by combining it with CNN, resulting in a stronger hybrid network to classify SAR images with a few labeled samples. In addition, considering the abundant structural information is crucial for SAR image classification, we design a sketch fitter and two structural constraints on both pixel and sketch spaces to assist the hybrid network to improve its classification performance. Finally, we evaluate the performance of our method on real-SAR images, and the experimental results demonstrate that the proposed framework outperforms related methods on classification while reducing the manual annotation substantially. Xiaoxue Qian, Fang Liu 0001, Licheng Jiao, Xiangrong Zhang, Puhua Chen, Lingling Li 0002, Yuanhao Cui |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Joint Siamese Attention-Aware Network for Vehicle Object Tracking in Satellite VideosabstractRemote sensing object tracking is a novel and challenging problem due to the negative effects of weak features and background noise. In this paper, from the perspective of attention-focus deep learning, we propose a Joint Siamese Attention-Aware Network (JSANet) for efficient remote sensing tracking which contains both self-attention and cross-attention modules. First, the self-attention modules we propose emphasize the interdependent channel-wise coefficient via channel attention and conduct corresponding space transformation of spatial domain information with spatial attention. Second, the cross-attention is designed to aggregate rich contextual interdependencies between the siamese branches via channel attention and excavate association produces reliable correspondence with spatial attention. In addition, a composite feature combine strategy is designed to fuse multiple attention features. Experimental results on the Jilin-1 satellite video datasets demonstrate that the proposed JSANet achieves state-of-the-art performance in terms of precision and success rate, demonstrate the effectiveness of the proposed methods. Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | EMTCAL: Efficient Multiscale Transformer and Cross-Level Attention Learning for Remote Sensing Scene ClassificationabstractIn recent years, convolutional neural network (CNN)-based methods have been widely used for remote sensing (RS) scene classification tasks and achieved excellent results. However, CNNs are not good at exploring contextual information, which is essential for fully understanding RS scenes. A new model named transformer attracts researchers’ attention to address this problem, which is skilled in mining the latent contextual information in RS scenes. Nevertheless, since the contents of RS scenes are diverse in type and various in scale, the performance of the original transformer in RS scene classification cannot reach what we expect. In addition, due to the specific self-attention mechanism, the time costs of the transformer are high, which hinders its practicability in the RS community. To overcome the above limitations, we propose a new model named efficient multi-scale transformer and cross-level attention learning (EMTCAL) for RS scene classification in this paper. EMTCAL combines the advantages of CNN and transformer to mine information within RS scenes fully. First, it uses a multi-layer feature extraction module (MFEM) to acquire global visual features and multi-level convolutional features from RS scenes. Second, a contextual information extraction module (CIEM) is proposed to capture rich contextual information from multi-level features. In CIEM, taking the characteristics of RS scenes and the computational complexity into account, we propose an efficient multi-scale transformer (EMST). EMST can mine the abundant knowledge with various scales hidden in RS scenes and model their inherent relations at small-time costs. Third, a cross-level attention module (CLAM) is developed to aggregate and explore correlations of multi-level features. Finally, a class score fusion module (CSFM) is designed to integrate the contributions of global and aggregated multi-level features for the discriminative scene representations. Extensive experiments are conducted on three public RS scene data sets. The positive results demonstrate that our EMTCAL can achieve superior classification performance and outperform many state-of-the-art methods. Xu Tang 0004, Mingteng Li, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | LHNet: Laplacian Convolutional Block for Remote Sensing Image Scene ClassificationabstractRecently, many state-of-the-art results for remote sensing image scene classification have been achieved by convolutional neural networks (CNNs) due to their large learning capability. However, in the forward process of CNNs, the high-frequency/texture features are gradually blurred with hierarchical down-sampling and convolution operations. High-frequency features are important to capture the diversity within a class and the similarity between classes. For example, the line features are crucial to distinguish a tennis court from a basketball court. For tennis court in different scenes, the highlight of line features can effectively avoid the influence of diverse background. As a consequence, we propose a Laplacian high-frequency convolutional block (LHCB) based on CNN to extract useful high-frequency features by trainable Laplacian operator. To propagate high-frequency features, we embed LHCB into the existing CNN structures and obtain LHNet. In LHNet, there are two pathways. The original CNN architecture can be taken as the low-frequency pathway and we propose a high-frequency pathway based on LHCB that propagates the residual high-frequency features blurred in each low-frequency layer. Considering that the high-frequency features usually show large variance between images of the same class, we propose a new objective for high-frequency pathway to enhance the intra-class similarity of high-frequency features. The final objective function is obtained by combining the new objective and the baseline classification objective. Numerous experiments on three public available remote sensing image scene classification data sets NWPU-RESISC45, AID and UC Mercerd demonstrate the superior performance of the proposed method. Licheng Jiao, Fang Liu 0001, Jia Liu 0020, Zhen Cui 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | MBLT: Learning Motion and Background for Vehicle Tracking in Satellite VideosabstractRecently, satellite videos provide a new way to dynamically monitor the Earth’s surface. The interpretation of satellite videos has attracted more and more attentions. In this article, we focus on the problem of the vehicle tracking in satellite videos. Satellite videos usually own a lower resolution, which leads to the following phenomena: 1) the size of a vehicle target usually includes a few pixels and 2) vehicles are usually with similar appearance which easily results in the wrong tracking within the observing region. General popular tracking methods usually focus on the representation of the target and recognize it from background which are limited in this problem. As a consequence, in this article, we propose to learn motion and background of the target in order to help the trackers recognize the target with higher accuracy. A prediction network is proposed to predict the location probability of the target in each pixel in next frame based on fully convolutional network (FCN) which is learned from previous results. In addition, a segmentation method is introduced to generate the feasible region for target in each frame and assign high probability for such a region. For quantitative comparison, we manually annotate 20 representative vehicle targets from nine satellite videos taken by JiLin-1. In addition, we also selected two public satellite video datasets for experiments. Numerous experimental results demonstrate the superior of the proposed method. Licheng Jiao, Fang Liu 0001, Lingling Li 0002, Xu Liu 0006, Jia Liu 0020 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Sparse Feature Clustering Network for Unsupervised SAR Image Change DetectionabstractIn this article, we propose a sparse feature clustering network (SFCNet) for change detection in synthetic aperture radar (SAR) images. One of the principal problems in dealing with SAR images is to reduce the impact of speckle noise. Therefore, based on a neural network framework for change detection, we introduce the multiobjective sparse feature learning (MO-SFL) model where the sparsity of representation is adaptively learned in order to increase the robustness to different levels of noise. For learning the semantic information of changed and unchanged pixels, the network is fine-tuned by the correctly labeled samples selected from coarse results. The selection criterion influences the change detection result a lot. Therefore, we construct a novel cross-entropy clustering loss (CEC) by introducing a clustering regularization term to learn the discriminative representations. Experiments on simulate and real SAR images demonstrate the superiority of the proposed method over compared methods. Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Jia Liu 0020 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Adaptive Dual-Path Collaborative Learning for PAN and MS ClassificationabstractDue to the limitation of sensor technology, researchers tend to obtain high-quality image information from panchromatic (PAN) images and multispectral (MS) images with different resolutions. Therefore, the classification of remote sensing images of PAN and MS have become a research hotspot. In this paper, we propose an adaptive dual-path collaborative learning method for PAN and MS classification. In the stage of sample generation and training, we propose an adaptive neighborhood sample grading (ANSG) strategy in the establishing sample stage so that each pixel to be classified can obtain neighborhood information suitable for itself. Further, to simulate biological cognitive mechanisms, we divide the samples into different levels, and design the self-paced progressive loss (SPL), thus allowing the network to do preference training in different stages. The network’s training can quickly reach the optimal of the current stage and the overall convergence is more thorough. In the network structure, we propose a dual-path module (DPM) to effectively alleviate the gradient degradation in theresidual path, while ensuring maximum gradient loss information flow between every two layers in thedensely connected path. This module can extract more robust features to cope with the complex characteristics of remote-sensing images. Moreover, using the characteristics of the dual path to better fuse the features by the gradual collaborative fusion (GCF) way. The experimental results and theoretical analysis have demonstrated the proposed approach’s effectiveness, feasibility, and robustness. Our model are available at https://github.com/AIpy-nan/DBFI-Net. Hao Zhu 0009, Kenan Sun, Licheng Jiao, Fang Liu 0001, Biao Hou, Shuang Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | MFNet: A Novel GNN-Based Multi-Level Feature Network With Superpixel PriorsabstractSince the superpixel segmentation method aggregates pixels based on similarity, the boundaries of some superpixels indicate the outline of the object and the superpixels provide prerequisites for learning structural-aware features. It is worthwhile to research how to utilize these superpixel priors effectively. In this work, by constructing the graph within superpixel and the graph among superpixels, we propose a novel Multi-level Feature Network (MFNet) based on graph neural network with the above superpixel priors. In our MFNet, we learn three-level features in a hierarchical way: from pixel-level feature to superpixel-level feature, and then to image-level feature. To solve the problem that the existing methods cannot represent superpixels well, we propose a superpixel representation method based on graph neural network, which takes the graph constructed by a single superpixel as input to extract the feature of the superpixel. To reflect the versatility of our MFNet, we apply it to an image-level prediction task and a pixel-level prediction task by designing different prediction modules. An attention linear classifier prediction module is proposed for image-level prediction tasks, such as image classification. An FC-based superpixel prediction module and a Decoder-based pixel prediction module are proposed for pixel-level prediction tasks, such as salient object detection. Our MFNet achieves competitive results on a number of datasets when compared with related methods. The visualization shows that the object boundaries and outline of the saliency maps predicted by our proposed MFNet are more refined and pay more attention to details. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Puhua Chen, Xu Liu 0006, Lingling Li 0002 |
IEEE Trans. Image Process. | 2 |
| 2022 | Coarse-to-Fine Contrastive Self-Supervised Feature Learning for Land-Cover Classification in SAR Images With Limited Labeled DataabstractContrastive self-supervised learning (CSSL) has achieved promising results in extracting visual features from unlabeled data. Most of the current CSSL methods are used to learn global image features with low-resolution that are not suitable or efficient for pixel-level tasks. In this paper, we propose a coarse-to-fine CSSL framework based on a novel contrasting strategy to address this problem. It consists of two stages, one for encoder pre-training to learn global features and the other for decoder pre-training to derive local features. Firstly, the novel contrasting strategy takes advantage of the spatial structure and semantic meaning of different regions and provides more cues to learn than that relying only on data augmentation. Specifically, a positive pair is built from two nearby patches sampled along the direction of the texture if they fall into the same cluster. A negative pair is generated from different clusters. When the novel contrasting strategy is applied to the coarse-to-fine CSSL framework, global and local features are learned successively by forcing the positive pair close to each other and the negative pair apart in an embedding space. Secondly, a discriminant constraint is incorporated into the per-pixel classification model to maximize the inter-class distance. It makes the classification model more competent at distinguishing between different categories that have similar appearance. Finally, the proposed method is validated on four SAR images for land-cover classification with limited labeled data and substantially improves the experimental results. The effectiveness of the proposed method is demonstrated in pixel-level tasks after comparison with the state-of-the-art methods. Meijuan Yang, Licheng Jiao, Fang Liu 0001, Biao Hou, Shuyuan Yang 0001, Yake Zhang, Jianlong Wang |
IEEE Trans. Image Process. | 3 |
| 2022 | Adaptive Contourlet Fusion Clustering for SAR Image Change DetectionabstractIn this paper, a novel unsupervised change detection method called adaptive Contourlet fusion clustering based on adaptive Contourlet fusion and fast non-local clustering is proposed for multi-temporal synthetic aperture radar (SAR) images. A binary image indicating changed regions is generated by a novel fuzzy clustering algorithm from a Contourlet fused difference image. Contourlet fusion uses complementary information from different types of difference images. For unchanged regions, the details should be restrained while highlighted for changed regions. Different fusion rules are designed for low frequency band and high frequency directional bands of Contourlet coefficients. Then a fast non-local clustering algorithm (FNLC) is proposed to classify the fused image to generate changed and unchanged regions. In order to reduce the impact of noise while preserve details of changed regions, not only local but also non-local information are incorporated into the FNLC in a fuzzy way. Experiments on both small and large scale datasets demonstrate the state-of-the-art performance of the proposed method in real applications. Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Jia Liu 0020 |
IEEE Trans. Image Process. | 3 |
| 2022 | New Generation Deep Learning for Video Object Detection: A SurveyabstractVideo object detection, a basic task in the computer vision field, is rapidly evolving and widely used. In recent years, deep learning methods have rapidly become widespread in the field of video object detection, achieving excellent results compared with those of traditional methods. However, the presence of duplicate information and abundant spatiotemporal information in video data poses a serious challenge to video object detection. Therefore, in recent years, many scholars have investigated deep learning detection algorithms in the context of video data and have achieved remarkable results. Considering the wide range of applications, a comprehensive review of the research related to video object detection is both a necessary and challenging task. This survey attempts to link and systematize the latest cutting-edge research on video object detection with the goal of classifying and analyzing video detection algorithms based on specific representative models. The differences and connections between video object detection and similar tasks are systematically demonstrated, and the evaluation metrics and video detection performance of nearly 40 models on two data sets are presented. Finally, the various applications and challenges facing video object detection are discussed. Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou, Lingling Li 0002, Xu Tang 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | DPFL-Nets: Deep Pyramid Feature Learning Networks for Multiscale Change DetectionabstractDue to the complementary properties of different types of sensors, change detection between heterogeneous images receives increasing attention from researchers. However, change detection cannot be handled by directly comparing two heterogeneous images since they demonstrate different image appearances and statistics. In this article, we propose a deep pyramid feature learning network (DPFL-Net) for change detection, especially between heterogeneous images. DPFL-Net can learn a series of hierarchical features in an unsupervised fashion, containing both spatial details and multiscale contextual information. The learned pyramid features from two input images make unchanged pixels matched exactly and changed ones dissimilar and after transformed into the same space for each scale successively. We further propose fusion blocks to aggregate multiscale difference images (DIs), generating an enhanced DI with strong separability. Based on the enhanced DI, unchanged areas are predicted and used to train DPFL-Net in the next iteration. In this article, pyramid features and unchanged areas are updated alternately, leading to an unsupervised change detection method. In the feature transformation process, local consistency is introduced to constrain the learned pyramid features, modeling the correlations between the neighboring pixels and reducing the false alarms. Experimental results demonstrate that the proposed approach achieves superior or at least comparable results to the existing state-of-the-art change detection methods in both homogeneous and heterogeneous cases. Meijuan Yang, Licheng Jiao, Fang Liu 0001, Biao Hou, Shuyuan Yang 0001, Meng Jian |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | MRTA: Multi-Resolution Training Algorithm for Multitemporal Semantic Change DetectionabstractThe multitemporal semantic change detection challenge track (Track MSD) in the 2021 Data Fusion Contest is to extract the land cover changes of the US state of Maryland from 2013 to 2017, but only low-resolution label data is provided. We present a multi-resolution training algorithm (MRTA) to alleviate the overfitting of the model on the coarse labels. First, using low-resolution coarse labels to train FCN, the average IoU can reach 0.5253. Generating pseudo-labels using this network, they are combined with coarse labels to form a multi-resolution label combination, and perform iterative fine-tuning and retrain. After that, by analyzing the loss and gain indicators of specific categories, it was found that the discrimination effect of the water area was poor, so a strong classifier was trained for the water area. We also implemented strategies such as model voting and weighted training to improve model performance. Finally, our method achieves 0.6445 mIoU on the test set, ranking 3rd in Track 2 of the IEEE Data Fusion Contest of 2021. Qianyue Bao, Yang Liu 0349, Zixiao Zhang, Dafan Chen, Yuting Yang 0008, Licheng Jiao, Fang Liu 0001 |
IGARSS | 7 |
| 2021 | DO-UNet, DO-LinkNet: UNet, D-LinkNet with DO-Conv for the Detection of Settlements without Electricity ChallengeabstractIn this paper, two semantic segmentation models, DO-UNet and DO-LinkNet, are presented for the detection of human settlements, and a threshold-based model is proposed to detect areas with electricity. In DO-UNet and DO-LinkNet, the conventional convolutional layer is replaced with depthwise over-parameterized convolutional layer. Also, an extra pooling operation is carried out in the last layer since the size of the input images is different from that of the labels. Depthwise over-parameterized convolutional layer enhances the convolutional layer with an additional depthwise convolution. Pooling operation can accelerate training speed, increase the receptive field in feature extraction, and reduce the requirement of network complexity. In the detection of settlements without electricity challenge track, our best F1-score on the validation set and the test set are 0.8820 and 0.8798, respectively. Ruoxian Feng, Xuanming Zhang, Jun Zhang 0045, Licheng Jiao, Xu Liu 0006, Fang Liu 0001 |
IGARSS | 7 |
| 2021 | Multisource Data Fusion for the Detection of Settlements Without ElectricityabstractThe international charity SolarAid aims to provide access to lights in areas without electricity, and it is a challenge to accurately and efficiently transmit the lights to the areas in need. Multisource, multitemporal, and multimodal remote sensing images can provide rich information about the target area, so using multisource remote sensing images for accurate detection of human settlements without electricity is a feasible solution. In this paper two separate detection tasks are formulated: building two attention SENet for settlements detection and light detection using the Sentinel-2 dataset and the Suomi Visible Infrared Imaging Radiometer Suite (VIIRS) night time dataset, respectively. In addition, we study a new outlier removal method based on the pixel distribution characteristics of the VIIRS dataset for data pre-processing, and propose a post-processing method based on region continuity for further correction of the results. Experiments show that our method can maximize the use of multisource data information and rank first in the detection of settlements without electricity challenge track (Track DSE) of the 2021 IEEE GRSS Data Fusion Contest. Yanbiao Ma, Kexin Feng, Xueli Geng, Licheng Jiao, Fang Liu 0001, Yuting Yang 0008 |
IGARSS | 6 |
| 2021 | Deep associative learning for neural networks
Jia Liu 0020, Fang Liu 0001, Liang Xiao 0001 |
Neurocomputing | 3 |
| 2021 | Ridgelet-Nets With Speckle Reduction Regularization for SAR Image Scene ClassificationabstractWith powerful feature representations, convolutional neural networks (CNNs) have produced tremendous achievements in image classification tasks and, typically, entail millions of labeled samples to train massive parameters. However, the sample labeling of synthetic aperture radar (SAR) images is extremely difficult, especially pixelwise labels, and has, sometimes, required field trips to accomplish labeling. Moreover, the inherent speckle noise may weaken the ability of networks to extract effective features from SAR images. In this article, we address these issues by labeling a few patchwise samples and propose Ridgelet-Nets with speckle reduction regularization for SAR image scene classification by combining deep learning with multiscale geometric analysis and statistical modeling of SAR images. First, we design Ridgelet-Nets with convolutional kernels constructed by ridgelet filters to reduce the training parameters and learn more discriminative features. Then, we embed speckle reduction regularization in the Ridgelet-Nets to restrain the influence of speckle noise and smooth the classification maps, in which the prior information of SAR image statistical modeling is introduced. Finally, we propose an adaptive SAR image scene classification framework based on an extended hierarchical visual semantic model, considering the differences in the structures and spatial relationships of different regions in the SAR images, particularly large-scale and complex scenes. Experimental results on real SAR images demonstrate that the proposed framework can achieve preferable classification performance using very limited labeled samples. Xiaoxue Qian, Fang Liu 0001, Licheng Jiao, Xiangrong Zhang, Yuwei Guo 0001, Xu Liu 0006, Yuanhao Cui |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Selective Adversarial Adaptation-Based Cross-Scene Change Detection Framework in Remote Sensing ImagesabstractSupervised change detection methods always face a big challenge that the current scene (target domain) is fully unlabeled. In remote sensing, it is common that we have sufficient labels in another scene (source domain) with a different but related data distribution. In this article, we try to detect changes in the target domain with the help of the prior knowledge learned from multiple source domains. To achieve this goal, we propose a change detection framework based on selective adversarial adaptation. The adaptation between multisource and target domains is fulfilled by two domain discriminators. First, the first domain discriminator regards each scene as an individual domain and is designed for identifying the domain to which each input sample belongs. According to the output of the first domain discriminator, a subset of important samples is selected from multisource domains to train a deep neural network (DNN)-based change detection model. As a result, not only the positive transfer is enhanced but also the negative transfer is alleviated. Second, as for the second domain discriminator, all the selected samples are thought from one domain. Adversarial learning is introduced to align the distributions of the selected source samples and the target ones. Consequently, it further adapts the knowledge of change from the source domain to the target one. At the fine-tuning stage, target samples with reliable labels and the selected source ones are used to jointly fine-tune the change detection model. As the target domain is fully unlabeled, homogeneity- and boundary-based strategies are exploited to make the pseudolabels from a preclassification map reliable. The proposed method is evaluated on three SAR and two optical data sets, and the experimental results have demonstrated its effectiveness and superiority. Meijuan Yang, Licheng Jiao, Biao Hou, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Residual Spectral-Spatial Attention Network for Hyperspectral Image ClassificationabstractIn the last five years, deep learning has been introduced to tackle the hyperspectral image (HSI) classification and demonstrated good performance. In particular, the convolutional neural network (CNN)-based methods for HSI classification have made great progress. However, due to the high dimensionality of HSI and equal treatment of all bands, the performance of these methods is hampered by learning features from useless bands for classification. Moreover, for patchwise-based CNN models, equal treatment of spatial information from the pixel-centered neighborhood also hinders the performance of these methods. In this article, we propose an end-to-end residual spectral-spatial attention network (RSSAN) for HSI classification. The RSSAN takes raw 3-D cubes as input data without additional feature engineering. First, a spectral attention module is designed for spectral band selection from raw input data by emphasizing useful bands for classification and suppressing useless bands. Then, a spatial attention module is designed for the adaptive selection of spatial information by emphasizing pixels from the same class as the center pixel or those are useful for classification in the pixel-centered neighborhood and suppressing those from a different class or useless. Second, two attention modules are also used in the following CNN for adaptive feature refinement in spectral-spatial feature learning. Third, a sequential spectral-spatial attention module is embedded into a residual block to avoid overfitting and accelerate the training of the proposed model. Experimental studies demonstrate that the RSSAN achieved superior classification accuracy compared with the state of the art on three HSI data sets: Indian Pines (IN), University of Pavia (UP), and Kennedy Space Center (KSC). Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Jianing Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | IPGN: Interactiveness Proposal Graph Network for Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) Detection is an important task to understand how humans interact with objects. Most of the existing works treat this task as an exhaustive triplet 〈 human, verb, object 〉 classification problem. In this paper, we decompose it and propose a novel two-stage graph model to learn the knowledge of interactiveness and interaction in one network, namely, Interactiveness Proposal Graph Network (IPGN). In the first stage, we design a fully connected graph for learning the interactiveness, which distinguishes whether a pair of human and object is interactive or not. Concretely, it generates the interactiveness features to encode high-level semantic interactiveness knowledge for each pair. The class-agnostic interactiveness is a more general and simpler objective, which can be used to provide reasonable proposals for the graph construction in the second stage. In the second stage, a sparsely connected graph is constructed with all interactive pairs selected by the first stage. Specifically, we use the interactiveness knowledge to guide the message passing. By contrast with the feature similarity, it explicitly represents the connections between the nodes. Benefiting from the valid graph reasoning, the node features are well encoded for interaction learning. Experiments show that the proposed method achieves state-of-the-art performance on both V-COCO and HICO-DET datasets. Haoran Wang 0008, Licheng Jiao, Fang Liu 0001, Lingling Li 0002, Xu Liu 0006, Deyi Ji, Weihao Gan |
IEEE Trans. Image Process. | 3 |
| 2021 | C-CNN: Contourlet Convolutional Neural NetworksabstractExtracting effective features is always a challenging problem for texture classification because of the uncertainty of scales and the clutter of textural patterns. For texture classification, spectral analysis is traditionally employed in the frequency domain. Recent studies have shown the potential of convolutional neural networks (CNNs) when dealing with the texture classification task in the spatial domain. In this article, we try combining both approaches in different domains for more abundant information and proposed a novel network architecture named contourlet CNN (C-CNN). The network aims to learn sparse and effective feature representations for images. First, the contourlet transform is applied to get the spectral features from an image. Second, the spatial-spectral feature fusion strategy is designed to incorporate the spectral features into CNN architecture. Third, the statistical features are integrated into the network by the statistical feature fusion. Finally, the results are obtained by classifying the fusion features. We also investigated the behavior of the parameters in contourlet decomposition. Experiments on the widely used three texture data sets (kth-tips2-b, DTD, and CUReT) and five remote sensing data sets (UCM, WHU-RS, AID, RSSCN7, and NWPU-RESISC45) demonstrate that the proposed approach outperforms several well-known classification methods in terms of classification accuracy with fewer trainable parameters. Mengkun Liu, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | Deep Adaptive Proposal Network in Optical Remote Sensing Images Objective DetectionabstractIt is difficult to distinguish complicated distribution characteristics of objects, which limits the performance of two-stage detectors in the field of optical remote sensing images object detection. In this paper, we propose a deep adaptive proposal network (DAPNet), which is a new category prior network (CPN) on the basis of the existing faster region convolutional neural network (Faster RCNN) architecture. The adaptive candidate boxes for each image is obtained by combining the candidate regions and the object number, which are generated by the fine-region proposal network (F-RPN) and the CPN respectively. These adaptive candidate boxes can satisfy the detection tasks in sparse and dense scenes. A set of experimental results verify the superiority of the proposed approach. Lingling Li 0002, Xiaohui Guo, Xu Liu 0006, Licheng Jiao, Fang Liu 0001 |
IGARSS | 6 |
| 2020 | Feature Correlation Analysis of Two-Branch Convolutional Networks for Multi-Source Image ClassificationabstractWith the development of multi-sensor imaging technology, constructing a multi-branch network model has become an important requirement for fusion decision. In the literature, many two-branch networks are proposed to interpret multisensor data and also get the satisfactory results. In this paper, we divide these models into two types and study them by mining changes in data and features. The main method used is analysis the correlation between the features of the same layer from the first branch and the second branch. The task of dual-source image classification serves as a means of experimentation. When the classification network is trained, the features of each layer are extracted for experimental analysis. Extensive experiments and analysis on the dataset IEEE_grss_dfc_2017 show that the analysis is meaningful. It is found that the correlation of the features from the low level to the high level is more and more consistent, a quantitative analysis is given in this paper. Xu Liu 0006, Licheng Jiao, Fang Liu 0001 |
IGARSS | 3 |
| 2020 | Remote Sensing Images Feature Learning based On Multi-Branch NetworksabstractRemote sensing (RS) images feature learning, plays a crucial role in many RS images application, and attracts scholars' attention. However, since RS images contain complex contents, how to extract robust features that can fully represent RS images becomes an important and tough task. In this paper, we develop a feature learning method based on multi-branch networks, named M-Net, which consists of fine-grained branch and coarse branch. Considering the objects within RS images are diverse in type and resolution, the fine-grained branch is developed to capture rich object-level information. First, the RS images convolutional features are extracted by fine-grained branch. Second, through encoding the score maps which can highlight the important regions, the fine-grained structure mapping are obtained. Finally, the object-level features are generated by transforming the convolutional features through mapping. The coarse branch is developed to transform the obtained object-level features into global structure for representing images. The positive experimental results counted on RS benchmark data set demonstrate that the proposed M-Net can learn more powerful features. Chao Liu 0042, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0001, Junyong Ma, Licheng Jiao |
IGARSS | 5 |
| 2020 | Weakly Supervised Land Cover Classification Method For Large-Scale Multi-Resolution Labeled Satellite Images Data SetsabstractThe global land cover map obtained by satellite is of vital importance to record land information. Despite the large amount of data acquired, rare annotations are available. Weakly-supervised learning methods can help to get use of extensive existing data resources and reduce the cost of manual labeling, which is of great significance for large-scale land cover mapping. Considering that the low-resolution labels are inaccurate and inexact and that high-resolution labels are rare, which is likely to cause training overfitting, we propose a method that can effectively apply multi-resolution labels, combining machine learning and deep learning. First, high-resolution labels of the validation set are used to train DeepLabv3+ (DLv3) and Random Forest (RF), which are applied to predict the training set then. The intersection of the prediction results is then used to supervise the training of DLv3 as pseudo-labels together with the high-resolution labels. We also added strategies such as model fusion, data augmentation and weighted training to improve classification accuracy. In the end, our method achieves the average accuracy (AA) of 0.609, ranking 3rd in Track 2 of the IEEE Data Fusion Contest (DFC) of 2020. Shuting Yin, Dafan Chen, Chengconghui Ma, Yanchao Lian, Licheng Jiao, Fang Liu 0001 |
IGARSS | 6 |
| 2020 | AttAN: Attention Adversarial Networks for 3D Point Cloud Semantic Segmentationabstract3D point cloud semantic segmentation has attracted wide attention with its extensive applications in autonomous driving, AR/VR, and robot sensing fields. However, in existing methods, each point in the segmentation results is predicted independently from each other. This property causes the non-contiguity of label sets in three-dimensional space and produces many noisy label points, which hinders the improvement of segmentation accuracy. To address this problem, we first extend adversarial learning to this task and propose a novel framework Attention Adversarial Networks (AttAN). With high-order correlations in label sets learned from the adversarial learning, segmentation network can predict labels closer to the real ones and correct noisy results. Moreover, we design an additive attention block for the segmentation network, which is used to automatically focus on regions critical to the segmentation task by learning the correlation between multi-scale features. Adversarial learning, which explores the underlying relationship between labels in high-dimensional space, opens up a new way in 3D point cloud semantic segmentation. Experimental results on ScanNet and S3DIS datasets show that this framework effectively improves the segmentation quality and outperforms other state-of-the-art methods. Qinghua Ma, Licheng Jiao, Fang Liu 0001, Qigong Sun |
IJCAI | 4 |
| 2020 | Complex Contourlet-CNN for polarimetric SAR image classification
Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Qigong Sun, Jin Zhao 0002 |
Pattern Recognit. | 4 |
| 2020 | Multi-Feature Weighted Sparse Graph for SAR Image AnalysisabstractSparse representation (SR) method has the advantages of good category distinguishing performance, noise robustness, and data adaptiveness. In this article, a multi-feature weighted sparse graph (MWSG) is presented for synthetic aperture radar (SAR) image analysis. First, multiple types of features are extracted to fully describe the characteristics of SAR image. Then, multiple SRs of samples in multiple feature spaces are obtained by solving a weighted joint SR model, in which the weight is the Gaussian kernel distance among samples. Moreover, a new fusion mechanism is given to integrate multiple weighted SRs, which aims to eliminate the negative influence of the singular data, so the MWSG is obtained. Afterward, the brief steps of the SAR image segmentation and semisupervised classification based on MWSG are stated. A series of experiments on the simulated and real SAR images shows that the MWSG has better performance than other existing relevant methods. Licheng Jiao, Fang Liu 0001, Xiangrong Zhang, Xu Tang 0004, Puhua Chen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Online Active Extreme Learning Machine With Discrepancy Sampling for PolSAR ClassificationabstractThe extreme learning machine (ELM) has drawn increasing attention in the field of machine learning due to its high accuracy and efficient learning. However, classical ELM works in batch and passive learning paradigms, which cannot deal with sequential data effectively. ELM has been extended to online sequential learning form (OS-ELM) and active learning form (AL-ELM), in which the former is used to improve training efficiency and the latter is mainly adopted to improve accuracy. In order to solve the problem of labeling samples difficulty and costly, poor sample validity, and continuous iterative learning in polarimetric synthetic aperture radar (PolSAR) image classification, we propose an online active extreme learning machine (OA-ELM) algorithm to combine the strengths and make up the weaknesses of OS-ELM and AL-ELM, which improves both efficiency and generalization ability. OA-ELM can learn from sequential data dynamically with low computational complexity and good generalization ability. Specifically, OA-ELM reduces time and memory cost for training via extended recursive least squares for optimization. It also improves accuracy using informative training samples selected by proposed discrepancy sampling (DS), which modifies an active query method called margin sampling (MS). Before applying MS to ELM, real-valued outputs of ELM need to be converted into probabilistic outputs first. Instead, the proposed DS can be applied to ELM directly by calculating the difference between the two largest actual nonprobabilistic outputs of ELM. Experimental results of PolSAR classification demonstrate that OA-ELM is effective and efficient compared with other algorithms in terms of accuracy and running time. Lingling Li 0002, Licheng Jiao, Pujiang Liang, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | A Stepwise Method for Change Detection in Large-Scale Polarimetric SAR ImagesabstractIn this paper, a stepwise method which consists of two main steps is proposed to tackle change detection in large-scale polarimetric Synthetic Aperture Radar (SAR) images. First, down-sample two registered polarimetric SAR images and calculate the corresponding Difference Image (DI), then spatial localization is conducted to orient sub-region which contains changed area in high probability. Second, polarimetric SAR sub-images are selected and they are trained together with down-sampled whole images by Convolutional Neural Network (CNN), where changed areas in sub-region are revealed and the next sub-region is generated accordingly. Repeat these two steps until all interested regions are detected. In general, it collects the whole but coarse information at first to locate important domains with changed areas and then analyzes them for accurate detection result and generate the next sub-region for further detection. Experiment results show that the proposed method performs well in detecting changed areas in large-scale polarimetric SAR images. Fang Liu 0001, Xu Tang 0004 |
IGARSS | 1 |
| 2019 | Polsar Image Classification Based on Polarimetric Scattering Coding and Sparse Support Matrix MachineabstractPOLSAR image has an advantage over optical image because it can be acquired independently of cloud cover and solar illumination. PolSAR image classification is a hot and valuable topic for the interpretation of POLSAR image. In this paper, a novel POLSAR image classification method is proposed based on polarimetric scattering coding and sparse support matrix machine. First, we transform the original POLSAR data to get a real value matrix by the polarimetric scattering coding, which is called polarimetric scattering matrix and is a sparse matrix. Second, the sparse support matrix machine is used to classify the sparse polarimetric scattering matrix and get the classification map. The combination of these two steps takes full account of the characteristics of POLSAR. The experimental results show that the proposed method can get better results and is an effective classification method. Xu Liu 0006, Licheng Jiao, Fang Liu 0001 |
IGARSS | 4 |
| 2019 | Pixel Dag-Recurrent Neural Network for Spectral-Spatial Hyperspectral Image ClassificationabstractExploiting rich spatial and spectral features contributes to improve the classification accuracy of hyperspectral images (HSIs). In this paper, based on the mechanism of the population receptive field (pRF) in human visual cortex, we further utilize the spatial correlation of pixels in images and propose pixel directed acyclic graph recurrent neural network (Pixel DAG-RNN) to extract and apply spectral-spatial features for HSIs classification. In our model, an undirected cyclic graph (UCG) is used to represent the relevance connectivity of pixels in an image patch, and four DAGs are used to approximate the spatial relationship of UCGs. In order to avoid overfitting, weight sharing and dropout are adopted. The higher classification performance of our model on HSIs classification has been verified by experiments on three benchmark data sets. Xiufang Li, Qigong Sun, Lingling Li 0002, Zhongle Ren, Fang Liu 0001, Licheng Jiao |
IGARSS | 5 |
| 2019 | Semi-Supervised Complex-Valued GAN for Polarimetric SAR Image ClassificationabstractPolarimetric synthetic aperture radar (PolSAR) images are widely used in disaster detection and military reconnaissance and so on. However, their interpretation faces some challenges, e.g., deficiency of labeled data, inadequate utilization of data information and so on. In this paper, a complex-valued generative adversarial network (GAN) is proposed for the first time to address these issues. The complex number form of model complies with the physical mechanism of PolSAR data and in favor of utilizing and retaining amplitude and phase information of PolSAR data. GAN architecture and semi-supervised learning are combined to handle deficiency of la-beled data. GAN expands training data and semi-supervised learning is used to train network with generated, labeled and unlabeled data. Experimental results on two benchmark data sets show that our model outperforms existing state-of-the-art models, especially for conditions with fewer labeled data. Qigong Sun, Xiufang Li, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Licheng Jiao |
IGARSS | 5 |
| 2019 | Fast unsupervised deep fusion network for change detection of multitemporal SAR images
Huan Chen 0006, Licheng Jiao, Miaomiao Liang, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
Neurocomputing | 4 |
| 2019 | Video reconstruction based on Intrinsic Tensor Sparsity model
Fang Liu 0001, Licheng Jiao, Shuyuan Yang 0001 |
Signal Process. Image Commun. | 2 |
| 2019 | A Pareto-Based Sparse Subspace Learning FrameworkabstractHigh-dimensionality is a common characteristic of real-world data, which often results in high time and space complexity or poor performance of ensuing methods. Subspace learning, as one kind of dimension reduction method, provides a way to overcome the aforementioned problem. In this paper, we introduce multiobjective evolutionary optimization into subspace learning, and propose a Pareto-based sparse subspace learning algorithm for classification tasks. The proposed algorithm aims at minimizing two conflicting objective functions, the reconstruction error and the sparsity. A kernel trick derived from Gaussian kernel is implemented to the sparse subspace learning for the nonlinear phenomena of nature. In order to speed up the convergence, an entropy-driven initialization scheme and a gradient-descent mutation scheme are designed specifically. At last, a knee point is selected from the Pareto front to guarantee that we can obtain a solution with good classification performance, and yet as sparse as possible. The experiments and detailed analysis on real-life datasets and the hyperspectral images demonstrated that the proposed model achieves comparable results with the existing conventional subspace learning and evolutionary feature selection algorithms. Hence, this paper provides a more flexible and efficient approach for sparse subspace learning. Juanjuan Luo, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Wenping Ma 0001 |
IEEE Trans. Cybern. | 3 |
| 2019 | Transferred Deep Learning-Based Change Detection in Remote Sensing ImagesabstractSupervised deep neural networks (DNNs) have been extensively used in diverse tasks. Generally, training such DNNs with superior performance requires a large amount of labeled data. However, it is time-consuming and expensive to manually label the data, especially for tasks in remote sensing, e.g., change detection. The situation motivates us to resort to the existing related images with labels, from which the concept of change can be adapted to new images. However, the distributions of the related labeled images (source domain) and unlabeled new images (target domain) are similar but not identical. It impedes a change detection model learned from source domains being well applied to the target domain. In this paper, we propose a transferred deep learning-based change detection framework to solve this problem. It consists of pretraining and fine-tuning stages. In the pretraining process, we propose two tasks to be learned simultaneously, namely, change detection for the source domain with labels and reconstruction of the unlabeled target data. The auxiliary task aims to reconstruct the difference image (DI) for the target domain. DI is an effective feature, such that the auxiliary task is of much relevance to change detection. The lower layers are shared between these two tasks in the training process. It mitigates the distribution discrepancy between the source and target domains and makes the concept of change from the source domain adapt to the target domain. In addition, we evaluate three modes of the U-net architecture to merge the information for a pair of patches. To fine-tune the change detection network (CDN) for the target domain, two strategies are exploited to select the pixels that have a high possibility of being correctly classified by an unsupervised approach. The proposed method demonstrates an excellent capacity for adapting the concept of change from the source domain to the target domain. It outperforms the state-of-the-art change detection methods via experimental results on real remote sensing data sets. Meijuan Yang, Licheng Jiao, Fang Liu 0001, Biao Hou, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | A Novel Neural Network for Remote Sensing Image MatchingabstractRapid development of remote sensing (RS) imaging technology makes the acquired images have larger size, higher resolution, and more complex structure, which goes beyond the reach of classical hand-crafted feature-based matching. In this paper, we propose a feature learning approach based on two-branch networks to transform the image matching task into a two-class classification problem. To match two key points, two image patches centered at the key points are entered into the proposed network. The network aims to learn discriminative feature representations for patch matching, so that more matching pairs can be obtained on the premise of maintaining higher subpixel matching accuracy. The proposed network adopts a two-stage training mode to deal with the complex characteristics of RS images. An adaptive sample selection strategy is proposed to determine the size of each patch by the scale of its central key point. Thus, each patch can preserve the texture structure around its key point rather than all patches have a predetermined size. In the matching prediction stage, two strategies, namely, superpixel-based sample graded strategy and superpixel-based ordered spatial matching, are designed to improve the matching efficiency and matching accuracy, respectively. The experimental results and theoretical analysis demonstrate the feasibility, robustness, and effectiveness of the proposed method. Hao Zhu 0009, Licheng Jiao, Wenping Ma 0001, Fang Liu 0001, Wei Zhao 0014 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Circular Relevance Feedback for Remote Sensing Image RetrievalabstractRelevance feedback (RF) is a popular reranking technique, which aims at improving the performance of image retrieval by taking the user's opinions into account. In this paper, we introduce a new RF method, named circular relevance feedback (CRF), to enhance the behavior of remote sensing image retrieval (RSIR). Instead of the manual selection used in the common RF method, we adopt the active learning (AL) algorithm to select the samples from the initial results automatically in each RF iteration. Moreover, to ensure the selected images are representative and informative enough, we choose different AL algorithms to complete the different RF processes. Finally, the contributions of all AL-driven RF methods are integrated using a circular fusion scheme. The encouraging experimental results on the ground truth RS image archive illustrate that our CRF is useful for enhancing the performance of RSIR. In addition, compared with many existing RF methods, our CRF achieves improved behavior. Xu Tang 0004, Xiangrong Zhang, Fang Liu 0001, Licheng Jiao |
IGARSS | 3 |
| 2018 | Random subspace based ensemble sparse representation
Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Rongfang Wang, Puhua Chen, Yuanhao Cui, Junhu Xie, Yake Zhang |
Pattern Recognit. | 3 |
| 2018 | A modified convolutional neural network for face sketch synthesis
Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Wenping Ma 0001 |
Pattern Recognit. | 4 |
| 2018 | Fuzzy Sparse Autoencoder Framework for Single Image Per Person Face RecognitionabstractThe issue of single sample per person (SSPP) face recognition has attracted more and more attention in recent years. Patch/local-based algorithm is one of the most popular categories to address the issue, as patch/local features are robust to face image variations. However, the global discriminative information is ignored in patch/local-based algorithm, which is crucial to recognize the nondiscriminative region of face images. To make the best of the advantage of both local information and global information, a novel two-layer local-to-global feature learning framework is proposed to address SSPP face recognition. In the first layer, the objective-oriented local features are learned by a patch-based fuzzy rough set feature selection strategy. The obtained local features are not only robust to the image variations, but also usable to preserve the discrimination ability of original patches. Global structural information is extracted from local features by a sparse autoencoder in the second layer, which reduces the negative effect of nondiscriminative regions. Besides, the proposed framework is a shallow network, which avoids the over-fitting caused by using multilayer network to address SSPP problem. The experimental results have shown that the proposed local-to-global feature learning framework can achieve superior performance than other state-of-the-art feature learning algorithms for SSPP face recognition. Yuwei Guo 0001, Licheng Jiao, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001 |
IEEE Trans. Cybern. | 5 |
| 2018 | The Overcomplete Dictionary-Based Directional Estimation Model and Nonconvex Reconstruction MethodsabstractIn this paper, it is proposed the directional estimation model on the overcomplete dictionary, which bridges the compressed measurements of the image blocks and the directional structures of the dictionary. In the model, it is established the analytical method to estimate the structure type of a block as either smooth, single-oriented, or multioriented. Furthermore, the structures of each type of blocks are described by the structured subdictionaries. Then based on the obtained estimations and the constrains on the sparse dictionaries, the original image will be estimated. To verify the model, the nonconvex methods are designed for compressed sensing. Specifically, the greedy pursuit-based methods are established to search the subdictionaries obtained by the model, which achieve better local structural estimation than the methods without the directional estimation. More importantly, it is proposed the nonconvex image reconstruction method with direction-guided dictionaries and evolutionary searching strategies (NR_DG), where the evolutionary searching strategies are delicately designed for each type of the blocks based on the directional estimation. By the experimental results, it is shown that the NR_DG method performs better than the available two-stage evolutionary reconstruction method. Leping Lin, Fang Liu 0001, Licheng Jiao, Shuyuan Yang 0001, Hongxia Hao |
IEEE Trans. Cybern. | 2 |
| 2018 | Global Low-Rank Image Restoration With Gaussian Mixture ModelabstractLow-rank restoration has recently attracted a lot of attention in the research of computer vision. Empirical studies show that exploring the low-rank property of the patch groups can lead to superior restoration performance, however, there is limited achievement on the global low-rank restoration because the rank minimization at image level is too strong for the natural images which seldom match the low-rank condition. In this paper, we describe a flexible global low-rank restoration model which introduces the local statistical properties into the rank minimization. The proposed model can effectively recover the latent global low-rank structure via nuclear norm, as well as the fine details via Gaussian mixture model. An alternating scheme is developed to estimate the Gaussian parameters and the restored image, and it shows excellent convergence and stability. Besides, experiments on image and video sequence datasets show the effectiveness of the proposed method in image inpainting problems. Licheng Jiao, Fang Liu 0001, Shuang Wang 0001 |
IEEE Trans. Cybern. | 3 |
| 2018 | Fuzzy Double C-Means Clustering Based on Sparse Self-RepresentationabstractThis paper introduces the popular sparse representation method into the classical fuzzy c-means clustering algorithm, and presents a novel fuzzy clustering algorithm, called fuzzy double c-means based on sparse self-representation (FDCM_SSR). The major characteristic of FDCM_SSR is that it can simultaneously address two datasets with different dimensions, and has two kinds of corresponding cluster centers. The first one is the basic feature set that represents the basic physical property of each sample itself. The second one is learned from the basic feature set by solving a spare self-representation model, referred to as discriminant feature set, which reflects the global structure of the sample set. The spare self-representation model employs dataset itself as dictionary of sparse representation. It has good category distinguishing ability, noise robustness, and data-adaptiveness, which enhance the clustering and generalization performance of FDCM_SSR. Experiments on different datasets and images show that FDCM_SSR is more competitive than other state-of-the-art fuzzy clustering algorithms. Licheng Jiao, Shuyuan Yang 0001, Fang Liu 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2018 | Fuzzy Superpixels for Polarimetric SAR Images ClassificationabstractSuperpixels technique has drawn much attention in computer vision applications. Each superpixels algorithm has its own advantages. Selecting a more appropriate superpixels algorithm for a specific application can improve the performance of the application. In the last few years, superpixels are widely used in polarimetric synthetic aperture radar (PolSAR) image classification. However, no superpixel algorithm is especially designed for image classification. It is believed that both mixed superpixels and pure superpixels exist in an image. Nevertheless, mixed superpixels have negative effects on classification accuracy. Thus, it is necessary to generate superpixels containing as few mixed superpixels as possible for image classification. In this paper, first, a novel superpixels concept, named fuzzy superpixels, is proposed for reducing the generation of mixed superpixels. In fuzzy superpixels, not all pixels are assigned to a corresponding superpixel. We would rather ignore the pixels than assigning them to improper superpixels. Second, a new algorithm, named FuzzyS (FS), is proposed to generate fuzzy superpixels for PolSAR image classification. Three PolSAR images are used to verify the effect of the proposed FS algorithm. Experimental results demonstrate the superiority of the proposed FS algorithm over several state-of-the-art superpixels algorithms. Yuwei Guo 0001, Licheng Jiao, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001, Wenqiang Hua |
IEEE Trans. Fuzzy Syst. | 5 |
| 2018 | Adaptive Hierarchical Multinomial Latent Model With Hybrid Kernel Function for SAR Image Semantic SegmentationabstractSynthetic aperture radar (SAR) images have been one of the important tools to support earth observations and topographic measurements. It means that SAR images are essentially rich in structures. However, the single spatial relationship is difficult to deal with the heterogeneous structures of the SAR images. In this paper, we propose an adaptive hierarchical multinomial latent model with hybrid kernel function for SAR image semantic segmentation. In the proposed approach, we design a hybrid kernel function combing Gaussian radial basis function (GRBF) and ridgelet kernel function to adaptively describe the spatial relationships between the central pixel and the surrounding pixels. Then, based on the hybrid kernel function, adaptive methods are proposed for semantic segmentation. Specifically, an SAR image is divided into different characteristics subspaces, homogeneous, structural, and aggregated subspaces, by SAR hierarchical semantic model. For the homogeneous subspace, GRBF is used to describe the isotropic spatial relationships. Then, multilayer multinomial latent model with GRBF is used for segmentation to improve the labeling consistency and reduce the wrong segmentation. For the structural subspace, the ridgelet kernel function is used to describe the anisotropic spatial relationships. Then, we adopt the single-layer multinomial latent model with ridgelet kernel function for segmentation to preserve the details (such as edge, lines, and small objects). For aggregated subspace, bag-of-words model is used to extract the features of the aggregated portions, and then affinity propagation cluster is used for segmentation. Finally, the segmentation results of different subspaces are integrated together to obtain the final segmentation result. Comprehensive experiments on both synthetic and real SAR images demonstrate that the segmentation results by our proposed approach achieve the semantic consistency, labeling consistency, and detail preservation simultaneously. Yiping Duan, Fang Liu 0001, Licheng Jiao, Xiaoming Tao 0001, Jie Wu 0016, Cheng Shi 0002, Martin O. Wimmers |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Deep Multiple Instance Learning-Based Spatial-Spectral Classification for PAN and MS ImageryabstractPanchromatic (PAN) and multispectral (MS) imagery classification is one of the hottest topics in the field of remote sensing. In recent years, deep learning techniques have been widely applied in many areas of image processing. In this paper, an end-to-end learning framework based on deep multiple instance learning (DMIL) is proposed for MS and PAN images’ classification using the joint spectral and spatial information based on feature fusion. There are two instances in the proposed framework: one instance is used to capture the spatial information of PAN and the other is used to describe the spectral information of MS. The features obtained by the two instances are concatenated directly, which can be treated as simple fusion features. To fully fuse the spatial–spectral information for further classification, the simple fusion features are fed into a fusion network with three fully connected layers to learn the high-level fusion features. Classification experiments carried out on four different airborne MS and PAN images indicate that the classifier provides feasible and efficient solution. It demonstrates that DMIL performs better than using a convolutional neural network and a stacked autoencoder network separately. In addition, this paper shows that the DMIL model can learn and fuse spectral and spatial information effectively, and has huge potential for MS and PAN imagery classification. Xu Liu 0006, Licheng Jiao, Jiaqi Zhao 0001, Jin Zhao 0002, Fang Liu 0001, Shuyuan Yang 0001, Xu Tang 0004 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2018 | A Hybrid Method of SAR Speckle Reduction Based on Geometric-Structural Block and Adaptive NeighborhoodabstractGiven the improvement of synthetic aperture radar (SAR) imaging technologies, the resolution of SAR image is largely improved and the variation of backscatter amplitude should be considered in SAR image processing. In this paper, considering the spatial geometric properties of SAR image in gray pixel space and the sample selection in the estimation of true signal, local directional property of each pixel is explored with the help of SAR sketching method, and two specially designed filters are integrated for adaptive speckle reduction of SAR images. Specifically, based on the sketch map of a SAR image, the orientation of the sketch point lying at each sketch segment is assigned to the corresponding pixel, and thus all pixels of the SAR image are classified as the directional pixels and the nondirectional pixels. For the directional pixels, given the significant directionality of its neighborhood, a geometric-structural block (GB) is built to center on it and GB-wised nonlocal means filter is designed to estimate the true values of all pixels contained in the GB. Moreover, using the local orientation, the whole image is adopted as the searching range to search the similar GBs. For the nondirectional pixels, based on the locally estimated equivalent number of looks, a novel pixel-based metric is proposed to determine the local adaptive neighborhood (AN) with which an AN-based filter is developed to estimate its true value. Besides, since some nondirectional pixels are contained in GBs, a Bayesian-based fusion strategy is designed for the fusion of their estimated values. In the experiments, three synthetic speckled images and five real SAR images [obtained with different resolutions (e.g., 3, 1, and 0.1 m) and different bands (e.g., X-band, C-band, and Ka-band)] are used for evaluation and analysis. Owing to the usage of local spatial geometric property and the combination of two different filters, the proposed method shows a reasonable performance among the comparison methods, in terms of the speckle reduction and the details' preservation. Fang Liu 0001, Jie Wu 0016, Lingling Li 0002, Licheng Jiao, Hongxia Hao, Xiangrong Zhang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | Semi-supervised double sparse graphs based discriminant analysis for dimensionality reduction
Puhua Chen, Licheng Jiao, Fang Liu 0001, Jiaqi Zhao 0001, Shuai Liu 0016 |
Pattern Recognit. | 3 |
| 2017 | SAR Image segmentation based on convolutional-wavelet neural network and markov random field
Yiping Duan, Fang Liu 0001, Licheng Jiao, Lu Zhang 0028 |
Pattern Recognit. | 2 |
| 2017 | A Novel Image Representation Framework Based on Gaussian Model and Evolutionary OptimizationabstractWe propose a novel image representation framework based on Gaussian model and evolutionary optimization (EO). In this framework, image patches are categorized into smooth and nonsmooth ones, and the two categories are treated distinctively. For a smooth patch, we formulate it as the summation of a direct component and a variation component (VC). We observe that the values of all VCs in an image can be well fitted by a Gaussian distribution, according to which we present an efficient reconstruction approach based on maximizing the logarithm a posteriori probability. For a nonsmooth patch, we introduce the mechanism of EO to solve a combinatorial optimization over a principal component analysis dictionary. In addition, we develop two approaches for estimating the coefficients of the atoms. Experiment results demonstrate that the proposed framework obtains the state-of-the-art results in several image inverse problems. Licheng Jiao, Lingling Li 0002, Shuyuan Yang 0001, Fang Liu 0001, Hongxia Hao |
IEEE Trans. Evol. Comput. | 5 |
| 2017 | Two-Stage Reranking for Remote Sensing Image RetrievalabstractImage reranking is a popular postprocessing method for remote sensing image retrieval (RSIR), which aims at enhancing the initial retrieval performance. In general, it takes either users' opinions or the relationships between images into consideration to find an optimal reranked list based on the initial retrieved results. In this paper, we present a reranking method for improving RSIR, which is named two-stage reranking (TSR). Suppose the k-nearest neighbors of a query RS image have been obtained by the initial retrieval. The first step of our TSR is to edit these neighbors using the editing scheme. A handful of informative and representative RS images are selected by the active learning algorithm, and their binary labels are provided by the users relative to the query image. Then, a binary classifier is trained using the selected RS images and their labels to classify the rest of the neighbors. Finally, both classification results and rank information in the initial retrieval results are considered to decide which neighbor should be excluded. In the next step, the remaining RS images are reranked by the proposed reranking scheme, i.e., multisimilarity fusion reranking. Both the user's experience and image relationships are taken into account in TSR to ensure the performance of the reranking. The efficiency and the robustness of our method are validated by experiments conducted on two different types of RS images. Compared with the existing visual reranking approaches, our method achieves improved performance. Xu Tang 0004, Licheng Jiao, William J. Emery, Fang Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2016 | Prediction of missing links based on community relevance and ruler inference
Jingyi Ding, Licheng Jiao, Jianshe Wu, Fang Liu 0001 |
Knowl. Based Syst. | 4 |
| 2016 | Sketching Model and Higher Order Neighborhood Markov Random Field-Based SAR Image SegmentationabstractThe Markov random field (MRF) model has been successfully applied to synthetic aperture radar (SAR) image segmentation because of its excellent ability of capturing the local contextual information in the prior model. However, the geometric structures of the SAR image are always ignored when capturing the contextual information in the prior model. Therefore, this letter presents a new SAR image segmentation method based on the sketching model and higher order neighborhood MRF. In this approach, the sketching model is utilized to represent the geometric structures of the SAR image. Meanwhile, a higher order neighborhood is constructed to capture the complex priors. Then, according to the structure fluctuation in the higher order neighborhood, the homogeneous and heterogeneous neighborhoods are distinguished. Finally, the local energy function in the prior model is constructed in the higher order neighborhood with different characteristics. Specifically, the energy functions considering the labeling consistency and focusing on the structure preservations are designed for the homogeneous and heterogeneous neighborhoods, respectively. In this way, the ability of the prior model is improved by adding the geometric structures into the energy functions. Experiments on the real SAR images demonstrate the effectiveness of the proposed method in labeling consistency and structure preservations. Yiping Duan, Fang Liu 0001, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | Detecting Cars in VHR SAR Images via Semantic CFAR AlgorithmabstractIn this letter, a novel semantic constant-false-alarm-rate (CFAR) method for the detection of cars from a very high resolution synthetic-aperture-radar (SAR) image is presented. The method not only employs the strong scattering features of the target which is used in CFAR but also employs the shadow features of the target. Furthermore, the semantic relationship between the strong scattering features and the shadow features is established to partly reduce the false alarm targets. By the experiments on the MiniSAR image of 4-in resolution, it is shown that the proposed semantic CFAR method outperforms the CFAR algorithm by a much lower false alarm rate. Fang Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | A Nonlocal Means for Speckle Reduction of SAR Image With Multiscale-Fusion-Based Steerable Kernel FunctionabstractFor the robustness of a patch-based metric, the nonlocal means method is widely applied for speckle reduction of synthetic aperture radar (SAR) images, where the similarity computed by the patch-based metric is used as weight, and weighted averaging is used to obtain the true value. However, not knowing the local spatial property, a fixed kernel (e.g., Gaussian kernel or uniform kernel) is always used to compute the weight. This is not good for the preservation of geometrical features (e.g., edges, lines, and points). In this letter, considering the characteristics of SAR imagery, a multiscale-fusion-based steerable kernel function was formed to explore the local spatial property of SAR images. In addition, by combining the kernel function with a ratio-based similarity metric designed with the distribution of the speckle's ratio, a new patch-based metric was formed and used with the nonlocal scheme for speckle reduction. In the experiments, by comparing with two state-of-the-art methods, a reasonable performance was obtained by our method, in terms of speckle reduction and detail preservation. Jie Wu 0016, Fang Liu 0001, Hongxia Hao, Lingling Li 0002, Licheng Jiao, Xiangrong Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | Unsupervised feature selection based on maximum information and minimum redundancy for hyperspectral images
Jie Feng 0003, Licheng Jiao, Fang Liu 0001, Tao Sun 0007, Xiangrong Zhang |
Pattern Recognit. | 3 |
| 2016 | Hierarchical semantic model and scattering mechanism based PolSAR image classification
Fang Liu 0001, Junfei Shi, Licheng Jiao, Hongying Liu 0001, Shuyuan Yang 0001, Jie Wu 0016, Hongxia Hao, Jialing Yuan |
Pattern Recognit. | 1 |
| 2016 | Learning simultaneous adaptive clustering and classification via MOEA
Juanjuan Luo, Licheng Jiao, Ronghua Shang, Fang Liu 0001 |
Pattern Recognit. | 4 |
| 2016 | MOEA/D with biased weight adjustment inspired by user preference and its application on multi-objective reservoir flood control problem
Xiaoliang Ma 0001, Fang Liu 0001, Yutao Qi, Lingling Li 0002, Licheng Jiao, Xiaozheng Deng, Xiaodong Wang 0011, Bei Dong, Zhanting Hou, Yongxiao Zhang, Jianshe Wu |
Soft Comput. | 2 |
| 2016 | A group matching pursuit for image reconstruction
Fang Liu 0001, Licheng Jiao, Hongxia Hao, Shuyuan Yang 0001 |
Signal Process. Image Commun. | 2 |
| 2016 | Geometric structure guided collaborative compressed sensing
Leping Lin, Fang Liu 0001, Licheng Jiao |
Signal Process. Image Commun. | 2 |
| 2016 | A Multiobjective Evolutionary Algorithm Based on Decision Variable Analyses for Multiobjective Optimization Problems With Large-Scale VariablesabstractState-of-the-art multiobjective evolutionary algorithms (MOEAs) treat all the decision variables as a whole to optimize performance. Inspired by the cooperative coevolution and linkage learning methods in the field of single objective optimization, it is interesting to decompose a difficult high-dimensional problem into a set of simpler and low-dimensional subproblems that are easier to solve. However, with no prior knowledge about the objective function, it is not clear how to decompose the objective function. Moreover, it is difficult to use such a decomposition method to solve multiobjective optimization problems (MOPs) because their objective functions are commonly conflicting with one another. That is to say, changing decision variables will generate incomparable solutions. This paper introduces interdependence variable analysis and control variable analysis to deal with the above two difficulties. Thereby, an MOEA based on decision variable analyses (DVAs) is proposed in this paper. Control variable analysis is used to recognize the conflicts among objective functions. More specifically, which variables affect the diversity of generated solutions and which variables play an important role in the convergence of population. Based on learned variable linkages, interdependence variable analysis decomposes decision variables into a set of low-dimensional subcomponents. The empirical studies show that DVA can improve the solution quality on most difficult MOPs. The code and supplementary material of the proposed algorithm are available athttp://web.xidian.edu.cn/fliu/paper.html. Xiaoliang Ma 0001, Fang Liu 0001, Yutao Qi, Xiaodong Wang 0011, Lingling Li 0002, Licheng Jiao, Minglei Yin, Maoguo Gong |
IEEE Trans. Evol. Comput. | 2 |
| 2016 | SAR Image Segmentation Based on Hierarchical Visual Semantic and Adaptive Neighborhood Multinomial Latent ModelabstractA synthetic aperture radar (SAR) imaging system usually produces pairs of bright area and dark area when depicting the ground objects, such as a building or tree and its shadow. Many buildings (trees) are aggregated together to form urban areas (forests). It means that the pairs of bright and dark areas often exist in the aggregated scenes. Conventional unsupervised segmentation approaches usually segment the scenes (e.g., urban areas and forests) into different regions simply according to the gray values of the image. However, a more convincing way is to regard them as the consistent regions. In this paper, we aim at addressing this issue and propose a new SAR image segmentation approach via a hierarchical visual semantic and adaptive neighborhood multinomial latent model. In this approach, the hierarchical visual semantic of SAR images is proposed, which divides SAR images into aggregated, structural, and homogeneous regions. Based on the division, different segmentation methods are chosen for these regions with different characteristics. For the aggregated region, locality-constrained linear coding-based hierarchical clustering is used for segmentation. For the structural region, visual semantic rules are designed for line object location, and a geometric structure window-based multinomial latent model is proposed for segmentation. For the homogeneous region, a multinomial latent model with adaptive window selection is proposed for segmentation. Finally, these results are integrated together to obtain the final segmentation. Experiments on both synthetic and real SAR images indicate that the proposed method achieves promising performances in terms of the consistencies of the regions and the preservations of the edges and line objects. Fang Liu 0001, Yiping Duan, Lingling Li 0002, Licheng Jiao, Jie Wu 0016, Shuyuan Yang 0001, Xiangrong Zhang, Jialing Yuan |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | CRIM-FCHO: SAR Image Two-Stage Segmentation With Multifeature EnsembleabstractThis paper investigates the synthetic aperture radar (SAR) image segmentation in terms of feature analysis and fusion and develops a new algorithm based on multifeature ensemble accordingly. This paper is characterized by two aspects. First, multiple heterogeneous features are extracted to accurately describe the objects in SAR images. These features are then integrated in the feature level and the similarity level, respectively, to avoid the mutual influences between different kinds of features and maximize the discriminability of the similarity measure between objects. Second, a two-stage algorithm consisting of a coarse merging stage and a fine classification stage is proposed. In the coarse merging stage, a context-based region iterative merging algorithm is designed to merge most of the unambiguous superpixels in image domain at a high speed. In the fine classification stage, a fuzzy clustering algorithm incorporating hybrid optimization is developed to balance the efficiency and the robustness of the algorithm by simultaneously searching heuristically in the complete high-dimension feature space and searching along the direction of the gradient steepest descent in each feature subspace. The effectiveness of the proposed method has been successfully validated on synthetic and real SAR images. Licheng Jiao, Fang Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | Semisupervised Discriminant Feature Learning for SAR Image Category via Sparse EnsembleabstractTerrain scene classification plays an important role in various synthetic aperture radar (SAR) image understanding and interpretation. This paper presents a novel approach to characterize SAR image content by addressing category with a limited number of labeled samples. In the proposed approach, each SAR image patch is characterize by a discriminant feature which is generated in a semisupervised manner by utilizing a spare ensemble learning procedure. In particular, a nonnegative sparse coding procedure is applied on the given SAR image patch set to generate the feature descriptors first. The set is combined with a limited number of labeled SAR image patches and an abundant number of unlabeled ones. Then, a semisupervised sampling approach is proposed to construct a set of weak learners, in which each one is modeled by a logistic regression procedure. The discriminant information can be introduced by projecting SAR image patch on each weak learner. Finally, the features of SAR image patches are produced by a sparse ensemble procedure which can reduce the redundancy of multiple weak learners. Experimental results show that the proposed discriminant feature learning approach can achieve a higher classification accuracy than several state-of-the-art approaches. Licheng Jiao, Fang Liu 0001, Jiaqi Zhao 0001, Puhua Chen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2015 | A Memetic Optimization Strategy Based on Dimension Reduction in Decision SpaceabstractThere can be a complicated mapping relation between decision variables and objective functions in multi-objective optimization problems (MOPs). It is uncommon that decision variables influence objective functions equally. Decision variables act differently in different objective functions. Hence, often, the mapping relation is unbalanced, which causes some redundancy during the search in a decision space. In response to this scenario, we propose a novel memetic (multi-objective) optimization strategy based on dimension reduction in decision space (DRMOS). DRMOS firstly analyzes the mapping relation between decision variables and objective functions. Then, it reduces the dimension of the search space by dividing the decision space into several subspaces according to the obtained relation. Finally, it improves the population by the memetic local search strategies in these decision subspaces separately. Further, DRMOS has good portability to other multi-objective evolutionary algorithms (MOEAs); that is, it is easily compatible with existing MOEAs. In order to evaluate its performance, we embed DRMOS in several state of the art MOEAs to facilitate our experiments. The results show that DRMOS has the advantage in terms of convergence speed, diversity maintenance, and portability when solving MOPs with an unbalanced mapping relation between decision variables and objective functions. Handing Wang, Licheng Jiao, Ronghua Shang, Shan He 0001, Fang Liu 0001 |
Evol. Comput. | 5 |
| 2015 | Joint sparse regularization based Sparse Semi-Supervised Extreme Learning Machine (S3ELM) for classification
Xiao-Zhuo Luo, Fang Liu 0001, Shuyuan Yang 0001, Xiaodong Wang 0011 |
Knowl. Based Syst. | 2 |
| 2015 | Learning compressive sampling via multiscale and steerable support value transform
Shuyuan Yang 0001, Min Wang 0007, Shigang Wang 0001, Fang Liu 0001, Licheng Jiao |
Knowl. Based Syst. | 5 |
| 2015 | Coupled compressed sensing inspired sparse spatial-spectral LSSVM for hyperspectral image classification
Lixia Yang, Shuyuan Yang 0001, Sujing Li, Rui Zhang 0045, Fang Liu 0001, Licheng Jiao |
Knowl. Based Syst. | 5 |
| 2015 | A cooperative belief rule based decision support system for lymph node metastasis diagnosis in gastric cancer
Fang Liu 0001, Lingling Li 0002, Licheng Jiao, Zhi-Jie Zhou 0001, Jian-Bo Yang, Zhi-Long Wang |
Knowl. Based Syst. | 2 |
| 2015 | Imbalanced Hyperspectral Image Classification Based on Maximum MarginabstractHyperspectral remote sensing images own rich spectral information to distinguish different land-cover classes. Sometimes, it may encounter the case that some classes have much fewer pixels than other classes. In this case, traditional classification methods are not appropriate because they are prone to assign all the pixels to the classes with a large number of pixels. For such an imbalanced problem, ensemble learning is a good method by partitioning the majority classes into different groups with small sizes. However, the existing ensemble schemes are independent of classifiers, which will not get the best performance for a certain classifier. In this letter, the selected classifier, i.e., a support vector machine (SVM), is considered in an ensemble procedure to improve the classification accuracy. Specifically, the criterion of the SVM, i.e., the maximum margin, is adopted to guide the ensemble learning procedure for imbalanced hyperspectral image classification. Experiments state that our method obtains higher classification accuracy than the SVM and several representative imbalanced classification methods for hyperspectral images. Tao Sun 0007, Licheng Jiao, Jie Feng 0003, Fang Liu 0001, Xiangrong Zhang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2015 | Pan-sharpening via regional division and NSST
Cheng Shi 0002, Fang Liu 0001, Qiguang Miao |
Multim. Tools Appl. | 2 |
| 2015 | A novel dynamic rough subspace based selective ensemble
Yuwei Guo 0001, Licheng Jiao, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001, Kaixuan Rong |
Pattern Recognit. | 5 |
| 2015 | A new patch based change detector for polarimetric SAR data
Ganchao Liu, Licheng Jiao, Fang Liu 0001, Hua Zhong 0003, Shuang Wang 0001 |
Pattern Recognit. | 3 |
| 2015 | Nonconvex Compressed Sensing by Nature-Inspired Optimization AlgorithmsabstractThe l 0 regularized problem in compressed sensing reconstruction is nonconvex with NP-hard computational complexity. Methods available for such problems fall into one of two types: greedy pursuit methods and thresholding methods, which are characterized by suboptimal fast search strategies. Nature-inspired algorithms for combinatorial optimization are famous for their efficient global search strategies and superior performance for nonconvex and nonlinear problems. In this paper, we study and propose nonconvex compressed sensing for natural images by nature-inspired optimization algorithms. We get measurements by the block-based compressed sampling and introduce an overcomplete dictionary of Ridgelet for image blocks. An atom of this dictionary is identified by the parameters of direction, scale and shift. Of them, direction parameter is important for adapting to directional regularity. So we propose a two-stage reconstruction scheme (TS_RS) of nature-inspired optimization algorithms. In the first reconstruction stage, we design a genetic algorithm for a class of image blocks to acquire the estimation of atomic combinations in all directions; and in the second reconstruction stage, we adopt clonal selection algorithm to search better atomic combinations in the sub-dictionary resulted by the first stage for each image block further on scale and shift parameters. In TS_RS, to reduce the uncertainty and instability of the reconstruction problems, we adopt novel and flexible heuristic searching strategies, which include delicately designing the initialization, operators, evaluating methods, and so on. The experimental results show the efficiency and stability of the proposed TS_RS of nature-inspired algorithms, which outperforms classic greedy and thresholding methods. Fang Liu 0001, Leping Lin, Licheng Jiao, Lingling Li 0002, Shuyuan Yang 0001, Biao Hou, Hongmei Ma, Jinghuan Xu |
IEEE Trans. Cybern. | 1 |
| 2015 | Mutual-Information-Based Semi-Supervised Hyperspectral Band Selection With High Discrimination, High Information, and Low RedundancyabstractThe large number of spectral bands in hyperspectral images provides abundant information to distinguish different land covers. However, these spectral bands have much redundancy and bring an extra computational burden. Thus, band selection is important for hyperspectral images. Since the labeled samples are difficult to obtain, a semi-supervised criterion based on maximum discrimination and information (MDI) is defined by using both limited labeled samples and sufficient unlabeled samples. This MDI criterion aims to select the most highly discriminative and informative bands, but it is hard to accurately calculate. Therefore, a novel criterion based on high discrimination, high information, and low redundancy (DIR) is proposed as its low-order approximation. Moreover, from an information theory perspective, a theoretical proof is given that many traditional semi-supervised feature selection criteria are the low-order approximations of this MDI criterion. Compared with them, the proposed criterion needs more relaxed approximation conditions. To search and optimize the proposed criterion, a novel clonal selection algorithm is proposed, where the adaptive clone and mutation operators are devised to speed up the convergence. Experimental results on hyperspectral images demonstrate the effectiveness of the proposed semi-supervised band selection method. Jie Feng 0003, Licheng Jiao, Fang Liu 0001, Tao Sun 0007, Xiangrong Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2015 | Learning Interpolation via Regional Map for Pan-SharpeningabstractAlthough the bandwidth of the high-resolution panchromatic (HR PAN) image is wide, it is narrow in each band of the low-resolution multispectral (LR MS) image. Hence, the spatial resolution of the HR PAN image is much higher than that of the LR MS image. However, HR PAN image only has a single band. The purpose of the Pan-sharpening algorithm is to make the Pan-sharpened image with both high spatial resolution and good spectral information. In this paper, a novel learning interpolation method for Pan-sharpening is proposed by expanding the sketch information in the HR PAN image. The sketch information contains the edges and lines features of the image, and each segment of the sketch information has its own direction. According to the primal sketch graph of the HR PAN image, a regional map is obtained by a designed geometrical template. Since the size of the HR PAN image is different from that of the LR MS image, the LR MS image is interpolated into an interpolated multispectral (IMS) image by the nearest interpolation method. In addition, the IMS image can be mapped into the structure and the nonstructure regions by this regional map. The nonstructure regions are divided into the smooth and the texture regions by a variance value. For the structure and texture regions, the interpolated pixels in the IMS image are relearned and readjusted by the proposed structure and texture learning interpolation method, respectively. Experimental results show that the proposed Pan-sharpening method can provide superior performance in both visual effect and quality metrics, particularly for the images with a large spectral difference. Cheng Shi 0002, Fang Liu 0001, Lingling Li 0002, Licheng Jiao, Yiping Duan, Shuang Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2014 | MOEA/D with Adaptive Weight AdjustmentabstractRecently, MOEA/D (multi-objective evolutionary algorithm based on decomposition) has achieved great success in the field of evolutionary multi-objective optimization and has attracted a lot of attention. It decomposes a multi-objective optimization problem (MOP) into a set of scalar subproblems using uniformly distributed aggregation weight vectors and provides an excellent general algorithmic framework of evolutionary multi-objective optimization. Generally, the uniformity of weight vectors in MOEA/D can ensure the diversity of the Pareto optimal solutions, however, it cannot work as well when the target MOP has a complex Pareto front (PF; i.e., discontinuous PF or PF with sharp peak or low tail). To remedy this, we propose an improved MOEA/D with adaptive weight vector adjustment (MOEA/D-AWA). According to the analysis of the geometric relationship between the weight vectors and the optimal solutions under the Chebyshev decomposition scheme, a new weight vector initialization method and an adaptive weight vector adjustment strategy are introduced in MOEA/D-AWA. The weights are adjusted periodically so that the weights of subproblems can be redistributed adaptively to obtain better uniformity of solutions. Meanwhile, computing efforts devoted to subproblems with duplicate optimal solution can be saved. Moreover, an external elite population is introduced to help adding new subproblems into real sparse regions rather than pseudo sparse regions of the complex PF, that is, discontinuous regions of the PF. MOEA/D-AWA has been compared with four state of the art MOEAs, namely the original MOEA/D, Adaptive-MOEA/D, [Formula: see text]-MOEA/D, and NSGA-II on 10 widely used test problems, two newly constructed complex problems, and two many-objective problems. Experimental results indicate that MOEA/D-AWA outperforms the benchmark algorithms in terms of the IGD metric, particularly when the PF of the MOP is complex. Yutao Qi, Xiaoliang Ma 0001, Fang Liu 0001, Licheng Jiao, Jianyong Sun, Jianshe Wu |
Evol. Comput. | 3 |
| 2014 | MOEA/D with opposition-based learning for multiobjective optimization problem
Xiaoliang Ma 0001, Fang Liu 0001, Yutao Qi, Maoguo Gong, Minglei Yin, Lingling Li 0002, Licheng Jiao, Jianshe Wu |
Neurocomputing | 2 |
| 2014 | MOEA/D with Baldwinian learning inspired by the regularity property of continuous multiobjective problem
Xiaoliang Ma 0001, Fang Liu 0001, Yutao Qi, Lingling Li 0002, Licheng Jiao, Meiyun Liu, Jianshe Wu |
Neurocomputing | 2 |
| 2014 | Incomplete variables truncated conjugate gradient method for signal reconstruction in compressed sensing
Xiaodong Wang 0011, Fang Liu 0001, Licheng Jiao, Jiao Wu 0002, Jianrui Chen 0002 |
Inf. Sci. | 2 |
| 2014 | A compressed sensing approach for efficient ensemble learning
Lin Li 0016, Rustam Stolkin, Licheng Jiao, Fang Liu 0001, Shuang Wang 0001 |
Pattern Recognit. | 4 |
| 2014 | Compressed sensing by collaborative reconstruction on overcomplete dictionary
Leping Lin, Fang Liu 0001, Licheng Jiao |
Signal Process. | 2 |
| 2014 | MOEA/D with uniform decomposition measurement for many-objective problems
Xiaoliang Ma 0001, Yutao Qi, Lingling Li 0002, Fang Liu 0001, Licheng Jiao, Jianshe Wu |
Soft Comput. | 4 |
| 2014 | Local Maximal Homogeneous Region Search for SAR Speckle Reduction With Sketch-Based Geometrical Kernel FunctionabstractWith the flourish of the nonlocal mean method, the neighborwise similarity metric is widely applied in speckle reduction for its robust performance on the search of similar samples. In this metric, an isotropic kernel function is usually chosen to aggregate the corresponding pixels' distance between two neighborhoods. It means that the kernel function is considered as the explanation of the local spatial relationship at each pixel. However, for anisotropic features (such as edges and lines), a strong relationship exists along their directions rather than across them, so the isotropic kernel is not suitable to explain the spatial relationship around these features. Meanwhile, due to the inherent speckle in synthetic aperture radar (SAR) images, the discrimination and exploration of the geometrical properties of anisotropic features are important for the construction of adaptive kernel function. In this paper, the sketch map which is a representation of the sketch information of SAR images is extracted as the criterion for designing the kernel function. Meanwhile, due to the properties of symmetric and maximal self-similarity, a modified ratio distance is proposed and used jointly with the constructed kernel function as a similarity metric. Then, under the local stationary assumption, the local maximal homogeneous region of each pixel is searched by using the region growing method with the proposed metric. Moreover, maximal likelihood rule is used within the region for the estimation of true value. From the experiments on the synthetic and real SAR images, a promising performance in terms of speckle reduction and preservation of the details is achieved by our proposed method. Jie Wu 0016, Fang Liu 0001, Licheng Jiao, Xiangrong Zhang, Hongxia Hao, Shuang Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2013 | A novel selection evolutionary strategy for constrained optimization
Licheng Jiao, Lin Li 0016, Ronghua Shang, Fang Liu 0001, Rustam Stolkin |
Inf. Sci. | 4 |
| 2013 | A co-evolutionary multi-objective optimization algorithm based on direction vectors
Licheng Jiao, Handing Wang, Ronghua Shang, Fang Liu 0001 |
Inf. Sci. | 4 |
| 2013 | A bi-level belief rule based decision support system for diagnosis of lymph node metastasis in gastric cancer
Fang Liu 0001, Licheng Jiao, Zhi-Jie Zhou 0001, Jian-Bo Yang, Maoguo Gong, Xiao-Peng Zhang |
Knowl. Based Syst. | 2 |
| 2013 | An efficient matrix bi-factorization alternative optimization method for low-rank matrix recovery and completion
Yuanyuan Liu 0001, Licheng Jiao, Fanhua Shang, Fang Liu 0001 |
Neural Networks | 5 |
| 2013 | Selective multiple kernel learning for classification with ensemble strategy
Tao Sun 0007, Licheng Jiao, Fang Liu 0001, Shuang Wang 0001, Jie Feng 0003 |
Pattern Recognit. | 3 |
| 2013 | Multivariate pursuit image reconstruction using prior information beyond sparsity
Jiao Wu 0002, Fang Liu 0001, Licheng Jiao, Xiaodong Wang 0011 |
Signal Process. | 2 |
| 2013 | Immune optimization algorithm for solving vertical handoff decision problem in heterogeneous wireless network
Fang Liu 0001, Si-Feng Zhu, Yutao Qi, Jianshe Wu |
Wirel. Networks | 1 |
| 2012 | Immune optimization algorithm for solving joint call admission control problem in next-generation wireless network
Si-Feng Zhu, Fang Liu 0001, Yutao Qi, Jianshe Wu |
Eng. Appl. Artif. Intell. | 2 |
| 2012 | An evidential reasoning based classification algorithm and its application for face recognition with class noise
Xiaodong Wang 0011, Fang Liu 0001, Licheng Jiao, Jingjing Yu 0001, Bing Li 0001, Jianrui Chen 0002, Jiao Wu 0002, Fanhua Shang |
Pattern Recognit. | 2 |
| 2012 | A Novel Immune Clonal Algorithm for MO ProblemsabstractResearch on multiobjective optimization (MO) becomes one of the hot points of intelligent computation. Compared with evolutionary algorithm, the artificial immune system used for solving MO problems (MOPs) has shown many good performances in improving the convergence speed and maintaining the diversity of the antibody population. However, the simple clonal selection computation has some difficulties in handling some more complex MOPs. In this paper, the simple clonal selection strategy is improved and a novel immune clonal algorithm (NICA) is proposed. The improvements in NICA are mainly focus on four aspects. 1) Antibodies in the antibody population are divided into dominated ones and nondominated ones, which is suitable for the characteristic of one multiobjective optimization problem has a series Pareto-optimal solutions. 2) The entire cloning is adopted instead of different antibodies having different clonal rate. 3) The clonal selection is based on the Pareto-dominance and one antibody is selected or not depending on whether it is a nondominated one, which is different from the traditional clonal selection manner. 4) The antibody population updating operation after the clonal selection is adopted, which makes antibody population under a certain size and guarantees the convergence of the algorithm. The influences of the main parameters are analyzed empirically. Compared with the existed algorithms, simulation results on MOPs and constrained MOPs show that NICA in most problems is able to And much better spread of solutions and better convergence near the true Pareto-optimal front. Ronghua Shang, Licheng Jiao, Fang Liu 0001, Wenping Ma 0001 |
IEEE Trans. Evol. Comput. | 3 |
| 2011 | Interactive MOEA/D for multi-objective decision makingabstractIn this paper, an interactive version of the decomposition based multiobjective evolutionary algorithm (iMOEA/D) is proposed for interaction between the decision maker (DM) and the algorithm. In MOEA/D, a multi-objective problem (MOP) can be decomposed into several single-objective sub-problems. Thus, the preference incorporation mechanism in our algorithm is implemented by selecting the preferred sub-problems rather than the preferred region in the objective space. At each interaction, iMOEA/D offers a set of current solutions and asks the DM to choose the most preferred one. Then, the search will be guided to the neighborhood of the selected. iMOEA/D is tested on some benchmark problems, and various utility functions are used to simulate the DM's responses. The experimental studies show that iMOEA/D can handle the preference information very well and successfully converge to the expected preferred regions. Maoguo Gong, Fang Liu 0001, Wei Zhang 0009, Licheng Jiao, Qingfu Zhang 0001 |
GECCO | 2 |
| 2011 | Artificial immune multi-objective SAR image segmentation with fused complementary features
Licheng Jiao, Maoguo Gong, Fang Liu 0001 |
Inf. Sci. | 4 |
| 2011 | An improved multi-agent genetic algorithm for numerical optimization
Xiaoying Pan, Licheng Jiao, Fang Liu 0001 |
Nat. Comput. | 3 |
| 2011 | Compressive Sensing SAR Image Reconstruction Based on Bayesian Framework and Evolutionary ComputationabstractCompressive sensing (CS) is a theory that one may achieve an exact signal reconstruction from sufficient CS measurements taken from a sparse signal. However, in practical applications, the transform coefficients of SAR images usually have weak sparsity. Exactly reconstructing these images is very challenging. A new Bayesian evolutionary pursuit algorithm (BEPA) is proposed in this paper. A signal is represented as the sum of a main signal and some residual signals, and the generalized Gaussian distribution (GGD) is employed as the prior of the main signal and the residual signals. BEPA decomposes the residual iteratively and estimates the maximum a posteriori of the main signal and the residual signals by solving a sequence of subproblems to achieve the approximate CS reconstruction of the signal. Under the assumption of GGD with the parameter 0 < p < 1, the evolutionary algorithm (EA) is introduced to CS reconstruction for the first time. The better reconstruction performance can be achieved by searching the global optimal solutions of subproblems with EA. Numerical experiments demonstrate that the important features of SAR images (e.g., the point and line targets) can be well preserved by our algorithm, and the superior reconstruction performance can be obtained at the same time. Jiao Wu 0002, Fang Liu 0001, Licheng Jiao, Xiaodong Wang 0011 |
IEEE Trans. Image Process. | 2 |
| 2011 | Multivariate Compressive Sensing for Image Reconstruction in the Wavelet Domain: Using Scale Mixture ModelsabstractMost wavelet-based reconstruction methods of compressive sensing (CS) are developed under the independence assumption of the wavelet coefficients. However, the wavelet coefficients of images have significant statistical dependencies. Lots of multivariate prior models for the wavelet coefficients of images have been proposed and successfully applied to the image estimation problems. In this paper, the statistical structures of the wavelet coefficients are considered for CS reconstruction of images that are sparse or compressive in wavelet domain. A multivariate pursuit algorithm (MPA) based on the multivariate models is developed. Several multivariate scale mixture models are used as the prior distributions of MPA. Our method reconstructs the images by means of modeling the statistical dependencies of the wavelet coefficients in a neighborhood. The proposed algorithm based on these scale mixture models provides superior performance compared with many state-of-the-art compressive sensing reconstruction algorithms. Jiao Wu 0002, Fang Liu 0001, Licheng Jiao, Xiaodong Wang 0011, Biao Hou |
IEEE Trans. Image Process. | 2 |
| 2010 | A hybrid multiobjective immune algorithm with region preference for decision makersabstractRecently, one of the main tools of decision maker (DM) preference incorporation in the multiobjective optimization (MOO) has been using reference points and achievement scalarizing functions (ASF). The core idea of these methods is converting the original multiobjective problem (MOP) into single objective problem by using ASF to find a single preferred point. However, many DMs not only interest in a single point but also a set of efficient points in their preferred region. In this paper, we introduce a hybrid multiobjective immune algorithm (HMIA) for DM. It combines the immune inspired algorithm and region preference based on a novel dominance concept called region-dominance without ASF. The new algorithm can let DMs flexibly decide the number of reference points and accurately determine the preferred region with its simple and effective interactive methods. To exemplify its advantages, simulated results of HMIA are shown with some well-known problems. Licheng Jiao, Wei Zhang 0009, Ruochen Liu 0006, Fang Liu 0001 |
IEEE Congress on Evolutionary Computation | 4 |
| 2010 | Optimizing detector distribution in V-detector negative selection using a constrained multiobjective immune algorithmabstractIn this paper, a novel constrained multiobjective immune algorithm for optimizing detector distribution in V-detector negative selection is proposed. The theory of artificial immune system (AIS) and the spirit of population evolution are introduced to generate detectors. By combining the constraint handling technique and AIS-based multiobjective optimization, the algorithm is able to steadily maximize the anomaly coverage with little extra cost, which means the distribution with maximized coverage of the non-self space and minimized overlapping among detectors with fixed size will be well realized. Furthermore, the new approach is tested on some benchmark problems. The experimental results show that compared with some state-of-the-art methods, our algorithm can remarkably outperform them in terms of enhancing the detection rate by optimizing distribution without increasing the number of detectors. Fang Liu 0001, Maoguo Gong, Jingjing Ma 0001, Licheng Jiao, Wei Zhang 0009 |
IEEE Congress on Evolutionary Computation | 1 |
| 2010 | A sphere-dominance based preference immune-inspired algorithm for dynamic multi-objective optimizationabstractReal-world optimization involving multiple objectives in changing environment known as dynamic multi-objective optimization (DMO) is a challenging task, especially special regions are preferred by decision maker (DM). Based on a novel preference dominance concept called sphere-dominance and the theory of artificial immune system (AIS), a sphere-dominance preference immune-inspired algorithm (SPIA) is proposed for DMO in this paper. The main contributions of SPIA are its preference mechanism and its sampling study, which are based on the novel sphere-dominance and probability statistics, respectively. Besides, SPIA introduces two hypermutation strategies based on history information and Gaussian mutation, respectively. In each generation, which way to do hypermutation is automatically determined by a sampling study for accelerating the search process. Furthermore, The interactive scheme of SPIA enables DM to include his/her preference without modifying the main structure of the algorithm. The results show that SPIA can obtain a well distributed solution set efficiently converging into the DM's preferred region for DMO. Ruochen Liu 0006, Wei Zhang 0009, Licheng Jiao, Fang Liu 0001, Jingjing Ma 0001 |
GECCO | 4 |
| 2010 | Memetic computation based on regulation between neural and immune systems: the framework and a case study
Maoguo Gong, Licheng Jiao, Fang Liu 0001, Jie Yang 0011 |
Sci. China Inf. Sci. | 3 |
| 2010 | Immune algorithm with orthogonal design based initialization, cloning, and selection for global optimization
Maoguo Gong, Licheng Jiao, Fang Liu 0001, Wenping Ma 0001 |
Knowl. Inf. Syst. | 3 |
| 2010 | Multicontourlet-Based Adaptive Fusion of Infrared and Visible Remote Sensing ImagesabstractThis letter proposes a novel pixel-level adaptive remote sensing image fusion method based on multicontourlet transform. The multicontourlet that we constructed is a flexible multiscale and multidirection image decomposition. With better direction selectivity and energy convergence compared to that of a multiwavelet, a multicontourlet is suitable for representing remote sensing images bearing abundant detailed and directional information. The fusion weight of the low-pass coefficients is selected adaptively based on the golden section algorithm. For the high-frequency directional coefficients, the local energy feature is employed to select the better coefficients to fusion. Experimental results show that the proposed method achieves better visual quality and objective evaluation indexes than a wavelet-transform-based, a contourlet-transform-based, and a multiwavelet-transform-based weighted fusion method. Xia Chang, Licheng Jiao, Fang Liu 0001, Fangfang Xin |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2010 | SAR Image Despeckling Using Edge Detection and Feature Clustering in Bandelet DomainabstractTo effectively preserve the edges of a synthetic aperture radar (SAR) image when despeckling, an algorithm with edge detection and fuzzy clustering in the translation-invariant second-generation bandelet transform (TIBT) domain is proposed in this letter. A Canny operator is first utilized to detect and remove edges from the SAR image. Then, TIBT and fuzzy C-mean clustering are employed to decompose and despeckle the edge-removed image, respectively. Finally, the removed edges are added to the reconstructed image. The algorithm suggests each coefficient in high-frequency subbands as the clustering feature, proposes a calculation method of the best clustering number, and defines the signal and noise in the clustering results. Experimental results show that the visual quality and evaluation indexes outperform the other methods with no edge preservation. The proposed algorithm effectively realizes both despeckling and edge preservation and reaches the state-of-the-art performance. Wenge Zhang, Fang Liu 0001, Licheng Jiao, Biao Hou, Shuang Wang 0001, Ronghua Shang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2008 | Spectral Clustering Ensemble Applied to SAR Image SegmentationabstractSpectral clustering (SC) has been used with success in the field of computer vision for data clustering. In this paper, a new algorithm named SC ensemble (SCE) is proposed for the segmentation of synthetic aperture radar (SAR) images. The gray-level cooccurrence matrix-based statistic features and the energy features from the undecimated wavelet decomposition extracted for each pixel being the input, our algorithm performs segmentation by combining multiple SC results as opposed to using outcomes of a single clustering process in the existing literature. The random subspace, random scaling parameter, and Nystrom approximation for component SC are applied to construct the SCE. This technique provides necessary diversity as well as high quality of component learners for an efficient ensemble. It also overcomes the shortcomings faced by the SC, such as the selection of scaling parameter, and the instability resulted from the Nystrom approximation method in image segmentation. Experimental results show that the proposed method is effective for SAR image segmentation and insensitive to the scaling parameter. Xiangrong Zhang, Licheng Jiao, Fang Liu 0001, Liefeng Bo, Maoguo Gong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2005 | Directional self-learning of genetic algorithmabstractIn order to overcome the low convergence speed and prematurity of classical genetic algorithm, an improved method named directional self-learning of genetic algorithm (DSLGA) is proposed in this paper. Through the self-learning operator directional information was introduced in local search process. The search direction was guided by the false derivative of the function fitness. Using the four operators among the individuals, the best solution was updated continuously. In experiments, DSLGA was tested on 4 unconstrained benchmark problems, and the results were compared with the algorithms presented recently. It showed that DSLGA performs much better than the other algorithms both in the quality of the solutions and in the computational complexity. Yuheng Sha, Licheng Jiao, Fang Liu 0001 |
GECCO | 4 |
| 2005 | Immune Clonal Selection Wavelet Network Based Intrusion Detection
Fang Liu 0001 |
ICANN (1) | 1 |
| 2004 | Space-Time Multiuser Detection Combined with Adaptive Wavelet Networks over Multipath Channels
Ling Wang 0003, Licheng Jiao, Haihong Tao, Fang Liu 0001 |
ISNN (2) | 4 |
| 2004 | Performance Analysis of Recurrent Neural Networks Based Blind Adaptive Multiuser Detection in Asynchronous DS-CDMA Systems
Ling Wang 0003, Haihong Tao, Licheng Jiao, Fang Liu 0001 |
ISNN (2) | 4 |
| 2004 | A New Data Mining Method Using Organizational Coevolutionary Mechanism
Jing Liu 0006, Weicai Zhong, Fang Liu 0001, Licheng Jiao |
PAKDD | 3 |