VLDB 2026 Research / reviewers in the wild / expert
Xu Liu 0006
dblp:93/3167-6
· DBLP profile ↗
187ranked-venue papers
13as first author
177since 2021 · last 2026
0000-0002-8780-5455ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 74 · 5 first-author · 74 since 2021Graphics, computer vision, multimedia, augmented reality and games · 64 · 3 first-author · 63 since 2021Applied, interdisciplinary, general and emerging computing · 59 · 6 first-author · 50 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evolving Semantic Propagation for Aerial Semantic 3D Gaussian SplattingabstractSemantic understanding of large-scale aerial scenes represents a critical challenge in 3D computer vision, hindered by the prohibitive cost of dense annotation. This paper introduces EvoPropGS, a novel approach for the semantic segmentation of 3D Gaussian Splatting models that requires only minimal supervision. Our core insight is to leverage the inherent structural repetitions within aerial environments to propagate semantic information from a sparse set of annotations across the entire 3D scene. Our approach constructs a prompt library by pairing SAM-generated mask candidates with DINOv2 feature embeddings from annotated views. For unannotated regions, we generate pseudo-labels by matching region proposals with these featured prompts via cosine similarity. We then formulate optimal prompt selection as a discrete optimization problem solved via evolutionary search, guided by our novel fitness function that evaluates both 3D consistency and 2D semantic coherence. Extensive experiments demonstrate that EvoPropGS achieves accurate segmentation with only 2 percent annotated pixels. Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Licheng Jiao, Puhua Chen, Wenping Ma 0001, Shuyuan Yang 0001 |
AAAI | 3 |
| 2026 | SOAR: Semi-Supervised Open-Vocabulary Aerial Object Detection via Dual-Aware Enhanced Prior DenoisingabstractOpen-Vocabulary Object Detection (OVOD) shows promise in remote sensing (RS), but due to its unique value, there are challenges such as the predominance of background regions, sparse labels, limited semantic information, and difficulties in semi-supervised training. To tackle these challenges, we propose the Semi-Supervised Open-Vocabulary Aerial Object Detection with Dual-Perception Prior Denoising (SOAR), which explicitly models the background embeddings of each scene to indirectly construct foreground priors, thereby capitalizing on the abundant background information present in RS imagery. We further introduce a query enhancement module that integrates language and foreground prior information to enhance the effectiveness of query selection and feature augmentation. During the decoding stage of semi-supervised training, we perform denoising and reconstruction of the foreground priors to generate pseudo-labels that support the training process. Additionally, we address the sparsity of label information through expansion and aggregation techniques, further improving model performance. Experimental evaluations reveal that, in the open-vocabulary object detection task on the DIOR dataset, our method achieves a mean Average Precision (mAP) of 68.5% and Harmonic Mean (HM) of 55.9%, outperforming the previous state-of-the-art model’s mAP of 61.6% and HM of 53.6%. Our approach offers a novel solution to the open-vocabulary challenge in aerial object detection. Xu Liu 0006, Lingling Li 0002, Licheng Jiao |
AAAI | 1 |
| 2026 | HTTrack: Learning to Perceive Targets via Historical Trajectories in Satellite Video TrackingabstractIn recent years, the rapid progress of deep learning has driven notable advancements in satellite video tracking, a critical task for applications such as environmental monitoring, disaster management, and defense. Despite these strides, existing approaches remain constrained by their inability to handle dynamic challenges, such as target appearance variations, complex motion patterns, and occlusions. Traditional methods often suffer from static template matching or overly complex update mechanisms, compromising their robustness and practicality in real-world scenarios. To address these limitations, we propose a paradigm shift in satellite video tracking by integrating historical trajectory knowledge with visual features. This fusion enhances the tracker's perceptual understanding of targets over time, enabling more adaptive and resilient tracking. By aligning spatial, temporal, and cross-modal information, our approach effectively bridges the gap between fragmented observations and coherent tracking performance, even under challenging conditions like small target detection and cluttered backgrounds. Extensive experiments conducted on multiple satellite video tracking benchmarks demonstrate the superiority of our method, with HTTrack achieving success rates of 51.5% on SV248S, 52.9% on SatSOT, and 32.6% on VISO, significantly outperforming state-of-the-art trackers and marking a step forward in achieving robust, accurate, and scalable satellite video tracking. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
AAAI | 9 |
| 2026 | Rethinking Neuromorphic Object Detection with Hybrid Dynamic Interaction Transformers
Dianze Li, Jianing Li 0001, Xu Liu 0006, Zhaokun Zhou, Xiaopeng Fan 0001, Yonghong Tian 0001 |
Int. J. Comput. Vis. | 3 |
| 2026 | Enhancing few-shot segmentation via mask combination learning
Shuo Li 0010, Fang Liu 0034, Licheng Jiao, Xuejian Gou, Lingling Li 0002, Xu Liu 0006, Puhua Chen |
Neurocomputing | 6 |
| 2026 | Contrastive perception representation learning for image inpainting
Maoguo Gong, Jianzhao Li, Licheng Jiao, Xu Liu 0006, Fang Liu 0001 |
Knowl. Based Syst. | 5 |
| 2026 | Group Interaction Network With Wavelet Attention for Remote Sensing Image Change DetectionabstractRemote sensing image change detection (CD), as a pivotal technology for monitoring Earth’s surface dynamics, plays a crucial role in urban planning, resource management, and disaster assessment. Despite the success of deep learning-based methods, they still suffer from two significant limitations. Firstly, the inadequate exploitation of frequency-domain information restricts their ability to capture subtle structural and edge changes. Secondly, the suboptimal interaction strategies that predominantly rely on attention mechanisms or direct feature exchange that fails to fully model complex semantic differences between bi-temporal images. To address these challenges, we propose a Wavelet Attention-based Group Interaction Network (WAGINet), which leverages wavelet attention for joint frequency-spatial domain feature learning and uses a group-wise feature exchange mechanism to optimize bi-temporal interaction. The wavelet attention module decomposes features into high-low frequency components and emphasize important ones to improve edge-aware feature extraction. In the meanwhile, the group interaction strategy enables both channel-group and spatial-group feature exchange to capture the correlation between bi-temporal features while better protecting structural integrity, so that promotes more discriminative change representation. Experimental results on public LEVIR-CD and WHU-CD datasets show that WAGINet outperforms existing state-of-the-art methods, providing an effective solution for high-precision remote sensing image CD in complex scenarios. Code available at https://github.com/yizhilanmaodhh/WAGINet. Huihui Dong, Zongfang Ma, Sixiang Xu, Xu Liu 0006, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2026 | Like Human Rethinking: Contour Transformer AutoRegression for Referring Remote Sensing InterpretationabstractReferring remote sensing interpretation holds significant application value in various scenarios such as ecological protection, resource exploration, and emergency management. However, referring remote sensing expression comprehension and segmentation (RRSECS) faces critical challenges, including micro-target localization drift problem caused by insufficient extraction of boundary features in existing paradigms. Moreover, when transferred to remote sensing domains, polygon-based methods encounter issues such as contour-boundary misalignment and multi-task co-optimization conflicts problems. In this paper, we propose SeeFormer, a novel contour autoregressive paradigm specifically designed for RRSECS, which accurately locates and segments micro, irregular targets in remote sensing imagery. We first introduce a brain-inspired feature refocus learning (BIFRL) module that progressively attends to effective object features via a coarse-to-fine scheme, significantly boosting small-object localization and segmentation. Next, we present a language-contour enhancer (LCE) that injects shape-aware contour priors, and a corner-based contour sampler (CBCS) to improve mask-polygon reconstruction fidelity. Finally, we develop an autoregressive dual-decoder paradigm (ARDDP) that preserves sequence consistency while alleviating multi-task optimization conflicts. Extensive experiments on RefDIOR, RRSISD, and OPTRSVG datasets under varying scenarios, scales, and task paradigms demonstrate transformative performance gains: compared to the baseline PolyFormer, our proposed SeeFormer improves oIoU and mIoU by 27.58% and 39.37% for referring image segmentation and by 18.94% and 28.90% for visual grounding on the RefDIOR dataset. Jinming Chai, Licheng Jiao, Xiaoqiang Lu, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Wenping Ma 0001, Weibin Li 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Physics-Informed Matrix Factorization OperatorabstractMatrix factorization is a fundamental characterization model in machine learning and is usually solved using mathematical decomposition reconstruction loss. However, matrix factorization is a data-driven model whose results depend on data quality, making it susceptible to noise. Inspired by physics, the law of conservation of energy is used to introduce physical laws into matrix factorization, which is called Physics-informed Matrix Factorization operator (PiMF). The PiMF operator uses the heat conduction equation to construct the energy objective function for matrix factorization, thereby retaining the mathematical model's decomposition meaning and satisfying the interpretability of physics. The PiMF follows the physical laws, thereby suppressing irregular or sudden noise signals that violate these physical principles. The solutions of the PiMF operator include more comprehensive knowledge of mathematics and physics, which improves the ability to generalize complex data, especially for noisy data. We demonstrate the consistency of the energy objective function and the mathematical model, which verifies the feasibility of matrix factorization using physical energy laws. In addition, the physical interpretability of the PiMF operator is proved from the perspective of energy decline. This study proposes two practical algorithms for PiMF in classification and clustering tasks, enhancing the practicability of matrix factorization by incorporating task-specific prior information constraints. The experimental results of PiMF for classification and clustering demonstrate the advantages of the proposed operator. The importance of physics-informed matrix factorization is verified, especially for noisy data. Chenxi Tian, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Causality-inspired learning semantic segmentation in unseen domain
Pei He, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Ronghua Shang, Yuwei Guo 0001, Puhua Chen, Shuyuan Yang 0001 |
Pattern Recognit. | 4 |
| 2026 | Concept-Aware Learning for Weakly Supervised Video Anomaly Detection
Shuo Li 0010, Fang Liu 0034, Licheng Jiao, Jiahao Wang 0002, Xu Liu 0006, Lingling Li 0002, Puhua Chen |
Pattern Recognit. | 5 |
| 2026 | PC2F: Language-guided Progressive Calibration and Cascade Filtering for Remote Sensing Visual Grounding
Licheng Jiao, Xu Liu 0006, Xiaoqiang Lu, Shuo Li 0010, Xiaolin Tian 0002 |
Pattern Recognit. | 4 |
| 2026 | Deep semi-supervised relation preserving learning model
Chenxi Tian, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Fang Liu 0034, Shuyuan Yang 0001 |
Pattern Recognit. | 3 |
| 2026 | VCGPrompt: Visual Concept Graph-Aware Prompt Learning for Vision-Language Models
Mengjia Wang, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
Pattern Recognit. | 7 |
| 2026 | Vision-by-prompt: Context-aware dual prompts for composed video retrieval
Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
Pattern Recognit. | 8 |
| 2026 | TFBTrack: Target-Aware Foreground-Background Modeling for vision-language tracking
Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
Pattern Recognit. | 8 |
| 2026 | Language-guided modulation-update for semi-supervised semantic segmentation
Libo Yan, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Jiahao Wang 0002, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Xuejian Gou |
Pattern Recognit. | 8 |
| 2026 | ERFC: Energy-Aware Reinforcement Feedback Calibration for Zero-Shot CaptioningabstractZero-shot captioning aims to generate descriptive captions for unseen image and video data by leveraging the potential of visual language models (VLMs) and language models (LMs) without requiring task-specific training. It has emerged as a critical task, but its performance is often hindered by the inherent gap between the training distribution and unseen test data. The fundamental challenge lies in the model’s strong dependence on the marginal distribution of the training data, which leads to biased predictions when handling test samples. To address this issue, we propose an Energy-aware Reinforcement Feedback Calibration (ERFC) framework to calibrate the distribution and predictions of caption models from a novel energy perspective. The calibration process of ERFC is divided into two key components: 1) We first construct an Energy Stabilizer (ES) based on the caption model, where energy is considered a measure of the affinity between the input sample and the model’s learned distribution. ES iteratively adjusts the embedding features of the input sample using Langevin Dynamics, reducing its energy to implicitly align the model’s distribution with the unseen target domain. 2) We deploy a Reinforcement Calibrator (RC) to refine and calibrate the generated captions through a reward-feedback mechanism. RC leverages the expert CLIP model as a reward signal to assess the quality of the generated captions and employs the policy gradient algorithm to reward or penalize the model, thereby improving its performance. By iteratively combining energy-based optimization and reward-driven calibration, ERFC achieves superior zero-shot generalization capabilities, as demonstrated on image benchmarks such as MSCOCO, Flickr30K, and NoCaps, as well as video benchmarks such as MSR-VTT and MSVD. Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | KCI-Net: Knowledge-Based Contourlet Inference Network for Super-ResolutionabstractTextural details are useful for image super-resolution, but massive CNN methods ignored the high-frequency components and generated over-smoothed outputs. The knowledge-based contourlet inference network is proposed in this paper. Different from other CNN-based methods that are directly infer high-resolution (HR) images, our model learns to reconstruct the HR image through the series of corresponding contourlet coefficients. Specifically, first, we consider the low-pass subbands of the contourlet as the corresponding low-resolution (LR) image. Then, feed it to the embedding net with residual blocks to provide adequate information for the contourlet coefficients prediction. Finally, we innovatively convert the estimation of contourlet coefficients into the estimation of the generalized gaussian distribution (GGD) parameters, and design the corresponding loss function to ensure training stability, which explores the smoothness of the contour effectively and guarantees the general structure and details of images. Experiments on four remote sensing datasets, four natural scenes and human-made content datasets, and the outdoor dataset demonstrate the superiority of the proposed model quantitatively and qualitatively. Mengkun Liu, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Shuang Wang 0001, Biao Hou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Knowledge-Aware Evolutionary TransformerabstractWith the Transformer architecture achieving impressive results in the vision domain. It has become a current popular research to explore more potentials of Transformer mixed architectures and explore more suitable mixed combinations. In this paper, Transformer classification network is designed and explored by multi-task architecture search algorithm. A new paradigm for multi-task architecture is designed by combining convolution and Transformer. The designed search architecture can combine the respective advantages of convolution and Transformer and can obtain better performance. At the same time, corresponding knowledge-aware multi-task genetic operators are designed to generate offspring individuals. Inter-task and inter-experience knowledge-aware is utilised to facilitate evolutionary convergence. During the search process, the reference evaluation method is utilised to reduce the redundant computation and time during the search process. In the experimental section, the search results are compared with state-of-the-art architectures and search algorithms. The experimental results confirm the effectiveness and high generalisation of the searched architectures. The ablation experimental part proves the effectiveness of the proposed architectural paradigm, genetic operators and reference evaluation. Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Evol. Comput. | 5 |
| 2026 | RefZVC: Refinable Zero-Shot Video Captioning by Test-Time Reinforcement PolishingabstractRecently, the zero-shot image captioning (zero-shot IC) method based on pre-trained visual language models (VLMs) and large language models (LLMs) has made significant progress. However, how to adapt it to the zero-shot video captioning (zero-shot VC) scenario (without video-text paired supervision) has not been well explored. Inspired by various recent test-time strategies (sacrificing additional test time to improve performance), we try to introduce a new paradigm of Test-time Reinforcement Polishing in zero-shot VC scenario. We take temporal dependency modeling as the starting point and propose a novel framework for Refinable Zero-shot VC, called RefZVC. RefZVC can greatly cover the long-term context of the video and continuously polish and refine the generated captions in a reward-feedback manner. We first design an Adaptive Frame Skipping module (AdaSkip) to skip redundant frames and select diverse keyframe sequences. Subsequently, we propose a Multi-granularity Reinforcement Polishing (MRP) mechanism, which iteratively polishes captions by leveraging Gaussian Kernel Cache (GKC) to capture temporal dynamics, store and reuse relevant historical context. In addition, MRP calculates rewards for generated captions at both the sentence-level and entity-level to achieve test-time polishing. With the MRP mechanism, RefZVC achieves superior zero-shot generalization performance, outperforming previous zero-shot VC methods on benchmarks such as MSVD, MSR-VTT, and VATEX. Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Lingling Li 0002, Xu Liu 0006 |
IEEE Trans. Image Process. | 6 |
| 2026 | Image Singularity Scattering Representation Learning ClassificationabstractThe multi-scale geometric analysis is a great representation tool. It can be used to improve the feature representation and learning process of deep networks. In addition to extracting features, the multi-scale geometric prior knowledge can also be used for the structure improvement of deep networks. In this paper, we propose a multi-scale scattering representation learning network, abbreviated as MSRLN, for image classification tasks. The exploration of structure improvement can be made with multi-scale scattering operations. In this way, the better singularity representation learning process for networks can be achieved. Firstly, the filter banks and multi-scale scattering operator are introduced for non-linear and singularity representation. Secondly, the novel multi-scale scattering representation learning network structure is designed. The scaling- wise scattering process is deployed in the shallow layer as a non-linear layer. This structure essentially supplements deep networks with geometric prior knowledge. It can further improve the non-linear activation and singularity representation process. Thirdly, we put forward the multi-stage scattering representation strategy and the prior knowledge weakening mechanism. With flexible scaling factors and learning rates, the stepwise approximation and learning process of networks can be achieved. In sum, MSRLN is a kind of structural innovative, and the scattering singularity representation structure can be extended to other backbones or tasks. Extensive experimental results show that MSRLN can achieve better image classification accuracy. Finally, necessary convergence, insight, and adaptability analyses are provided in evaluation experiments. Jie Gao 0013, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Puhua Chen, Yuwei Guo 0001, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | Learning to Prompt With Refining Text Knowledge for Zero-Shot Video Action RecognitionabstractFoundational vision-language models (VLMs) like CLIP are redefining the vision domain with their exceptional generalization capabilities. Prompt-based learning methods adapt pre-trained VLMs to video action recognition tasks using task-specific learnable text tokens. However, these tokens often struggle to generalize to unseen categories, as they tend to forget general textual knowledge. To address this, we construct knowledge prompts composed of handcrafted and descriptive prompts and introduce a novel knowledge-guided context mapping to enhance the generalization of learnable prompts to unseen categories. This approach mitigates the forgetting of fundamental knowledge by reducing the discrepancy between learnable prompts and knowledge prompts while simultaneously allowing the prompts to extract rich contextual knowledge from LLM data. Then, incorporating the knowledge-guided context mapping into the contrastive loss enables zero-shot transfer of prompts to new categories and data, providing discriminative prompts for both seen and unseen tasks. In addition, we propose an advanced temporal aggregation method that refines uniform mean pooling by incorporating frame-level textual relevance scoring. Extensive evaluations on multiple benchmarks demonstrate that learning to prompt with refining text knowledge is an effective quick-tuning method, achieving superior sample generalization performance without increasing training parameters. Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
IEEE Trans. Multim. | 8 |
| 2026 | Adaptive Multi-Modal Visual Tracking With Dynamic Semantic PromptsabstractRGB-based object tracking is a fundamental task in computer vision, aiming to identify, locate, and continuously track objects of interest across sequential video frames. Despite the significant advancements in the performance of traditional RGB trackers, they still face challenges in maintaining accuracy and robustness in the presence of complex backgrounds, occlusions, and rapid movements. To tackle these challenges, combining visual auxiliary modalities has gained significant attention. Beyond this, integrating natural language information offers additional advantages by providing high-level semantic context, enhancing robustness, and clarifying target priorities, further elevating tracker performance. This work proposes theAdaptiveMulti-modalVisual Tracking with Dynamic Semantic Prompts (AMVTrack) tracker, which efficiently incorporates image descriptions and avoids text dependency during tracking to improve flexibility and adaptability. AMVTrack significantly reduces computational resource consumption by freezing the parameters of the image encoder, text encoder, and Box Head and only optimizing a few learnable prompt parameters. Additionally, we introduce the Adaptive Dynamic Semantic Prompt Generator (ADSPG), which dynamically generates semantic prompts based on visual features, and theVisual-LanguageFusionAdaptation (V-L FA) method, which integrates multi-modal features to ensure consistency and complementarity of information. Additionally, we partition the Image Encoder to conduct an in-depth investigation into the relationship between the importance of features across different depth and width regions. Experimental results demonstrate that AMVTrack achieves significant performance improvements on multiple benchmark datasets, proving its effectiveness and robustness in complex scenarios. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
IEEE Trans. Multim. | 8 |
| 2026 | Adaptive Visual Prompting for Effective Satellite Video TrackingabstractSatellite video tracking presents significant challenges due to unpredictable target variations, environmental disturbances, and occlusions. Existing approaches either rely on auxiliary modalities or require full fine-tuning of foundation models, resulting in excessive parameter sensitivity and poor generalization. Meanwhile, conventional prompt-based tuning only updates parameters at a single location, limiting its ability to adapt to complex appearance changes. To address these limitations, we propose Adaptive Visual Prompting for Effective Satellite Video Tracking (AVPTrack). Unlike conventional prompts, introduced Super Prompts dynamically refine the original template at multiple distinct positions. This multi-location adaptation allows for fine-grained representation learning, enabling the tracker to better capture target variations and resist environmental disturbances. Additionally, Dynamic Templates are introduced to mitigate tracking failures in highly challenging scenarios, such as occlusions and background clutter, ensuring robust target localization. Furthermore, the Template Selection Adapter (TSA) selects the most relevant templates in real-time, enhancing tracking efficiency. These components are optimized during training while keeping other parameters frozen, ensuring parameter efficiency. We also investigate the relationship between fine-tuning proportions and learning rates to optimize model performance. Extensive evaluations on the SV248S, SatSOT, and VISO datasets demonstrate the superior adaptability and robustness of AVPTrack compared to existing methods. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Yanbiao Ma, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Mengjia Wang |
IEEE Trans. Multim. | 9 |
| 2026 | Regularized-Aware Discriminative Transformer Tracker for Satellite Videos
Xin Zhang 0167, Licheng Jiao, Lingling Li 0002, Zhongjian Huang, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 6 |
| 2026 | Multiscale Spatial-Frequency Learning for Degradation Decoupling in RS Image RestorationabstractRemote sensing (RS) images are prone to various degradations, which poses challenges to downstream tasks. Although existing single-task remote sensing image restoration methods are effective, they lack generalizability across tasks. All-in-one methods can handle multiple degradation tasks, but they usually focus on spatial information, ignoring the physical properties of the degradation information. To address the above limitations, we propose a Multiscale Spatial-Frequency Degradation Decoupling framework for All-in-One remote sensing image restoration (SFD$^{2}$IR), which decouples degradation features across different tasks to guide the model in performing task-specific image restoration. Specifically, a task-specific instruction generator (TIG) is proposed first to transform degradation features into task-specific prompts. Then, a multi-scale multi-frequency enhancement (MME) module is designed to decouple degradation effects from both spatial and frequency perspectives, thus enhancing the model's adaptability to various degradation types. Finally, a prompt feature refinement (PFR) module is developed to further refine the model's response to degraded tasks. Extensive experiments demonstrate that the proposed method achieves excellent performance on different RSIR tasks, including cloud removal, deblurring, dehazing, and super-resolution. The source code will be publicly available at SFD$^{2}$IR. Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 4 |
| 2026 | Edit-Your-Motion: Space-Time Diffusion Decoupling Learning for Video Motion EditingabstractExisting diffusion-based methods have achieved impressive results in human motion editing. However, these methods often exhibit significant ghosting and body distortion in unseen in-the-wild cases. In this paper, we introduce Edit-Your-Motion, a video motion editing method that tackles these challenges through one-shot fine-tuning on unseen cases. Specifically, firstly, we utilized DDIM inversion to initialize the noise, preserving the appearance of the source video and designed a lightweight motion attention adapter module to enhance motion fidelity. DDIM inversion aims to obtain the implicit representations by estimating the prediction noise from the source video, which serves as a starting point for the sampling process, ensuring the appearance consistency between the source and edited videos. The Motion Attention Module (MA) enhances the model's motion editing ability by resolving the conflict between the skeleton features and the appearance features. Secondly, to effectively decouple motion and appearance of source video, we design a spatio-temporal two-stage learning strategy (STL). In the first stage, we focus on learning temporal features of human motion and propose recurrent causal attention (RCA) to ensure consistency between video frames. In the second stage, we shift focus on learning the appearance features of the source video. With Edit-Your-Motion, users can edit the motion of humans in the source video, creating more engaging and diverse content. Extensive qualitative and quantitative experiments, along with user preference studies, show that Edit-Your-Motion outperforms other methods. Yi Zuo 0003, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Wenping Ma 0001, Shuyuan Yang 0001, Yuwei Guo 0001 |
IEEE Trans. Multim. | 5 |
| 2026 | PromptVAD: Abnormal Prompt via Vision-Language ModelabstractWeakly supervised video anomaly detection (WSVAD) aims at predicting frame-level anomaly scores by modeling training videos with video-level annotations. The category names of abnormal events contain high-level knowledge abstracted by humans about abnormalities, which is of great help in identifying abnormal events. To utilize the knowledge implicit in category names, based on the visual-language pretraining model, we introduce a learnable abnormal prompt from three aspects: learnable domain prompt, learnable category prompt, and nonlearnable category definition prompt. Based on the learnable abnormal prompt, we propose a novel fine-grained WSVAD method: PromptVAD, which exploits a learnable abnormal prompt to reduce the semantic gap between visual images and anomaly categories. Through a similarity measure and our proposed coarse-grained two-class prompt module, our PromptVAD jointly learns coarse-grained and fine-grained VAD. Extensive experimental results on the ShanghaiTech, University of Central Florida (UCF)-Crime, and XD-Violence datasets show that our method achieves state-of-the-art performance. Specifically, our method achieves an area under the curve (AUC) of 88.62% on the UCF-Crime dataset. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Zehua Hao, Jiahao Wang 0002, Lingling Li 0002, Xu Liu 0006, Puhua Chen |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2026 | DGNMF: Dynamic Diffusion Graph Nonnegative Matrix FactorizationabstractIn feature learning (FL), structural information shows advantages in retaining information and maintaining stability. Graph diffusion, a graph learning method that can focus on neighborhood structure and transmit information, has great research potential. In this study, a novel dynamic diffusion graph nonnegative matrix factorization (DGNMF) method is proposed, which uses a diffusion graph to improve the performance of FL and further enhances the effectiveness and stability of downstream classification tasks. DGNMF aims to mine and retain structural information more deeply in FL to build a more powerful and stable FL method. First, the model embeds graph learning into FL to obtain features containing structural information. Second, dynamic diffusion graph learning is used to mine deeper and more global structural information. Finally, we construct an updateable indicator matrix to enhance the discriminability of features. The classification experimental results of DGNMF on six databases demonstrate its advantages, verify its effectiveness and stability, and prove the importance of diffusion graph in improving FL. Chenxi Tian, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2026 | Spatial-Temporal Diffusion Model for Matrix FactorizationabstractMatrix factorization (MF) is a fundamental problem in machine learning, which is usually used as a feature learning method in various fields. For complex data involving spatiotemporal interactions, MF that only handles 2-D data will disrupt spatial dependence or temporal dynamics, failing to effectively couple spatial information with temporal factors. According to Markov chain principle, the spatial information of the present time is related to the spatial state of the previous time. We propose a spatial-temporal diffusion model for MF (STDMF), which uses graph diffusion to couple spatial-temporal information. Then, MF is used to learn the joint feature of data and spatial-temporal diffusion graph. Specifically, STDMF utilizes the graph diffusion with physical laws to generate spatial-temporal structure information. It obtains the underlying core structure of complex systems from a global perspective, which enhances the generalization ability of MF in noisy time-series data. To learn the lowest rank subspace of MF in time-series data, STDMF uses structural learning to constrain the rank of the learned features. Finally, STDMF is applied to clustering and anomaly detection of dynamic graph. The effectiveness of this method is verified by sufficient experiments, especially for noisy data. Chenxi Tian, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Licheng Jiao, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Logits DeConfusion with CLIP for Few-Shot LearningabstractWith its powerful visual-language alignment capability, CLIP performs well in zero-shot and few-shot learning tasks. However, we found in experiments that CLIP’s logits suffer from serious inter-class confusion problems in down-stream tasks, and the ambiguity between categories seriously affects the accuracy. To address this challenge, we propose a novel method called Logits DeConfusion, which effectively learns and eliminates inter-class confusion in logits by combining our Multi-level Adapter Fusion (MAF) module with our Inter-Class Deconfusion (ICD) module. Our MAF extracts features from different levels and fuses them uniformly to enhance feature representation. Our ICD learnably eliminates inter-class confusion in logits with a residual structure. Experimental results show that our method can significantly improve the classification performance and alleviate the inter-class confusion problem. The code is available at https://github.com/LiShuo1001/LDC. Shuo Li 0010, Fang Liu 0001, Zehua Hao, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
CVPR | 6 |
| 2025 | Asynchronous Collaborative Graph Representation for Frames and EventsabstractIntegrating frames and events has become a widely accepted solution for various tasks in challenging scenarios. However, most multimodal methods directly convert events into image-like formats synchronized with frames and process each stream through separate two-branch backbones, making it difficult to fully exploit the spatiotemporal events while limiting inference frequency to the frame rate. To address these problems, we propose a novel asynchronous collaborative graph representation, namely ACGR, which is the first trial to explore a unified graph framework for asynchronously processing frames and events with high performance and low latency. Technically, we first construct uni-modal graphs for frames and events to preserve their spatiotemporal properties and sparsity. Then, an asynchronous collaborative alignment module is designed to align and fuse frames and events into a unified graph and the ACGR is generated through graph convolutional networks. Finally, we innovatively introduce domain adaptation to enable cross-modal interactions between frames and events by aligning their feature spaces. Experimental results show that our approach outperforms state-of-the-art methods in both object detection and depth estimation tasks, while significantly reducing computational latency and achieving real-time inference up to 200 Hz. Our code can be available at https://github.com/dianzl/ACGR. Dianze Li, Jianing Li 0001, Xu Liu 0006, Xiaopeng Fan 0001, Yonghong Tian 0001 |
CVPR | 3 |
| 2025 | Knowledge-Guided Part Segmentation
Xuejian Gou, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, Hao Wang 0211, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
ICCV | 7 |
| 2025 | Domain-Aware Category-Level Geometry Learning Segmentation for 3D Point Clouds
Pei He, Lingling Li 0002, Licheng Jiao, Ronghua Shang, Fang Liu 0001, Shuang Wang 0001, Xu Liu 0006, Wenping Ma 0001 |
ICCV | 7 |
| 2025 | FPEM: Face Prior Enhanced Facial Attractiveness Prediction for Live Videos with Face Retouching
Hongjiu Yu, Ying Chen 0011, Kai Li 0012, Xiongkuo Min, Huiyu Duan, Guangtao Zhai, Xu Liu 0006 |
ICCV | 10 |
| 2025 | Language-Guided Hybrid Representation Learning for Visual Grounding on Remote Sensing ImagesabstractVisual grounding (VG) refers to detecting the specific objects in images based on linguistic expressions, and it has profound significance in the advanced interpretation of natural images. In remote sensing image interpretation, visual grounding is limited by characteristics such as the complex scenes and diverse object sizes. To solve this problem, we propose a novel remote sensing visual grounding (RSVG) framework, named language-guided hybrid representation learning Transformer (LGFormer). Specifically, we designed a multimodal dual-encoder Transformer structure called the adaptive multimodal feature fusion module. This structure innovatively integrates text and visual features as hybrid queries, enabling early-stage decoding queries to perceive the target position accurately. Then, the different modal information from the dual encoders is aggregated by hybrid queries to obtain the final object embedding for coordinate regression. Besides, a multi-scale cross-modal feature enhancement module (MSCM) is designed to enhance the self-representation of the extracted text and visual features and align them semantically. As for the hybrid queries, we use linguistic guidance to select visual features as the visual part and sentence-level features as the textual part. Finally, the LGFormer model we designed achieved the best results compared to existing models on the DIOR-RSVG and OPT-RSVG datasets. Xu Liu 0006, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Youlin Huang |
IJCAI | 2 |
| 2025 | FA3T: Feature-Aware Adversarial Attacks for Multi-modal TrackingabstractMulti-modal visual tracking leverages complementary sensor information to enhance robustness under challenging conditions. However, the security of multi-modal tracking systems remains largely unexplored. Existing attacks primarily target single-modal trackers or independently disrupt each modality, failing to exploit the inherent feature interactions and fusion mechanisms that define multi-modal tracking. As a result, these methods exhibit limited attack effectiveness and fail to assess multi-modal tracking systems' vulnerabilities accurately. Understanding these security risks is crucial, as adversarial threats could lead to severe failures in safety-critical applications. To address these challenges, a feature-aware adversarial attack, termed FA3T is proposed. It is designed to explicitly disrupt feature extraction and cross-modal alignment, thereby weakening the fusion process that multi-modal trackers rely on. To achieve this, a Frequency-Spatial Feature Separation (FSFS) module is constructed to perturb feature representations at multiple levels, weakening the modality-complementary advantages of multi-modal tracking. Furthermore, a Target Confusion Attack (TCA) module is devised to manipulate the target-background-template relationships, making it increasingly difficult for the tracker to distinguish the true target, significantly impairing tracking performance. Extensive experiments on five benchmark datasets (i.e., LasHeR, RGBT234, DepthTrack, VOT-RGBD2022, VisEvent) across three different modalities (RGB-T, RGB-D, and RGB-E) demonstrate that our attack substantially degrades state-of-the-art multi-modal trackers, exposing their susceptibility to adversarial threats. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
ACM Multimedia | 8 |
| 2025 | Imagining Vision From Language for Few-Shot Class-Incremental Learning
Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Yanbiao Ma, Puhua Chen, Lingling Li 0002, Xu Liu 0006, Xuejian Gou |
ACM Multimedia | 10 |
| 2025 | High-Rate Monocular Depth Estimation via Cross Frame-Rate Collaboration of Frames and Events
Xu Liu 0006, Xiaopeng Fan 0001, Jianing Li 0001, Dianze Li, Wei Zhang 0161, Zhengyu Ma, Yonghong Tian 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | Preserving text space integrity for robust compositional zero-shot learning via mixture of pretrained experts
Zehua Hao, Fang Liu 0001, Licheng Jiao, Yaoyang Du, Shuo Li 0010, Hao Wang 0211, Pengfang Li, Xu Liu 0006, Puhua Chen |
Neurocomputing | 8 |
| 2025 | A Physics-Aware Collaborative Framework With Prototype Consistency for Noisy Label Signal Modulation ClassificationabstractSignal Modulation Classification (SMC) is a fundamental technique in wireless communications. However, the prevalence of label noise in practical scenarios severely constrains the advancement of SMC technology. Existing SMC methods heavily rely on high-quality labeled data and often underutilize the inherent physical prior knowledge of signals. To address these issues, this article proposes a Physics-Aware Collaborative Framework with Prototype Consistency (PhyCo-PC), designed for noisy label environments and operating without requiring reliable labels. Firstly, the framework leverages co-teaching for noise identification and incorporates a collaborative consensus-guided module for prototype learning and pseudo-label generation. Secondly, it constructs physics-guided downstream decision module that fuses deep learning features with instantaneous physical signal characteristics to enhance decision robustness. Thirdly, a Domain Knowledge-guided Adaptive Sample Selection (DKASS) strategy is introduced. DKASS parameterizes the selection rate scheduling function, incorporates domain knowledge to constrain the search space, and utilizes automated search for optimization. This enables the model to adaptively determine the optimal training strategy for varying noise environments. Finally, experimental results demonstrate that PhyCo-PC significantly improves SMC classification performance under complex label noise scenarios on the RML2016.10a/04c datasets, exhibiting excellent robustness and significant advantages. Lingling Li 0002, Jiadong Lin, Huaji Zhou, Xu Liu 0006, Fang Liu 0001, Licheng Jiao |
IEEE Internet Things J. | 5 |
| 2025 | LLM Knowledge-Driven Target Prototype Learning for Few-Shot Segmentation
Pengfang Li, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Xu Liu 0006, Puhua Chen, Lingling Li 0002, Zehua Hao |
Knowl. Based Syst. | 5 |
| 2025 | Knowledge-aware evolutionary graph neural architecture search
Chao Wang 0099, Jiaxuan Zhao, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
Knowl. Based Syst. | 6 |
| 2025 | Visual-Language Scene-Relation-Aware Zero-Shot CaptionerabstractZero-shot image captioning can harness the knowledge of pre-trained visual language models (VLMs) and language models (LMs) to generate captions for target domain images without paired sample training. Existing methods attempt to establish high-quality connections between visual and textual modalities in text-only pre-training tasks. These methods can be divided into two perspectives: sentence-level and entity-level. Although they achieve effective performance on some metrics, they suffer from hallucinations due to biased associations during training. In this paper, we propose a scene-relation-level pre-training task by considering relations as more valuable modal connection bridges. Based on this, we construct a novel Visual-Language Scene Relation Aware Captioner (SRACap), which expands the ability to predict scene relations while generating captions for images. In addition, SRACap possesses excellent cross-domain zero-shot generalization capability, which is driven by a well-designed scene reinforcement switching pipeline. We introduce a scene policy network to dynamically crop salient regions from images and feed them into a language model to generate captions. We integrate multiple expert CLIP models to form a mixture-of-rewards module (MoR) as a reward source, and deeply optimized SRACap through the policy gradient algorithm in the zero-shot inference stage. With the iteration of scene reinforcement switching, SRACap can gradually refine the generated caption details while maintaining high semantic consistency across visual-linguistic modalities. We conduct extensive experiments on multiple standard image captioning benchmarks, showing that SRACap can accurately understand scene structures and generate high-quality text, significantly outperforming other zero-shot inference methods. Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Unveiling and Mitigating Generalized Biases of DNNs Through the Intrinsic Dimensions of Perceptual ManifoldsabstractBuilding fair deep neural networks (DNNs) is a crucial step towards achieving trustworthy artificial intelligence. Delving into deeper factors that affect the fairness of DNNs is paramount and serves as the foundation for mitigating model biases. However, current methods are limited in accurately predicting DNN biases, relying solely on the number of training samples and lacking more precise measurement tools. Here, we establish a geometric perspective for analyzing the fairness of DNNs, comprehensively exploring how DNNs internally shape the intrinsic geometric characteristics of datasets-the intrinsic dimensions (IDs) of perceptual manifolds, and the impact of IDs on the fairness of DNNs. Based on multiple findings, we propose Intrinsic Dimension Regularization (IDR), which enhances the fairness and performance of models by promoting the learning of concise and ID-balanced class perceptual manifolds. In various image recognition benchmark tests, IDR significantly mitigates model bias while improving its performance. Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Lingling Li 0002, Wenping Ma 0001, Shuyuan Yang 0001, Xu Liu 0006, Puhua Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Predicting and Enhancing the Fairness of DNNs With the Curvature of Perceptual ManifoldsabstractTo address the challenges of long-tailed classification, researchers have proposed several approaches to reduce model bias, most of which assume that classes with few samples are weak classes. However, recent studies have shown that tail classes are not always hard to learn, and model bias has been observed on sample-balanced datasets, suggesting the existence of other factors that affect model bias. In this work, we first establish a geometric perspective for analyzing model fairness and then systematically propose a series of geometric measurements for perceptual manifolds in deep neural networks. Subsequently, we comprehensively explore the effect of the geometric characteristics of perceptual manifolds on classification difficulty and how learning shapes the geometric characteristics of perceptual manifolds. An unanticipated finding is that the correlation between the class accuracy and the separation degree of perceptual manifolds gradually decreases during training, while the negative correlation with the curvature gradually increases, implying that curvature imbalance leads to model bias. We thoroughly validate this finding across multiple networks and datasets, providing a solid experimental foundation for future research. We also investigate the convergence consistency between the loss function and curvature imbalance, demonstrating the lack of curvature constraints in existing optimization objectives. Building upon these observations, we propose curvature regularization to facilitate the model to learn curvature-balanced and flatter perceptual manifolds. Evaluations on multiple long-tailed and non-long-tailed datasets show the excellent performance and exciting generality of our approach, especially in achieving significant performance improvements based on current state-of-the-art techniques. Our work opens up a geometric analysis perspective on model bias and reminds researchers to pay attention to model bias on non-long-tailed and even sample-balanced datasets. Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Maoji Wen, Lingling Li 0002, Wenping Ma 0001, Shuyuan Yang 0001, Xu Liu 0006, Puhua Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | Text generation and multi-modal knowledge transfer for few-shot object detection
Yaoyang Du, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Zehua Hao, Pengfang Li, Jiahao Wang 0002, Hao Wang 0211, Xu Liu 0006 |
Pattern Recognit. | 9 |
| 2025 | Knowledge-Driven Compositional Action Recognition
Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006 |
Pattern Recognit. | 7 |
| 2025 | VLPA-CLIP: Video Language Prompting and Adapting CLIP for efficient video action recognition
Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
Pattern Recognit. | 8 |
| 2025 | Temporal-Noise-Aware Neural Networks for Suicidal Ideation Prediction Using Physiological DataabstractThe robust generalization of deep learning models in the presence of inherent noise remains a significant challenge, especially when labels are ambiguous due to their subjective nature and noise is indiscernible in natural settings. In this article, we address a specific and important scenario of monitoring suicidal ideation (SI), where time-series data, such as galvanic skin response (GSR) and photoplethysmography (PPG), are susceptible to such noise. Current methods predominantly focus on image and text data or address artificially introduced noise, neglecting the complexities of natural noise in time-series analysis. To tackle this, we introduce a novel neural network model tailored for analyzing noisy physiological time-series data, named DBN_ConvNet, which integrates advanced encoding techniques with confidence learning training to enhance prediction performance. Another main contribution of our work is the collection of a specialized dataset of GSR and PPG signals derived from real-world environments for SI prediction. By employing this dataset, our DBN_ConvNet achieves a prediction accuracy of 76.67% and an F1 score of 0.74 in a binary classification task, outperforming state-of-the-art methods. Furthermore, comprehensive evaluations have been conducted on three other well-known public datasets with artificially introduced noise to test the DBN_ConvNet’s capabilities rigorously. These tests consistently demonstrated DBN_ConvNet’s superior performance by achieving an improvement of more than 10% in both accuracy and F1 score compared to the baseline methods. Niqi Liu, Fang Liu 0035, Xinxin Du, Yezhi Shu, Xu Liu 0006, Guozhen Zhao, Wenting Mu, Yong-Jin Liu 0001 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | Knowledge-Aware Geometric Contourlet Semantic Learning for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) provides detailed spectral and spatial information, essential for precise earth observation and various applications. Deep learning has advanced HSI classification, but the scarcity of labeled data and large model parameters necessitate semi-supervised methods to enhance performance and generalization. In this paper, we propose a novel semi-supervised framework dubbed Knowledge-Aware Geometric Contourlet Semantic Learning (KGCSL), aiming to achieve high-precision HSI classification with limited samples leveraging geometric and semantic knowledge. Specifically, to fully leverage geometric knowledge, KGCSL incorporates multi-scale and multi-directional representations of the contourlet transform within the neural network, enhancing the robustness of feature extraction and interpretability. Furthermore, to fully utilize semantic knowledge, an entropy-weighted prototype loss function is designed that exploits the attribute relationships between labeled and unlabeled samples to guide the optimization of unlabeled samples, promoting comprehensive semantic learning. Comprehensive evaluations of the proposed KGCSL framework on three public HSI datasets show that it outperforms existing state-of-the-art HSI classification methods and exhibits excellent generalization capabilities in limited-sample scenarios. The source code is available athttps://github.com/ShirlySmile/KGCSL. Xueli Geng, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Prompt-Based Concept Learning for Few-Shot Class-Incremental LearningabstractFew-Shot Class-Incremental Learning (FSCIL) faces a huge stability-plasticity challenge due to continuously learning knowledge from new classes with a small number of training samples without forgetting the knowledge of previously seen old classes. To alleviate this challenge, we propose a novel method called Prompt-based Concept Learning (PCL) for FSCIL, which generalizes conceptual knowledge learned from old classes to new classes by simulating human learning capabilities. In our PCL, in the base session, we simultaneously learn common basic concepts from the training data and the class-concept weight of each class in a prompt learning manner, and in each incremental session, class-concept weights between new classes and previously learned basic concepts are learned to achieve incremental learning. Furthermore, in order to avoid catastrophic forgetting, we propose a distribution estimation module to retain feature distributions of previously seen classes and a data replay module to randomly sample features of previously seen classes in incremental sessions. We verify the effectiveness of our PCL on widely used benchmarks, such as miniImageNet, CIFAR-100, and CUB-200. Experimental results show that our PCL achieves competitive results compared with other state-of-the-art methods, especially we achieve an average accuracy of 94.02% across all sessions on the miniImageNet benchmark. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Contour Knowledge-Aware Perception Learning for Semantic SegmentationabstractThe diversity of contextual information is of great importance for accurate semantic segmentation. However, most methods focus on single spatial contextual information, which results in an overlap of the semantic content of categories and a loss of contour information of objects. In this article, we propose a novel contour knowledge-aware perception learning network (CKPL-Net) to capture diverse contextual information by space-category aggregation module (SCAM) and contour-aware calibration module (CACM). First, SCAM is introduced to enhance intraclass consistency and interclass differentiation of features. By integrating space-aware and category-aware attention, SCAM reduces the redundancy of features from a categorical perspective while maintaining spatial correlation of pixels, substantially avoiding the overlap of the semantic content in categories. Second, CACM is designed to maintain the integrity of objects by perceiving contour contextual information. It develops a novel contour-aware knowledge and adaptively transforms the grid structure of convolutions for boundary pixels, which effectively calibrates the representation of features near boundaries. Finally, the quantitative and qualitative analyses on the three public datasets: ISPRS Potsdam dataset, ISPRS Vaihingen dataset, and WHDLD dataset, demonstrate that the proposed CKPL-Net achieves superior performance compared with prevalent methods, which indicates diverse contextual information is beneficial for accurate segmentation. Chao You, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Fine-Grained Visual-Language Alignment for Remote Sensing Image-Text RetrievalabstractRemote sensing image-text retrieval (RSITR) is critical for applications, including environmental monitoring and disaster management. The main challenge in this field is that the multi-scale feature of remote sensing images and the semantic differences of professional texts make it difficult to achieve accurate alignment. Existing coarse-grained methods struggle to address the inherent difference between images and text. In light of this, we propose the Fine-Grained Visual-Language Alignment (FGVLA) method. Our FGVLA employs a hybrid loss function that combines coarse-grained contrastive and triplet loss with novel fine-grained loss. Fine-grained loss includes spatial mask loss and fine-grained contrastive loss to enhance semantic alignment. The method also introduces an inference process that works cooperatively with fine-grained loss to explicitly align image patches with textual nouns. Extensive experiments on RSICD, RSITMD, and UCM-Caption datasets demonstrate that FGVLA outperforms existing methods, achieving superior retrieval performance. The code of our FGVLA has been released at https://github.com/Ji-Haoyang/FGVLA. Shuo Li 0010, Haoyang Ji, Fang Liu 0001, Licheng Jiao, Xutong Min, Jiahao Wang 0002, Lingling Li 0002, Xu Liu 0006 |
IEEE Trans. Geosci. Remote. Sens. | 10 |
| 2025 | LSCF: Long-Term Semantic-Guidance ConvFormer for Referring Remote Sensing Image SegmentationabstractReferring Remote Sensing Image Segmentation (RRSIS) task aims to generate segmentation masks for target objects based on language descriptions. It requires precise localization while distinguishing between visually similar yet semantically distinct objects. Fusing vision-language features only during extraction causes information loss and semantic forgetting in the decoder, harming similar target distinction. Additionally, high-resolution remote sensing images present challenges, including complex backgrounds, diverse object scales, and intricate boundaries, limiting the effectiveness of previous methods. To address these issues, we propose the Long-term Semantic-guidance ConvFormer (LSCF) Network. First, we fuse multi-receptive-field local features extracted by the Multi-scale CoordConv (MCC) module with language-aware global features from the Cross-modal Attention (CA) module to obtain multi-modal representations. Second, the Sampling Attention (SA) module enables fine-grained vision context alignment under semantic guidance. Finally, the Global Language Fusion (GLF) module is incorporated in the decoder to maintain long-term vision-language alignment and mitigate semantic degradation. Experimental validation on the RefSegRS, RRSIS-D, and RISBench datasets demonstrates that LSCF achieves oIoU scores of 83.27%, 77.42%, and 74.88%, and mIoU scores of 77.44%, 64.25%, and 68.53%, respectively. On RefSegRS, LSCF surpasses the SOTA method FIANet by 5.53% (oIoU) and 9.58% (mIoU), while delivering competitive performance on RRSIS-D and RISBench. Code and experimental configurations will be released. Lingling Li 0002, Xiaoqiang Lu, Licheng Jiao, Fang Liu 0001, Wenping Ma 0001, Xu Liu 0006 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | A Mamba-Aware Spatial-Spectral Cross-Modal Network for Remote Sensing ClassificationabstractThis study introduces a novel cross-modal spatial-spectral interaction Mamba (CMS2I-Mamba) for remote sensing image fusion classification. Unlike convolution-based models focusing on local details and Transformer-based models with high computational complexity, CMS2I-Mamba efficiently models global long-range dependencies in a linear complexity manner. First, multispectral (MS) and panchromatic (PAN) images each have unique advantages in the spectral and spatial attributes. Given this, this paper innovatively designs the multi-path selective-scan mechanism (MPS2M), which applies different path scanning strategies to deeply capture the global features from both spectral and spatial dimensions, aiming to enhance the robustness and complementarity of spatial-spectral features. Secondly, to overcome the characterization differences between images acquired by different sensors, this paper further introduces the channel interaction alignment module (CIAM). This module employs efficient former-last and oddeven channel interaction strategies to achieve precise semantic alignment of deep features between modalities. Finally, to leverage the shared fusion features to guide the unique singular features, this paper proposes a semantic-aware calibration module (SACM), which accurately constraints and calibrates the same semantic information in deep features. This not only enhances the model’s ability to understand scene semantics, but also promotes the deep fusion and utilization of information between different modalities. Through experimental verification on multiple datasets, the CMS2I-Mamba proposed in this paper shows excellent recognition performance and computational efficiency (parameter quantity and running speed) in fusion classification tasks. The code for CMS2I-Mamba is available at: https://github.com/ru-willow/CMSI-Mamba. Mengru Ma, Jiaxuan Zhao, Wenping Ma 0001, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Change Knowledge-Guided Vision-Language Remote Sensing Change DetectionabstractRemote sensing image change detection plays a critical role in applications like video surveillance and geographic information systems. However, existing binary and semantic change detection methods often rely solely on visual information, neglecting language information, which limits interpretability and the ability to provide specific change details. This work proposes the Change Knowledge-Guided Vision-Language Remote Sensing Change Detection (CKCD) method to address these limitations. By introducing change knowledge as language information, CKCD enhances semantic understanding and change detail representation. A Cross-Modal Affinity (CMA) module is designed to effectively fuse visual and textual features, improving information complementarity and fusion coherence. CKCD further enhances data utilization efficiency by merging change area detection and change category information into a single output through endto- end learning. This design reduces redundant data representations and simplifies the detection process, leading to a more compact and efficient use of the input data without requiring additional branches or multiple output heads. Experimental results demonstrate consistent performance improvements over traditional methods across multiple change detection datasets. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | Dual Causal-Aware Detection Transformer for Remote Sensing ImagesabstractDeep neural networks often inherit biases from training data, compromising generalization. In visual recognition, distinguishing foreground from background is insufficient, as models tend to rely on spurious correlations rather than learning essential causal patterns. To address this issue, this paper proposes a novel transformer architecture, termed dual causal-aware detection transformer (DCDT), specifically designed for object detection in optical remote sensing images from a causal perspective. Specifically, we begin by constructing a structural causal model to intuitively analyze the causal effects inherent in the overall visual patterns. Building on this foundation, DCDT introduces causal constraints at the attention level by embedding dynamic multi-scale causal prototypes into the attention mechanism. The derived causal priors are subsequently used to enhance features at the representation level, thereby enforcing feature-level causal modulation. This dual causal-aware strategy enables the precise extraction and reinforcement of causally relevant features, improving both robustness and discriminative capability in complex detection scenarios. In addition, a sparse kernel-region mask is incorporated to decouple local information from global representations, effectively strengthening the modeling of fine-grained structures. Extensive experiments conducted on two challenging public datasets, DIOR and HRRSD, demonstrate that DCDT consistently outperforms existing methods and baselines. These results validate the effectiveness of DCDT in capturing both global causal semantics and local fine-grained features, highlighting its practicality in complex remote sensing scenarios. Yuhan Wang 0007, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Zhongjian Huang, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Local-Global Spectral Feature-Aware Learning for Hyperspectral Imagery ClassificationabstractEffective modeling of the relationship between local spectral details and global contextual information remains a core challenge in hyperspectral image (HSI) classification. In this paper, a local-global spectral feature-aware network (LGSFA-Net) is proposed, which achieves local and global spectral feature learning through synergistic integration of local convolutional inductive biases and global state-space models (SSMs). The architecture of LGSFA-Net comprises three sequential components, including an embedding stage using convolutions for fundamental feature extraction, an encoding stage that cascades standard Mamba blocks with specialized interactive Mamba (IMamba) blocks and an enhanced spatial-spectral feature fusion (ESSFF) module. The proposed IMamba blocks employ separable convolutions and feature interaction learning for explicitly modeling the cross-channel spectral correlations learning, which can be effective in awareness of the spectral feature. And then, the ESSFF module utilizes self-attention mechanisms to dynamically balance local and global spatial-spectral feature contributions. The final prediction stage incorporates a lightweight classification head for efficient inference. Experimental results validate the effectiveness of the proposed methods for HSI classification on four benchmark datasets, including the PaviaU, Houston, Honghu, and Hanchuan datasets. The proposed LGSFA-Net achieves approximately 1.48%-2.51% increased overall accuracy (OA), 1.34%-2.06% increased average accuracy(AA), and 1.37%-3.75% increased Kappa on the aforementioned four datasets, respectively, outperforming the contrasting methods. The code implementation will be available at https://github.com/yutinyang/LGSFA-Net. Yuting Yang 0008, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Fang Liu 0001, Shuo Li 0010, Wenping Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | CSCT: Channel-Spatial Coherent Transformer for Remote Sensing Image Super-ResolutionabstractRemote sensing image super-resolution (RSISR) techniques are crucial in practice as an economical approach to enhancing the resolution of remote sensing images (RSIs). The scale of structural information and the richness of texture details in RSIs far exceed those in natural images. Therefore, accurately restoring and preserving edge and detail information are a critical challenge in the super-resolution (SR) process. Currently, convolutional neural network (CNN)-based methods primarily rely on local feature extraction, which fails to effectively capture and integrate global contextual information. Generative adversarial network (GAN)-based methods, while improving the visual quality, often suffer from artifacts and training instability, adversely affecting image quality. Moreover, these approaches struggle to accurately represent high-frequency features, leading to blurriness or distortion when reconstructing fine details and edges. To address these limitations, we introduce the channel–spatial coherent transformer (CSCT). The core of CSCT includes the channel–spatial coherent attention (CSCA) and the frequency-gated feed-forward network (FGFN), which work synergistically to enhance edge and detail preservation while significantly improving overall image clarity. CSCA efficiently aggregates channel and spatial information, while FGFN adaptively adjusts frequency information to enhance high-frequency details and suppress low-frequency noise. Moreover, this article leverages advanced data augmentation methods that markedly boost RSISR performance, offering new avenues for further exploration. The empirical analysis across several remote sensing SR benchmark datasets reveals that our approach excels in detail restoration, effectively reduces artifacts and noise, and significantly enhances the quality of SR images. Kexin Zhang 0003, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Wenping Ma 0001, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Anomaly-Led Prompting Learning Caption Generating Model and BenchmarkabstractVideo anomaly detection (VAD) is an important intelligent system application, but most current research views it as a coarse binary classification task that lacks a fine-grained understanding of abnormal video sequences. We explore a new task for video anomaly analysis called Comprehensive Video Anomaly Caption (CVAC), which aims to generate comprehensive textual captions (containing scene information such as time, location, anomalous subject, anomalous behavior, etc.) for surveillance videos. CVAC is more consistent with human understanding than VAD, but it has not been well explored. We constructed a large-scale benchmark CVACBench to lead this research. For each video clip, we provide 6 fine-grained annotations, including scene information and abnormal keywords. A new evaluation metric Abnormal-F1 (A-F1) is also proposed to more accurately evaluate the caption generation performance of the model. We also designed a method called Anomaly-Led Generating Prompting Transformer (AGPFormer) as a baseline. In AGPFormer, we introduce an anomaly-led language modeling mechanism (Anomaly-Led MLM, AMLM) to focus on anomalous events in videos. To achieve more efficient cross-modal semantic understanding, we design the Interactive Generating Prompting (IGP) module and Scene Alignment Prompting (SAP) module to explore the divide between video and text modalities from multiple perspectives, and to improve the model's performance in understanding and reasoning about the complex semantics of videos. We conducted experiments on CVACBench by using traditional caption metrics and the proposed metrics, and the experimental results demonstrate the effectiveness of AGPFormer in the field of anomaly caption. Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Baoliang Chen |
IEEE Trans. Multim. | 7 |
| 2025 | Uncertainty Guided Progressive Few-Shot Learning Perception for Aerial View SynthesisabstractView synthesis of aerial scenes has gained attention in the recent development of applications such as urban planning, navigation, and disaster assessment. This development is closely connected to the recent advancement of the Neural Radiance Field (NeRF). However, when autonomousaerial vehicles(AAVs) encounter constraints such as limited perspectives or energy limitations, NeRF degrades with sparsely sampled views in complex aerial scenes. On this basis, we aim to solve this problem in a few-shot manner. In this paper, we propose Uncertainty Guided Perception NeRF (UPNeRF), an uncertainty-guided perceptual learning framework that focuses on applying and improving NeRF in few-shot aerial view synthesis (FSAVS). First, simply optimizing NeRF in complex aerial scenes with sparse input can lead to overfitting in training views, resulting in a collapsed model. To address this, we propose a progressive learning strategy that utilizes the uncertainty present in sparsely sampled views, enabling a gradual transition from easy to hard learning. Second, to take advantage of the inherent inductive bias in the data, we introduce an uncertainty-aware discriminator. This discriminator leverages convolutional capabilities to capture intricate patterns in the rendered patches associated with uncertainty. Third, direct optimization of NeRF lacks prior knowledge of the scene. This, coupled with a reduction in training views, can result in unrealistic rendering. To overcome this, we present a perceptual regularizer that incorporates prior knowledge through prompt tuning of a self-supervised pre-trained vision transformer. In addition, we adopt a sampled scene annealing strategy to enhance training stability. Finally, we conducted experiments with two public datasets, and the positive results indicate our method is effective. Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | LGSNet: Local-Global Semantics Learning Object DetectionabstractSelf-attention learns capturing the long-range dependencies between embeddings (e.g., image pixels). However, the memory overhead and computation cost are prohibitive due to being quadratic in term of the spatial resolution. The structure analysis reveals two crucial roles in the attention: the correlation-based dependency structure and feature normalization. In this work, an efficacious Local-Global Semantics (LGS) module is proposed to alleviate the above issues by modeling the local semantic aggregation and global semantic interaction. Our LGS module contains a group convolution and an Efficient Global Semantic Attention (EGSA). Firstly, the group convolution aggregates local semantics. Secondly, considering a feature map as a sequence of 2-D channel representations, EGSA formulates a general model for the global semantic interaction. The linear correlation is computed between global semantics. LGS has the linear memory overhead and computation cost in term of the spatial resolution. The LGS module can be smoothly incorporated into object detection frameworks. The experiment results verify its effectiveness on two popular detection datasets: the MS COCO and PASCAL VOC. Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Puhua Chen |
IEEE Trans. Multim. | 3 |
| 2025 | Semantic-Aware Wavelet Transformer for Pyramid Learning Object DetectionabstractTransformer displays the impressive capabilities on vision tasks. The built-in self-attention retains the quadratic computation burden in respect of the spatial resolution of image features. The traditional downsampling (e.g., average pooling) can reduce the resolution. Nonetheless, it may suffer from the dropping of detailed information. In this work, we propose an Efficient Wavelet Attention (EWA), which injects the wavelet transform and a Mean GELU (MGELU) function. Firstly, the wavelet transform enables the detailed information to participate in the efficient interaction modeling. Secondly, MGELU regards the statistical mean as reference and loosely passes the high relative responses. Building upon EWA, we present an effective Semantic-aware Wavelet Transformer (SWFormer), which is then employed for pyramid learning, including CNN feature hierarchy or Region of Interest (RoI) features. For the feature hierarchy, a Pyramid SWFormer (PSWFormer) incorporates SWFormer at each level to fit the bidirectional features. For RoIs, a Recognition-Localization SWFormer (RLSWFormer) is inserted into the head to fit their features from all levels. The effectiveness of our SWFormer is displayed experimentally on the MS COCO detection dataset and the Pascal VOC dataset. When exploiting Swin-small backbone, our SWFormer-based method acquires AP of 52.1 in the single-scale evaluation on the COCO test-dev set. This work will have the codes athttps://github.com/TimeIsFuture/Dt2_SWFormer. Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Puhua Chen |
IEEE Trans. Multim. | 3 |
| 2025 | Adaptive Complex Wavelet Informed Transformer OperatorabstractVisual transformers have achieved great success in representation learning. This is mainly due to efficient token dependency modeling via self-attention. However, the computational burden increases sharply as the input pixels increase. Although recent Fourier-based global frequency-domain mixing methods attempt to improve the efficiency of transformers for high-resolution image inputs, the Fourier operator has limited ability to capture the local geometric structure. Complex wavelets can perform local attention in both the spatial domain and the frequency domain. Therefore, we propose the complex wavelet informed transformer operator that uses the real and imaginary wavelets of the dual-tree complex wavelet transform to simulate the interaction in the attention kernel. In order to further reduce the computational burden of operators, we introduce an adaptive local block shared attention mechanism in the channel domain for our wavelet informed operators. Further, we construct the deep multi-head operator network consisting of a hybrid stack of complex wavelet informed transformer operators and self-attention layers. This enables the Transformer to more sparsely capture multi-scale and multi-directional structured features in the process of learning dependencies. Extensive experimental results show that our adaptive complex wavelet informed transformer operator under the Transformer architecture achieves highly competitive accuracy performance on multiple image classification benchmark datasets. And the proposed operators can be flexibly and effectively migrated to vision tasks in dynamic video scenarios. Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Hao Zhu 0009, Xu Liu 0006, Lingling Li 0002, Wenping Ma 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Uncertainty-Aware Semi-Supervised Learning Segmentation for Remote Sensing ImagesabstractDeep learning based remote sensing (RS) image segmentation significantly impacts several real application scenarios. Behind its success, massive labeled data plays an important role. However, annotating high-resolution RS images requires time-consuming and relevant expertise efforts. To address it, many works dive into semi-supervised learning which utilizes raw information embedded in unlabeled data to improve the segmentation model. Nevertheless, previous studies ignore the integrity and effectiveness of the potential context information hidden in RS data. In this work, we propose an uncertainty-aware masked consistency learning (U-MCL) framework that contains an uncertainty-aware masked denoising (U-MD) module and an uncertainty-aware masked image consistency (U-MIC) module. U-MCL initially generates a patch-wise uncertainty map for each unlabeled image during each training iteration, which is then used to derive an adaptive mask ratio for pseudo-label denoising in U-MD. Simultaneously, the uncertainty map is adopted to model a masked unlabeled image for reasoning unseen areas in U-MIC. Consequently, U-MCL is capable of enhancing model performance by engaging in accurate and stable consistency learning while preserving the integrity of the context and employing the context to infer the predictions of the masked regions safely. Extensive experiments on six RS datasets, i.e., ISPRS Vaihingen, FloodNet, MiniFrance, LoveDA, MER, and MSL, demonstrate the superiority of our U-MCL over recent most advanced methods, achieving new state-of-the-art performance under all benchmarks. Xiaoqiang Lu, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | A 3D Self-Awareness Diffusion Network for Multimodal ClassificationabstractAs imaging sensor technology in remote sensing has advanced quickly, multimodal fusion classification has become an important research direction in land cover and urban planning classification tasks. While generative models and image classification have greatly benefited from diffusion models, the present ones primarily concentrate on single-modality-driven diffusion processes. Therefore, this paper presents a 3D self-awareness diffusion network (3DSA-DiffNet) for multispectral (MS) and panchromatic (PAN) image fusion classification, which would make it easier to classify heterogeneous data from various sensors. First, in order to model the relationship between multi-channel spectra and multi-pixel spatial distributions as well as samples, respectively, a spatial-spectral joint denoising network (S$^{2}$JD-Net) is proposed. It can incorporate the diffusion process into the neural network to enhance the quality of diffusion features. Secondly, to imitate the brain's spatial-spectral coexistence learning mechanism, this work offers a 3D self-awareness module (3DSA-Module) that can learn the weight of each pixel in 3D space, resulting in extraordinarily high feature representation capabilities. Finally, experimental verification demonstrates that the 3D self-awareness diffusion fusion network driven by brain inspiration outperforms more sophisticated approaches on the Xi'an, Huhhot, and Muufl datasets. Mengru Ma, Wenping Ma 0001, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Yuwei Guo 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Tracking Like Human: Dynamic Scene Learning Reasoning Tracker in Satellite VideosabstractIn satellite video object tracking, the individual frame analysis method is usually used for target localization, ignoring informative cues of the dynamic scene. Temporal information could contribute to identifying the target from distractors. In this work, a novel dynamic scene learning reasoning tracker is proposed for satellite videos, which reasons over temporal dynamic information to derive the target location. It is inspired by the tracking pattern through human perception and reasoning. First, static-dynamic united analysis is designed to construct dynamic scenes by concatenating the static searching results along the temporal dimension. Second, the information of each response object is aggregated by wavelet transforms. Meanwhile, these scenes are projected into low-frequency and high-frequency subspaces, which could imitate different levels of perceptions of humans for scenes. Third, an object-aware reasoning transformer is proposed to utilize the temporal dynamics of input response objects. In each subspace, it models the mutual interactions between dynamic objects and further learns the intrinsic property of each object for target reasoning. Finally, to obtain the current reasoning result, inverse wavelet transforms are utilized to integrate the results of low-frequency and high-frequency subspaces. The effectiveness of the proposed method is validated on three public satellite video datasets, including SV248S, SkySat, and VISO. Qualitative and quantitative experimental results show that the proposed tracker outperforms 22 popular approaches in seven challenging tracking satellite scenarios. Licheng Jiao, Yangyang Li 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Heterogeneous Riemannian Few-Shot Learning NetworkabstractHow to learn and accurately distinguish new concepts from few samples, as humans do, is a long-standing concern in artificial intelligence (AI). Studies in brain science and neuroscience have shown that human brain perception is based on nonlinear manifolds, and high-dimensional manifolds can facilitate concept learning in neural circuits. Based on this inspiration, in this paper, we propose a heterogeneous Riemannian few-shot learning network (HRFL-Net), which is the first few-shot learning method to perform end-to-end deep learning on heterogeneous Riemannian manifolds. Specifically, to enhance the geometric invariance of the image representation, the image features are projected into three heterogeneous Riemannian manifold spaces. Then, the implicit Riemannian kernel function maps the manifolds to the separable high-dimensional reproducing Hilbert space. It is assumed that the embedded kernel features of the complementary manifolds are mapped to the same common subspace. Thus, a novel neural network-based Riemannian metric learning method is designed to solve the subspace feature vectors by imposing orthogonal normalized projection, which overcomes the data extension limitation of the Riemannian metric. Finally, with the optimization objective of increasing the interclass distance and decreasing the intraclass distance in Hilbert space, the HRFL-Net is trained with end-to-end stochastic optimization, and the optimal aggregation subspace is learned during the gradient descent process. Thus, the proposed HRFL-Net can be easily generalized to challenging nonconvex data. The evaluation of four public datasets shows that the proposed HRFL-Net has significant superiority and also achieves competitive results compared with the state-of-the-art methods. Jie Chen 0098, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Yuwei Guo 0001, Puhua Chen, Wenping Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | A Spatial-Spectral Relation-Guided Fusion Network for Multisource Optical RS Image ClassificationabstractMultisource optical remote sensing (RS) image classification has obtained extensive research interest with demonstrated superiority. Existing approaches mainly improve classification performance by exploiting complementary information from multisource data. However, these approaches are insufficient in effectively extracting data features and utilizing correlations of multisource optical RS images. For this purpose, this article proposes a generalized spatial-spectral relation-guided fusion network (S2RGF-Net) for multisource optical RS image classification. First, we elaborate on spatial- and spectral-domain-specific feature encoders based on data characteristics to explore the rich feature information of optical RS data deeply. Subsequently, two relation-guided fusion strategies are proposed at the dual-level (intradomain and interdomain) to integrate multisource image information effectively. In the intradomain feature fusion, an adaptive de-redundancy fusion module (ADRF) is introduced to eliminate redundancy so that the spatial and spectral features are complete and compact, respectively. In interdomain feature fusion, we construct a spatial-spectral joint attention module (SSJA) based on interdomain relationships to sufficiently enhance the complementary features, so as to facilitate later fusion. Experiments on various multisource optical RS datasets demonstrate that S2RGF-Net outperforms other state-of-the-art (SOTA) methods. Xueli Geng, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Brain-Inspired Learning, Perception, and Cognition: A Comprehensive ReviewabstractThe progress of brain cognition and learning mechanisms has provided new inspiration for the next generation of artificial intelligence (AI) and provided the biological basis for the establishment of new models and methods. Brain science can effectively improve the intelligence of existing models and systems. Compared with other reviews, this article provides a comprehensive review of brain-inspired deep learning algorithms for learning, perception, and cognition from microscopic, mesoscopic, macroscopic, and super-macroscopic perspectives. First, this article introduces the brain cognition mechanism. Then, it summarizes the existing studies on brain-inspired learning and modeling from the perspectives of neural structure, cognitive module, learning mechanism, and behavioral characteristics. Next, this article introduces the potential learning directions of brain-inspired learning from four aspects: perception, cognition, understanding, and decision-making. Finally, the top-ten open problems that brain-inspired learning, perception, and cognition currently face are summarized, and the next generation of AI technology has been prospected. This work intends to provide a quick overview of the research on brain-inspired AI algorithms and to motivate future research by illuminating the latest developments in brain science. Licheng Jiao, Mengru Ma, Pei He, Xueli Geng, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001, Biao Hou, Xu Tang 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Multiscale Deep Learning for Detection and Recognition: A Comprehensive SurveyabstractRecently, the multiscale problem in computer vision has gradually attracted people's attention. This article focuses on multiscale representation for object detection and recognition, comprehensively introduces the development of multiscale deep learning, and constructs an easy-to-understand, but powerful knowledge structure. First, we give the definition of scale, explain the multiscale mechanism of human vision, and then lead to the multiscale problem discussed in computer vision. Second, advanced multiscale representation methods are introduced, including pyramid representation, scale-space representation, and multiscale geometric representation. Third, the theory of multiscale deep learning is presented, which mainly discusses the multiscale modeling in convolutional neural networks (CNNs) and Vision Transformers (ViTs). Fourth, we compare the performance of multiple multiscale methods on different tasks, illustrating the effectiveness of different multiscale structural designs. Finally, based on the in-depth understanding of the existing methods, we point out several open issues and future directions for multiscale deep learning. Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Zhixi Feng, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Complex Dual-Tree Pyramid Scattering TransformerabstractAttention-based transformer networks have recently played an increasingly important role in computer vision tasks. However, since pixel-by-pixel attention multiplication does not involve constraint assumptions such as spatial invariance, the computational complexity grows quadratically with the increase of input pixels. Therefore, this article proposes a complex pyramid scattering Transformer in dense scale space, which introduces sparse scattering constraints with a small number of wavelet basis parameters. It enhances the Transformer's flexibility and sparsity in multiscale space and, to a certain extent, slows down the increase in computational complexity caused by multiresolution input. In addition, compared with the general single-tree real wavelet transform, the dual-tree complex scattering method improves the aliasing of the scattering attention layer and helps obtain a more robust feature representation. At the same time, the multihead stepwise pyramid scattering coupling mechanism helps increase the abundance of directional priors. We conduct experiments in image classification and video tracking scenarios and verify the reliability and superiority of our dual-tree complex pyramid scattering Transformer for visual tasks with different scale requirements. The performance is better than that of the baseline Transformer and other advanced wavelet scattering networks at the same parameter scale. The code is available at https://github.com/Dawn5786/CPSTFormer. Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Hao Zhu 0009, Xin Zhang 0167, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Chain-of-Situation Aware Progressive Inference LearningabstractThe grounded situation recognition (GSR) task aims to recognize the structured semantics of an image to achieve "human-like" event understanding. Most previous studies primarily focus on the visual features of the situation, overlooking the step-by-step cognitive reasoning process that humans employ in complex task settings. Recently, the emergence of multimodal large language models (MLLMs) has provided novel directions for addressing complex problems. However, directly deploying MLLMs on the GSR task is suboptimal due to their tendency to exhibit "hallucination" issues. Additionally, fine-tuning MLLMs for the GSR task incurs high training costs. To address these challenges, inspired by human cognitive theory and the chain-of-thought (CoT) strategy, we propose the chain-of-situation progressive inference learning (CoS-PIL) framework, a lightweight approach that progressively completes verb prediction, noun prediction, and role grounding. The prediction of each step depends on the historical information of the previous step. Specifically, we first design situation prompts tailored to the GSR task and utilize MLLMs to analyze the input image and language prompts, generating heuristic response text for the current situation in the image. Instead of fine-tuning the MLLM, we activate the reasoning capabilities of the frozen MLLM and adapt its generated responses into three lightweight modules: CoS-Verb, CoS-Noun, and CoS-Ground. Considering that MLLMs may generate redundant content, we carefully design the chain-of-interest predictor (CoI-Predictor) to extract key information from the extensive response text and inject it into the model as prompts to enhance the performance. Extensive experiments on the challenging SWiG benchmark demonstrate that CoS-PIL outperforms other state-of-the-art methods. The code is publically available at https://github.com/XDLiuyyy/CoS-PIL. Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | CoT: Contourlet Transformer for Hierarchical Semantic SegmentationabstractThe Transformer-convolutional neural network (CNN) hybrid learning approach is gaining traction for balancing deep and shallow image features for hierarchical semantic segmentation. However, they are still confronted with a contradiction between comprehensive semantic understanding and meticulous detail extraction. To solve this problem, this article proposes a novel Transformer-CNN hybrid hierarchical network, dubbed contourlet transformer (CoT). In the CoT framework, the semantic representation process of the Transformer is unavoidably peppered with sparsely distributed points that, while not desired, demand finer detail. Therefore, we design a deep detail representation (DDR) structure to investigate their fine-grained features. First, through contourlet transform (CT), we distill the high-frequency directional components from the raw image, yielding localized features that accommodate the inductive bias of CNN. Second, a CNN deep sparse learning (DSL) module takes them as input to represent the underlying detailed features. This memory- and energy-efficient learning method can keep the same sparse pattern between input and output. Finally, the decoder hierarchically fuses the detailed features with the semantic features via an image reconstruction-like fashion. Experiments demonstrate that CoT achieves competitive performance on three benchmark datasets: PASCAL Context [57.21% mean intersection over union (mIoU)], ADE20K (54.16% mIoU), and Cityscapes (84.23% mIoU). Furthermore, we conducted robustness studies to validate its resistance against various sorts of corruption. Our code is available at: https://github.com/yilinshao/CoT-Contourlet-Transformer. Yilin Shao, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationabstractPre-trained vision-language(V-L) models such as CLIP have demonstrated impressive Zero-Shot performance in many downstream tasks. Since adopting contrastive video-text pairs methods like CLIP to video tasks is limited by its high cost and scale, recent approaches focus on efficiently transferring the image-based CLIP to the video domain. A major finding is that fine-tuning the pre-trained model to achieve strong fully supervised performance leads to low zero shot, few shot, and base to novel generalization. Instead, freezing the backbone network to maintain generalization ability weakens fully supervised performance. Otherwise, no single prompt tuning branch consistently performs optimally. In this work, we proposed a multimodal prompt learning scheme that balances supervised and generalized performance. Our prompting approach contains three sections: 1) Independent prompt on both the vision and text branches to learn the language and visual contexts. 2) Inter-modal prompt mapping to ensure mutual synergy. 3) Reducing the discrepancy between the hand-crafted prompt (a video of a person doing [CLS]) and the learnable prompt, to alleviate the forgetting about essential video scenarios. Extensive validation of fully supervised, zero-shot, few-shot, base-to-novel generalization settings for video recognition indicates that the proposed approach achieves competitive performance with less commute cost. Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Zehua Hao, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
AAAI | 9 |
| 2024 | Multiplane Prior Guided Few-Shot Aerial Scene RenderingabstractNeural Radiance Fields (NeRF) have been successfully applied in various aerial scenes, yet they face challenges with sparse views due to limited supervision. The acquisition of dense aerial views is often prohibitive, as unmanned aerial vehicles (UAVs) may encounter constraints in perspective range and energy constraints. In this work, we introduce Multiplane Prior guided NeRF (MPNeRF), a novel approach tailored for few-shot aerial scene rendering-marking a pioneering effort in this domain. Our key insight is that the intrinsic geometric regularities specific to aerial imagery could be leveraged to enhance NeRF in sparse aerial scenes. By investigating NeRF's and Multiplane Image (MPI)'s behavior, we propose to guide the training process of NeRF with a Multiplane Prior. The proposed Multiplane Prior draws upon MPI's benefits and incorporates advanced image comprehension through a Swin V2 Transformer, pre-trained via SimMIM. Our extensive experiments demonstrate that MPN-eRF outperforms existing state-of-the-art methods applied in non-aerial contexts, by tripling the performance in SSIM and LPIPS even with three views available. We hope our work offers insights into the development of NeRF-based applications in aerial scenes with limited data. Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Puhua Chen, Yuwei Guo 0001 |
CVPR | 4 |
| 2024 | Multimodal Segformer for Flood Rapid Mapping with Sentinel-2 DataabstractFlood rapid mapping products play an important role in informing flood emergency response and management. To this end, the 2024 IEEE GRSS Data Fusion Contest Track 2 (DFC24-T2) establishes a multimodal benchmark for the segmentation of flood areas from Sentinel-2 multispectral images. However, the problems of imbalanced data distribution, data scarsity, and inter-modal differences severely inhibit the performance of deep-learning-based segmentation networks. In this work, we propose an end-to-end Multimodal Transformer-based Segmentation Network (MTSN) for accurate flood rapid mapping. MTSN first employs two Siamese encoders with shared parameters to accept multimodal inputs and output their respective hierarchical multiscale features, which are then enriched by several channel attention blocks. Subsequently, a Cross-modal Feature Fusion Module (CFFM) based on a gated mechanism is proposed to efficiently integrate the benefits of multimodal features, and generate informative representations. Finally, the fused features are decoded by a lightweight pure multilayer perception decoder to quickly generate mapping results of flood areas. Moreover, we introduce offline data augmentation, semi-supervised learning, test-time augmentation, and multimodal post-process to further boost the performance and generalization of our MTSN. Experimental results and extensive ablations show the effectiveness of our method. Code is available at https://github.com/xiaoqiang-lu/MMSegFormer. Xiaoqiang Lu, Tong Gou, Zhongjian Huang, Yuting Yang 0008, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001 |
IGARSS | 7 |
| 2024 | Multi-Scale Co-Attention Learning For SAR Image Change DetectionabstractSynthetic aperture radar (SAR)iamge change detection is an important and challenging task. Existing methods mainly extract features from the difference map and original images in a parallel structure, ignoring the correlations between features. In this paper, we propose a novel SAR image change detection method based onmulti-scale and co-attention learning, called the Multi-scale Co-attention Network (MCNet). Specifically, we extract features from the difference map and original images using co-attention mechanism, which establishes a relationship between the two kinds of features in the feature space while retaining key information from the original images. With this attention mechanism, the model pays more attention to the changed areas in SAR images, and the features contain richer information. Furthermore, we propose a multi-scale feature fusion method to combine high-level and low-level semantic information, improving the generality and robustness of the features. Finally, the performance of the proposed method has been validated for its effectiveness. Xu Liu 0006, Licheng Jiao, Fang Liu 0001 |
IGARSS | 2 |
| 2024 | Domain Generalization-Aware Uncertainty Introspective Learning for 3D Point Clouds Segmentation
Pei He, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0002, Shuyuan Yang 0001, Ronghua Shang |
ACM Multimedia | 4 |
| 2024 | Geometric Prior Guided Feature Representation Learning for Long-Tailed Classification
Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006, Puhua Chen |
Int. J. Comput. Vis. | 5 |
| 2024 | Saliency and Depth-Aware Full Reference 360-Degree Image Quality AssessmentabstractWith the widespread adoption of virtual reality and 360-degree video, there is a pressing need for objective metrics to assess quality in this immersive panoramic format reliably. However, existing image quality assessment models developed for traditional fixed-viewpoint content do not fully consider the specific perceptual issues involved in 360-degree viewing. This paper proposes a 360-degree image full-reference quality assessment (FR-IQA) methodology based on a multi-channel architecture. The proposed 360-degree FR-IQA method further optimizes and identifies the distorted image quality using two easily obtained useful saliency and depth-aware image features. The convolutional neural network (CNN) is designed for training. Furthermore, the proposed method accounts for predicting user viewing behaviors within 360-degree images, which will further benefit the multi-channel CNN architecture and enable the weighted average pooling of the predicted FR-IQA scores. The performance is evaluated on publicly available databases to demonstrate the advantages brought by the proposed multi-channel model in performance evaluation and cross-database evaluation experiments, where it outperforms other state-of-the-art ones. Moreover, an ablation study exhibits good generalization ability and robustness. Xuekai Wei, Qunyue Huang, Bin Fang 0001, Lei Ouyang, Weizhi Xian, Jun Luo 0003, Huayan Pu, Xueyong Xu, Chang Lu 0005, Hao Nan, Xu Liu 0006, Yachao Li 0001, Mingliang Zhou 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 11 |
| 2024 | Transformer with a Parallel Decoder for Image CaptioningabstractIn this paper, a parallel decoder and a word group prediction module are proposed to speed up decoding and improve the effect of captions. The features of the image extracted by the encoder are linearly projected to different word groups, and then a unique relaxed mask matrix is designed to improve the decoding speed and the caption effect. First, since image captioning is composed of many words, sentences can also be broken down into word groups or words according to their syntactic structure, and we achieve this function through constituency parsing. Second, we make full use of the extracted features to predict the size of word groups. Then, a new embedding representing the information of the word is proposed based on word embedding. Finally, with the help of word groups, we design a mask matrix to modify the decoding process so that each step of the model can produce one or more words in parallel. Experiments on public datasets demonstrate that our method can reduce the time complexity while maintaining competitive performance. Peilang Wei, Xu Liu 0006, Jun Luo 0006, Huayan Pu, Xiaoxu Huang, Shilong Wang 0001, Huajun Cao, Shouhong Yang, Xu Zhuang, Hong Yue, Cheng Ji 0002, Mingliang Zhou 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2024 | Global Shapes and Salient Joints Features Learning for Skeleton-Based Action RecognitionabstractGlobal shapes and local joints are significant cues to learn skeleton representations for human action recognition. However, most shapes-based methods ignore local joints information and thus perform less well when recognizing similar actions. In this letter, we propose a two-stream method of global shapes and salient joints features learning to address the above issues. Firstly, the global sparse shapes modeling stream (GSSM) explores the global skeletal shapes in Kendall space and codes the nonlinear shapes by a sparse coding and dictionary learning method. Then, the salient enhancing joints generation stream (SEJG) explores local joints information by extracting discriminative salient joints and generates frame-enhancing action sequences with invariant length by a bilinear frame interpolation module. To improve the temporal modeling ability of our method, we use similar multi-scale LSTM methods in both GSSM and SEJG, which explore global-local features in static, short-term, and long-term temporal scales. Extensive experiments on four challenging datasets verify the effectiveness of our proposed method, which achieves competitive results compared to state-of-the-art (SOTA) methods. Xingtong He, Xu Liu 0006, Licheng Jiao |
IEEE Signal Process. Lett. | 2 |
| 2024 | Satellite Video Object Tracking Based on Location PromptsabstractObject Tracking in satellite videos is a challenging task due to the small target size, low spatial resolution, limited appearance and texture information, and the potential for background confusion. While current state-of-the-art tracking methods perform well on natural images, they often produce unsatisfactory results when applied to satellite videos. In this paper, we address these challenges by leveraging location prompts and refining the feature extractor and bounding box refinement module. Furthermore, we integrate motion features to effectively handle illumination variations that frequently arise in satellite videos, thereby enhancing the overall robustness of the tracker. Our proposed approach, abbreviated as SVLPNet, has been thoroughly evaluated through extensive experiments conducted on two authentic satellite video datasets. The obtained results unequivocally showcase the promising potential of SVLPNet in facilitating object tracking on satellite videos. The source code and raw results will be released at https://github.com/Wprofessor/SVLPNet. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Yingjia Gao, Hao Wang 0211, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Shuo Li 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2024 | A Graph Association Motion-Aware Tracker for Tiny Object in Satellite VideosabstractSatellite video object tracking involves tracking a specified tiny object within a wide scene. The insufficient appearance features of these tiny objects pose significant challenges to appearance-based object trackers, particularly in situations involving occlusion, target blur, and similar interferences. In this paper, a novel Graph Association MOtion-aware tracker (GAMO) is proposed for tiny object in satellite videos, which integrates motion and spatial relationship information. First, a Gaussian motion estimator is proposed that decouples motion into velocity and direction, rather than using traditional x-y movement modeling. This estimator predicts the object’s position and estimates motion uncertainty with a directional motion probability map. Furthermore, the estimated motion serves as a prior to guide the proposal sampling. A probabilistic proposal sampling module is designed that samples candidate bounding boxes according to the directional motion probability map, focusing on the region where the target is most likely to appear. Additionally, we implement a graph association module to model and propagate the spatial relationships between the target and neighboring objects over time. This relationship information assists the appearance features in distinguishing the target from similar interferences. Experiments on the Skysat-1, SV248S, and VISO datasets demonstrate the superiority of the proposed tracker. GAMO leverages motion and surrounding information, resulting in significant improvements with minimal computational overhead. The code and results will be publicly available inhttps://github.com/Midkey/GAMO. Zhongjian Huang, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Xiangrong Zhang, Lingling Li 0002, Puhua Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Multi-Grained Gradual Inference Model for Multimedia Event ExtractionabstractWith the development of multimedia technology, events are usually presented in multimedia forms, thus multimedia event extraction (MEE) has become more and more important. Existing MEE works usually use simple strategies to align two modalities, making it difficult to precisely extract events and arguments in complex multimedia documents. To address this problem, we propose a novel Multi-grained Gradual Inference Model (MGIM) that focuses on inferring and interpreting events in complex multimedia structures in a coarse-to-fine manner. To efficiently integrate textual and visual modalities, we design a Coarse-grained Alignment (CA) module, which represents the two modalities in a graph structure and performs coarse-grained alignment. Based on the CA module, we further propose a Fine-grained Inference module (FI) that fine-grained aligns text and image by performing multiple rounds of gradual inference. MGIM provides a comprehensive interpretation of multimedia events at two information granularities (coarse and fine). Extensive experiments on the M2E2 dataset demonstrate the effectiveness of MGIM. Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2024 | Event-Based Monocular Depth Estimation With Recurrent TransformersabstractEvent cameras, offering high temporal resolutions and high dynamic ranges, have brought a new perspective to address common challenges in monocular depth estimation (e.g., motion blur and low light). However, existing CNN-based methods insufficiently exploit global spatial information from asynchronous events, while RNN-based methods show a limited capacity for effective temporal cues utilization for event-based monocular depth estimation. To this end, we propose a event-based monocular depth estimator with recurrent transformers, namely EReFormer. Technically, we first design a transformer-based encoder-decoder that utilizes multi-scale features to model global spatial information from events. Then, we propose a Gate Recurrent Vision Transformer (GRViT), introducing a recursive mechanism into transformers, to leverage rich temporal cues from events. Finally, we present a Cross Attention-guided Skip Connection (CASC), performing cross attention to fuse multi-scale features, to improve global spatial modeling capabilities. The experimental results show that our EReFormer outperforms state-of-the-art methods by a margin on both synthetic and real-world datasets. Our open-source code is available at https://github.com/liuxu0303/EReFormer. Xu Liu 0006, Jianing Li 0001, Jinqiao Shi, Xiaopeng Fan 0001, Yonghong Tian 0001, Debin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Self Pseudo Entropy Knowledge Distillation for Semi-Supervised Semantic SegmentationabstractRecently, semi-supervised semantic segmentation methods based on weak-to-strong consistency learning have achieved the most advanced performance. The key to such a technique lies in strong perturbations and multi-objective co-training. However, CutMix, the most commonly used data augmentation in this field, limits the strength of perturbations as it only focuses on single random local context. Besides, complex optimization targets also reduce computational efficiency. In this work, we propose an efficient consistency learning based framework. Specifically, a novel unsupervised data augmentation strategy, EntropyMix, is present for semi-supervised semantic segmentation. Patches of unlabeled data from multi-view augmentations are combined into new training samples based on their prediction entropy, which provides more informative and powerful perturbations for consistency regularization and impels the model to focus on cross-view local context. On this basis, we further propose Self Pseudo Entropy Knowledge Distillation (SPEED) to learn global pixel relations from multi- and cross-view perturbations by optimizing a linear combination of feature-and logit-level distillation loss, enhancing model performance without additional auxiliary segmentation heads or a complex pre-trained teacher model. The collocation of the two ideas above is a plug-and-play technique without additional modification. Extensive experimental results on PASCAL VOC and Cityscapes datasets under various training settings demonstrate the superiority of the proposed data augmentation strategy and self-distillation loss, achieving new state-of-the-art performance. Remarkably, our method reaches mIoU of 75.16% using only 0.87% labeled data on PASCAL VOC and mIoU of 76.98% using only 6.25% labeled data on Cityscapes. The code is available at https://github.com/xiaoqiang-lu/SPEED. Xiaoqiang Lu, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | MBSI-Net: Multimodal Balanced Self-Learning Interaction Network for Image ClassificationabstractA growing number of earth observation satellites are able to simultaneously gather multimodal images of the same area due to the expanding availability and resolution of satellite remote sensing data. This paper proposes a novel multimodal balanced self-learning interaction network (MBSI-Net) for the classification task. It involves a dual-branch teacher-student network that enables knowledge interaction and transfer between the multimodalities. Firstly, in order to introduce statistical information in addition to local and global structural information, a texture feature equalization module (TFE-Module) is proposed. This can enhance the texture information of features through histogram equalization and further improve the representation ability of features. Secondly, to enable the student network to provide timely feedback questions, the paper proposes a feature fusion module (F2-Module) that models and enhances teacher features through the student network. This helps to raise the classification’s accuracy by incorporating information from multimodal images. Finally, the paper proposes a loss function based on structural similarity analysis to ensure balanced self-learning between the student and the teacher networks. Taking the multispectral (MS) and the panchromatic (PAN) images of the same scene as examples, through experimental verification, the proposed method can achieve good results on multiple datasets compared with other methods. Therefore, it offers an effective method for classifying and fusing multimodal data. Mengru Ma, Wenping Ma 0001, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Shuyuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Visual and Language Collaborative Learning for RGBT Object TrackingabstractDespite the extensive research on RGBT object tracking, there are still several challenges and issues in practical applications, such as modality differences, lighting variations and disappearance of the target, and changes in viewpoint. Existing methods mostly address these issues by fusing image features, while neglecting a significant amount of target label information. To address these challenges, this paper introduces text to drive the alignment of visible and infrared image features, transforming features from different modalities into the same feature space and fully using complementary features between different modalities. Furthermore, inspired by the success of prompt learning in various tasks, we utilize prior boxes and language as prompts to further guide the model in tracking the target. Extensive experiments demonstrate that the proposed VLCTrack tracker has excellent potential in RGBT object tracking. Compared to previous methods developed for this purpose, our approach achieves state-of-the-art performance on three benchmark datasets. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Yingjia Gao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2024 | Domain Adaptation-Aware Transformer for Hyperspectral Object TrackingabstractVisual object tracking in natural scenes is a popular but challenging task, owing to the difficulties of feature representation from various changes of the targets, such as size change, deformation, illumination change, rotations, motion blur, background clutter, etc. High-speed hyperspectral imaging systems capture hyperspectral videos (HSVs) in wide spectral ranges and provide abundant spectral and spatial information to tell targets apart from backgrounds, alleviating the model drift in appearance-based tracking methods. However, different hyperspectral imagers, such as near-infrared (NIR), red-to-near-infrared (RedNIR), and visible (VIS), obtain heterogeneous types of data that could not be handled by common object trackers. In this paper, a domain adaptive Transformer framework is proposed for hyperspectral object tracking. Considering the HSVs are from different types of sensors, their heterogeneous features are learned in an adversarial way by domain label reverse learning with a gradient reversed layer. To fully utilize the spectral information in HSV frames, a band-wise spatial attention module (BSAM) is designed to emphasize the salient area near the target of interest. We adopt a Siamese-like Transformer tracker as the main structure for tracking. Our tracker outperforms top-ranking methods on a hyperspectral object tracking benchmark dataset containing three types, 87 hyperspectral videos in total. The comparison experiments validate the effectiveness of the proposed method. The source code and trained models of this work will be publicly available soon at https://github.com/LianYi233/Trans-DAT. Yinan Wu 0001, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Lingling Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Efficient LWPooling: Rethinking the Wavelet Pooling for Scene ParsingabstractExisting wavelet pooling methods discard the high-frequency sub-bands, which can improve the noise-robustness of convolutional neural networks (CNNs) but lose the essential detailed features. Besides, most of them depend on different wavelets, which is not adaptive. In this paper, a novel efficient lifting-based wavelet pooling (LWPooling) is proposed to alleviate the problems above. Firstly, wavelet pooling is rethought based on the equivalence of 2D discrete wavelet transform (DWT) and standard average pooling (SAP), which suggests the lack of detailed information on traditional wavelet pooling. Secondly, the efficient LWPooling module is proposed to adaptively capture and preserve the critical high-frequency features via lifting-based wavelets. It can constrain the features linear independence, which efficiently makes important features salient. Thirdly, the lifting-based wavelet collaborative network (LWCNet) is constructed for classification and segmentation tasks based on the efficient LWPooling module. Experiments are validated on Cifar10, Cifar100, and ADE20K datasets. It suggests that the efficient LWPooling can enhance CNN’s representation and achieve a particular performance advantage compared to average, maximum, and original wavelet pooling. Besides, the proposed LWCNet shows the potential for scene parsing. The code implementation will be available at https://github.com/yutinyang/LWCNet. Yuting Yang 0008, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Xiangrong Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Knowledge Guided Evolutionary Transformer for Remote Sensing Scene ClassificationabstractSolving the complex challenges of sophisticated terrain and multi-scale targets in remote sensing (RS) images requires a synergistic combination of Transformer and convolutional neural network (CNN). However, crafting effective CNN architectures remains a major challenge. To address these difficulties, this study introduces the knowledge guided evolutionary Transformer for RS scene classification (Evo RSFormer). It amalgamates adaptive evolutionary CNN (Evo CNN) with Transformers in a hybrid strategy synergistically, which combines fine-grained local feature extraction of CNNs with long-range contextual dependency modeling of Transformers. Furthermore, for the development of Evo CNN blocks, this paper presents a knowledge-guided adaptive efficient multi-objective evolutionary neural architecture search (MOE2-NAS) strategy. This approach markedly diminishes the labor-intensive characteristics associated with traditional CNN design, striking a balance for both accuracy and compactness. Additionally, by leveraging domain knowledge from natural scene analysis into the RS field, MOE2-NAS facilitates the efficiency of classical NAS. It utilizes a priori knowledge to generate promising initial solutions and constructs a surrogate model for efficient search. The effectiveness of the proposed Evo RSFormer has been rigorously tested on various benchmark RS datasets, including UC Merced, NWPU45, and AID. Empirical results strongly support the superiority of Evo RSFormer over existing methods. Furthermore, experiments on MOE2-NAS have been studied to confirm the important role of knowledge guidance in improving the efficiency of NAS. Jiaxuan Zhao, Licheng Jiao, Chao Wang 0099, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Mengru Ma, Shuyuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Evolutionary Dual-Stream TransformerabstractVision transformers (ViTs) are rapidly evolving and are widely used in computer vision. However, high-performance ViTs require many computations, which limit their further development in the vision field. In this article, a novel evolutionary dual-stream transformer (E-DST) model is proposed to alleviate the computational resource demand problem. A hybrid attention mechanism structure is proposed for a DST model. The DST model uses a dual-branch structure to fuse convolutional and transformer features. Combining the features learned by the transformer and convolution effectively saves model computational resources. In addition, an evolutionary optimizer is proposed to optimize the parameters of the model. The excellent search ability of the evolutionary algorithm is utilized to optimize the transformer model parameters. The convergence of the evolutionary optimizer is proved in this article. In addition, the proposed E-DST model is experimentally compared with a variety of classic models and their deformations based on three datasets. And, the evolutionary optimizer proves its generality in convolutional and recurrent neural networks. The experimental results show that the E-DST model can effectively reduce computational resources and that the evolutionary optimizer can solve large-scale optimization problems. In conclusion, our proposed method is feasible and effective. Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Cybern. | 5 |
| 2024 | Bi-Level Multiobjective Evolutionary Learning: A Case Study on Multitask Graph Neural Topology SearchabstractThe construction of machine learning models involves many bi-level multiobjective optimization problems (BL-MOPs), where upper-level (UL) candidate solutions must be evaluated via training weights of a model in the lower level (LL). Due to the Pareto optimality of subproblems and the complex dependency across UL solutions and LL weights, a UL solution is feasible if and only if the LL weight is Pareto optimal. It is computationally expensive to determine which LL Pareto weight in the LL Pareto weight set is the most appropriate for each UL solution. This article proposes a bi-level multiobjective learning framework (BLMOL), coupling the above decision-making process with the optimization process of the upper-level MOP (UL-MOP) by introducing LL preference$\boldsymbol {r}$. Specifically, the UL variable and$\boldsymbol {r}$are simultaneously searched to minimize multiple UL objectives by evolutionary multiobjective algorithms. The LL weight with respect to$\boldsymbol {r}$is trained to minimize multiple LL objectives via gradient-based preference multiobjective algorithms. In addition, the preference surrogate model is constructed to replace the expensive evaluation process of the UL-MOP. We consider a novel case study on multitask graph neural topology search. It aims to find a set of Pareto topologies and their Pareto weights, representing different tradeoffs across tasks at UL and LL, respectively. The found graph neural network is employed to solve multiple tasks simultaneously, including graph classification, node classification, and link prediction. Experimental results demonstrate that BLMOL can outperform some state-of-the-art algorithms and generate well-representative UL solutions and LL weights. Chao Wang 0099, Licheng Jiao, Jiaxuan Zhao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Evol. Comput. | 5 |
| 2024 | A Quantum Evolutionary Learning Tracker for VideoabstractVideo object tracking has been a popular area in the field of computer vision. As video data evolves, more special perspectives and challenging video data are constantly kept up to date. This poses challenges for object tracking tasks and places higher demands on the generalization capabilities of the models. In this article, we propose a novel quantum evolutionary learning tracker (QELT) for video. The model combines quantum evolution with deep networks for tracking video objects. The model uses a QELT to generate a reliable population of candidate regions and a deep network for classification. In particular, the quantum evolutionary predictor predicts the object motion state through rotation operator and trajectory inference, and provides motion state information for the tracker. The predictor can incorporate object history contextual information and can provide stable candidate estimation populations for the model in case of failure of appearance features. Both quantum evolution and deep networks are combined to form an end-to-end online video object tracker. In addition, we propose a new video object tracking evaluation algorithm, Balanced Intersection over Union. The evaluation algorithm uses aspect ratios to balance the share of overlap and distance. Finally, we test the model on the OTB 2015 dataset for natural video and on the SV248A10-SOT dataset for satellite video. The performance of the proposed model is also analyzed and validated by comparing it with more than 20 classical tracker models. The experimental results show that our model has high generalization ability and robustness. Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Evol. Comput. | 4 |
| 2024 | Mask-Guided Correlation Learning for Few-Shot Segmentation in Remote Sensing ImageryabstractFew-shot segmentation aims to segment specific objects in a query image based on a few densely annotated images and has been extensively studied in recent years. In remote sensing, image segmentation faces challenges such as less training data, large intraclass diversity, and low foreground-background contrast. In this work, we propose a novel few-shot segmentation method in remote sensing imagery based on mask-guided correlation learning (MGCL) to alleviate the above challenges. In our MGCL, a novel mask-guided feature enhancement (MGFE) module is proposed, which makes features have intramask consistency by leveraging oversegmented masks. In order to enhance the contrast between foreground and background, a novel foreground-background correlation (FBC) module is proposed, which enhances background correlation representation by learning foreground correlation and background correlation separately. Furthermore, a novel mask-guided correlation decoder (MGCD) module is proposed to guide the decoder to focus on the consistency within the mask, thereby learning how to segment complete objects and improving segmentation accuracy. Sufficient experiments on the iSAID-$5^{i}$and DLRSD-$5^{i}$datasets show that our MGCL outperforms all comparative methods. In particular, in the one-shot setting of the iSAID-$5^{i}$dataset, we achieve an mIoU of 39.92 based on ResNet50, which is an improvement of 4.25 over the state-of-the-art (SOAT) method. The visualization of features before and after the MGFE module further concretely demonstrates the motivation and advantages of our MGCL. The code is available athttps://github.com/LiShuo1001/MGCL. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Xu Liu 0006, Puhua Chen, Lingling Li 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | MutSimNet: Mutually Reinforcing Similarity Learning for RS Image Change DetectionabstractChange detection involves analysis of discrepancies between two phases. However, when the unchanged elements are known, the changed features to be identified become straightforward. In addition, remote sensing image is constrained by limited spectral information, which leads to blurred boundaries between different semantics. Based on these two prior knowledge, in this artical, we introduce a novel change detection framework, named the mutually reinforcing similarity network (MutSimNet). This architecture aims to minimize false alarms along changing boundaries and reduce misjudgment rates among outliers. First, similarity learning is applied to change detection. The relationship between the two phases is considered when deriving the change feature maps. Second, we devise a mutually reinforcing loss function that integrates initial features with final features. Third, a self-attention module is connected in the feature pyramid network. This design mitigates information loss during the down-sampling process. Fourth, an attention feature fusion strategy is proposed for the integration of multi-layer features. This strategy takes into account the interaction between layer-by-layer features. Fifth, experimental results validate MutSimNet’s efficiency, particularly its ability to focus on edge contour learning. The MutSimNet also achieves superior performance on two benchmark datasets and predicts positive samples with higher probability. The codebase is accessible at https://github.com/ly-yu/MutSimNet. Xu Liu 0006, Yu Liu 0005, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | MGPACNet: A Multiscale Geometric Prior Aware Cross-Modal Network for Images Fusion ClassificationabstractConvolutional neural networks (CNNs) and self-attention (SA) are highly effective techniques used for the fusion of multisource remote sensing (RS) data, and they have found extensive application in Earth observation (EO) tasks. Nevertheless, CNNs are insufficient for the comprehensive extraction of contextual information and the representation of the sequential properties of spectral features. Furthermore, the loss of edge geometry information is often a consequence of information mining, which limits its application in RS. To address the abovementioned limitations, we propose a method called “multiscale geometric prior aware cross-modal network (MGPACNet)” for RS image fusion classification. First, a geometric prior feature enhanced residual module (GPFEResM) is created to extract shallow multiscale geometric edge prior features and detailed information from multimodal RS data to enhance feature boundary information. Second, a multiscale global-local spatial-spectral feature extraction module (MG-LS2FEM) uses multiscale spatial modeling and global-local spectral modeling to perceive rich semantic information in the spatial-spectral domain. Finally, a dual attention fusion module (DAFM) is designed to use pixel-level SA and cross-attention between heterogeneous data to achieve deep aggregation and cross-focusing of cross-modal information in two branches, and enhance the complementarity of heterogeneous data. A comprehensive examination of public RS data (hyperspectral-synthetic aperture radar (HS-SAR) Augsuburg/Berlin, hyperspectral-light detection and ranging (HS-LiDAR) Trento/MUUFL) from four distinct modalities (HS/SAR/LiDAR) has revealed that our method outperforms alternative models. Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Lighter and Robust: A Rotation-Invariant Transformer for VHR Image Change DetectionabstractIn recent years, change detection (CD) has emerged as an increasingly intricate research domain. However, in natural images, the orientation of objects is often aligned with the image boundaries, whereas in RS images, the imaging angles are random. As a result, existing CD methods encounter limitations when effectively representing vector features. In this article, we propose a rotation-invariant CD architecture named RFormer. It effectively utilizes direction-sensitive position embedding (DSPE) to represent features in RS images. To address the challenge of the quadratic growth in attention mechanism complexity with sequence length, we introduce low-cost cross attention (LC2A) to reduce its complexity to$1/{C^{2}}$. Furthermore, we employ the implicit timing extraction process (TEP) to represent interframe bitemporal features. TEP plays a crucial role in mitigating prediction biases caused by seasonal changes in land cover and prevents overconfident discrimination by the classifier in CD tasks. Experimental results demonstrate that RFormer achieves competitive performance on WHU, deeply supervised image fusion network (DSIFN)-CD, CDD, and LEVIR-CD datasets. Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | TrTr-CMR: Cross-Modal Reasoning Dual Transformer for Remote Sensing Image CaptioningabstractRemote sensing image captioning (RSIC) is an interesting but challenging cross-modal reasoning task for computer vision and natural language processing. Most of the recent popular approaches for RSIC utilize encoder-decoder architectures, which focus on visual features captured by convolutional neural network (CNN)-based encoder and semantic information by recurrent neural network (RNN)-based or long short-term memory (LSTM)-based decoder, but encounter difficulties with multiscale, multicategories, and direction ambiguity challenges. To make the most of semantic understanding ability of Transformers, in this article, we propose a new attention-based visual-linguistic reasoning framework with dual Transformer for RSIC. Specifically, Swin Transformer (SwinT) encoder with shifted window partitioning scheme is introduced for multiscale visual feature extraction to discover the intrinsic relationship in the objects, and then, a Transformer language model (TLM) with self-attention and cross attention is designed as the decoder to generate a well-formed sentence for the image. Extensive experiments are conducted on the public RSIC benchmark datasets, including UCM-Captions, Sydney-Captions, and RSICD. The impressive performance verifies the effectiveness and superiority of the proposed method. In addition, the source code and models of this work are publicly available athttps://github.com/LianYi233/TrTr-CMR. Yinan Wu 0001, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | High-Order Relation Learning Transformer for Satellite Video Object TrackingabstractSurrounding contexts are generally perceived as interfering with object tracking in satellite videos, leading to model drift. From another perspective, they can also be seen as reference objects of the tracked target, the dynamic interactions between them could provide essential information. In this article, a high-order relation learning transformer (HRLT) is proposed for satellite video object tracking, which not only models the high-order interactions of different target-context pairs but also reasons the associations between these high-order relations across multiple frames. First, a spatial high-order relation reasoning (SHR2) module is designed to model the high-order interactions between the target and scene contexts. Second, a temporal high-order relation reasoning (THR2) module is proposed to associate and reason these spatial high-order relations across multiple frames. Third, historical high-order relations are collected to provide more reasoning bases for the current frame prediction. Finally, qualitative and quantitative evaluations are performed on the SV248S, SkySat, and VISO datasets. The results show that HRLT outperforms 20 popular methods in different challenging scenarios. Licheng Jiao, Yangyang Li 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | LGLFormer: Local-Global Lifting Transformer for Remote Sensing Scene ParsingabstractIn deep learning, convolutional neural networks (CNNs) and transformers have gained excellent achievements in remote sensing scene parsing. Strong feature representation ability is still a challenge for them. Besides, the complex scenes are still essential challenges for deep learning in remote sensing scene parsing. In this article, an efficient local–global lifting transformer (LGLFormer) framework is proposed to ease the challenges above. It effectively combines CNNs, transformer, and wavelet transform to build a strong local–global (LG) feature representation network. Besides, global feature learning driven by LG adaptive features is proposed based on the 2-D LG adaptive feature extractor (LGAFE) and refined global feature attention module. The 2-D LG lifting feature extractor is inspired by the lifting scheme, which introduces local and global dependency. Furthermore, two LG lifting schemes are proposed, including the series and parallel modes, which can effectively learn LG relations between pixels. Finally, experiments are validated on three remote sensing benchmark datasets. The proposed LGLFormer achieves the state-of-the-art with 99.02%, 99.2%, and 99.48% overall accuracy (OA) on AID, WHU-RS19, and UCM datasets, respectively. In addition, LGLFormer shows good convergence with competitive parameters. The experimental code will be available athttps://github.com/yutinyang/LGLFormer. Yuting Yang 0008, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Relation Learning Reasoning Meets Tiny Object Tracking in Satellite VideosabstractTiny objects in satellite videos are usually not independent individuals, there exist rich semantic and temporal relations with each other. Thus, modeling and reasoning the variation of such intrinsic relationships can be beneficial for tiny object tracking. In this paper, a relation learning reasoning method is proposed for tiny object tracking in satellite videos. The core of the proposed is the relation reasoning network that consists of a key context module, a global semantic module, and a relation reasoning module sequentially. First, the key context module exploits global key contexts which explicitly or implicitly contribute to the target object, modeling the intrinsic relations with the target. Second, to reason the contribution, the global semantic module analyses the interaction between them in the same frame. Third, the relation reasoning module deduces the target based on the variation of the semantic relations among different frames. Such a relation learning reasoning approach which takes the target as the core is aligned with the satellite tiny object tracking task, significantly improves the identification performance in dense similarity scenes and the retrieval ability after completely occluded. Furthermore, the proposed method is shown to report improved qualitative and quantitative results on Jilin-1 and SkySat satellite video datasets. Licheng Jiao, Yangyang Li 0001, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Adaptive Multi-Scale Transformer Tracker for Satellite VideosabstractSatellite video tracking tasks are often characterized by blurred foreground boundaries in vast scenes, a wide range of targets varying in scale, and irregular changes in appearance. These challenges significantly impact the optimization of robust tracker performance. Therefore, it is imperative to extract diverse features with dynamic adaptive learning capabilities for the target being tracked in each sequence. In this article, we explore a novel adaptive multi-scale Transformer (MT) tracker for satellite videos to explore the potential spatiotemporal information of the target effectively. Specifically, a multi-scale spatial Transformer (MSST) is designed to leverage stage-by-stage spatial reduction and channel doubling, thereby enhancing the representation capabilities for the tracked target. In dynamic feature learning, an adaptive temporal Transformer (ATT) is then introduced based on multiple cross attentions, which analyzes the adaptive learning capacity for the dynamic target. It analyzes the weight proportion of different attentions automatically in the specific sequence through the learnable parameters. Finally, a multi-scale feature (MSF) regression module is crafted to improve the positioning accuracy of targets with low pixel counts in satellite scenes. This module accomplishes precise annotation of target boxes by effectively fusing features from diverse stages. We evaluate the proposed tracker performance on several public satellite datasets, including SatSOT, SV248S, and VISO. Experimental results show that the performance of our model can be comparable to the state-of-the-art trackers. Xin Zhang 0167, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Globally-Aware Continuous-Time Redistribution Learning for RS Image Change DetectionabstractChange detection (CD) based on deep learning has achieved excellent performance in recent years. However, these models exhibit limited capability in complete temporal modeling or face problems with fine-grained spatial features being overshadowed by the temporal context. Pure CNN-based CD pipelines also struggle to establish long-range connections. In this article, a globally aware continuous-time redistribution network (GCRNet) is proposed for RSCD. First, a boundary extraction branch is designed to preserve the semantic invariance of objects within the same boundary. This is achieved by providing boundary attention to adaptively guide the integration of temporal and spatial information. Then, a globally aware operator (GAO) is developed to obtain global interaction features. GAO utilizes the convolution theorem, which combines the Fourier transform and inverse Fourier transform, achieving it with low computational costs. Finally, an adaptive feature redistribution (AFR) module is designed to increase the distance between positive and negative samples in the latent space with change perception. It alleviates the effects of the severe class imbalance issue. Experimental results demonstrate that our proposed GCRNet surpasses 13 state-of-the-art CD methods. It achieves F1-score 0.33%, 0.62%, 0.84%, 0.17%, and 1.54% higher than the second-best model on the LEVIR-CD, LEVIR-CD+, WHU, CDD, and DSIFN datasets. The code of GCRNet is available athttps://github.com/XiaowenZhang-kuku/GCRNet. Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Effective and Robust: A Discriminative Temporal Learning Transformer for Satellite VideosabstractRobust feature learning has always been a research hotspot in dynamic temporal tasks. It makes the model almost unaffected by some challenging properties. The sequential nature of the transformer means attractive for temporal learning tasks, making it perform well in the video field. It is a current research hotspot for learning effective features by utilizing the target motion trends in satellite videos with multiple attributes, such as similar objects (SOBs) interference and occlusion. In this article, a novel discriminative temporal learning transformer tracker (DTLTracker) is introduced to characterize the dynamic target information for satellite videos. A discriminative transformer (DT) is proposed to comprehensively explore the dynamic target features with multiple attention mechanisms. It focuses on the primary information of the search area, making the target more discriminative. A fast convergence (FC) filter is designed to accelerate the weights convergence in calculating the target correlation operation, thereby ensuring the efficiency of model learning. The effectiveness and convergence have been demonstrated for the proposed optimization method. Additionally, a motion prior correction (MPC) module is constructed to utilize temporal information for target tracklet prediction, assisting the tracker in predicting the correct target. Numerous experiments are performed on three satellite videos to verify the effectiveness and feasibility of the proposed DTLTracker. It shows robustness compared to the state-of-the-art trackers on some challenging properties. Xin Zhang 0167, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Renormalized Connection for Scale-Preferred Object Detection in Satellite ImageryabstractSatellite imagery, due to its long-range imaging, brings with it a variety of scale-preferred tasks, such as the detection of tiny/small objects, making the precise localization and detection of small objects of interest a challenging task. In this article, we design a knowledge discovery network (KDN) to implement the renormalization group theory in terms of efficient feature extraction (FE). Renormalized connection (RC) on the KDN enables “synergistic focusing” of multiscale features. Based on our observations of KDN, we abstract a class of RCs with different connection strengths, called$n21$C, and generalize it to feature pyramid network (FPN)-based multibranch detectors. In a series of FPN experiments on the scale-preferred tasks, we found that the “divide-and-conquer” idea of FPN severely hampers the detector’s learning in the right direction due to the large number of large-scale negative samples and interference from background noise. Moreover, these negative samples cannot be eliminated by the focal loss function. The RCs extends the multilevel feature’s “divide-and-conquer” mechanism of the FPN-based detectors to a wide range of scale-preferred tasks, and enables synergistic effects of multilevel features on the specific learning goal. In addition, interference activations in two aspects are greatly reduced and the detector learns in a more correct direction. Extensive experiments of 17 well-designed detection architectures embedded with$n21$Cs on five different levels of scale-preferred tasks validate the effectiveness and efficiency of the RCs. Especially the simplest linear form of RC—E421C performs well in all tasks, and it satisfies the scaling property of renormalization group theory. All experiments can be trained and tested on a graphics card with 8 GB of video memory, which greatly enhances the applicability of our methodology. We hope that our approach will transfer a large number of well-designed detectors from the computer vision community to the remote sensing community. Datasets and codes will be available at:https://github.com/rabbitme/ Fan Zhang 0041, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Robust Instance-Based Semi-Supervised Learning Change Detection for Remote Sensing ImagesabstractSemi-supervised change detection (SSCD) has experienced rapid development, with numerous semi-supervised methods being proposed to reduce the reliance on labeled data in change detection. Existing approaches typically rely on manually set high-confidence thresholds to select robust pseudo-labels. However, the single-pixel threshold filtering method for pseudo-labels (STFP) lacks context correlation, cannot eliminate high-confidence false positive samples, and leads to erroneously filtering out low-confidence true positive samples. To address this issue, we propose robust instance-based semi-supervised learning change detection (RISL) for remote sensing images. RISL evaluates the reliability of each instance object by linking the semantic information of the context, thereby generating robust pseudo-labels. In RISL, firstly, a simple boundary trimming module (BT) as a preprocessing method for change prediction map is introduced. BT can effectively remove low-confidence false positive samples while avoiding confusion in the category of instance objects, thereby improving the quality of instance objects. Then, we propose a reliable instance evaluation module (RIEM) to evaluate the reliability of each instance object. RIEM combines the semantic information of the entire instance and establishes correlations between sample contexts to determine the reliability of the instance, effectively eliminating high false positive samples. In addition, the consistency regularization (CR) is integrated into RISL, and a new strategy suitable for RIEM is constructed. This strategy enhances the model’s generalization ability by mining and hiding semantic information from different views of unlabeled data. Experimental results on the challenging WHU-CD, LEVIR-CD, and CDD-CD datasets show that the proposed method achieves 89.80%, 90.01%, and 87.56% F1 scores on labeled data with 5% distribution. RISL achieves state-of-the-art performance compared to other methods. Yi Zuo 0003, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Hierarchical Dynamic Graph Clustering NetworkabstractConnections between visual components are ubiquitous. Graphs, as a highly flexible data structure, not only allow imposing relational induction bias on data, but can provide a completely distinct learning perspective for regular image data. In this paper, we propose a hierarchical dynamic graph clustering network (HDGCN) for visual feature learning. We construct hierarchical graph representations in graph domain in an adaptive, data-adaptive and task-adaptive manner. First, the initial graph is constructed in high-dimensional feature domain of images. To mine the hierarchical geometric features in latent graph space, adaptive clustering network (ClusterNet) is performed to learn discriminative clusters and generates cluster-based coarse graph. Then, graph convolutional networks (GCNs) are used to diffuse, transform and aggregate information among clusters. So, the intra-class and inter-class information is fully explored to increase the discriminativity of graph representations. Next, coarsened graph representations are mapped to grid based on its affinity with linear projection features. To further improve the task adaptation of clusters and hierarchical graph representations, ClusterNet and GCNs are fused in the same framework for end-to-end training and clusters is updated dynamically. We have conducted extensive experiments on classification and segmentation tasks. The experimental results fully validate the robustness of the proposed algorithm. Jie Chen 0098, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Puhua Chen, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Fast and Effective: Progressive Hierarchical Fusion Classification for Remote Sensing ImagesabstractMultisource remote sensing image fusion classification aims to produce accurate pixel-level classification maps by combining complementary information from different sources of remote sensing data. Existing methods based on Convolutional Neural Networks (CNN-based) utilize a patch-based learning framework, which has a high computational cost, leading to poor real-time performance. In contrast, methods based on Fully Convolutional Networks (FCN-based) can process the entire image directly, achieving fast inference. However, FCN-based methods require high computational resources and exhibit shortcomings in feature fusion, hindering practical applications. In this paper, a lightweight FCN-based Progressive Hierarchical Fusion Network (PHFNet) is tailored for multisource remote sensing image classification. PHFNet comprises a pyramid dual-path encoder and a pyramid decoder. In the encoder, cross-source features are hierarchically fused via the adaptive modulation fusion module (AMF), which leverages style calibration for cross-source alignment and promotes the complementarity of the fusion feature. In the decoder, we introduced an improved convolutional gated recurrent unit (iConvGRU) to progressively integrate the semantic and detailed information of hierarchical features, producing a context-enhanced global representation. In addition, we consider the relation between the channel number, convolutional kernel size, and parameter count to make the model as lightweight as possible. Comprehensive evaluations on three multisource remote sensing datasets demonstrate that PHFNet improves overall accuracy by 1.5% to 2.8% with a low computational overhead compared to state-of-the-art methods. The source code is avaliable athttps://github.com/ShirlySmile/PHFNet. Xueli Geng, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Cross-Domain Scene Unsupervised Learning Segmentation With Dynamic SubdomainsabstractUnsupervised cross-domain scene segmentation approach adapts the source model to the target domain, which utilizes two-stage strategies to minimize the inter-domain and intra-domain gap. However, the accumulation of errors in the previous stages affects the training of the subsequent stages. In this paper, a framework called statistical and structural domain adaptation (SSDA) is proposed to optimize inter-domain and intra-domain adaptation jointly. Firstly, the statistical inter-domain adaptation (StaIA) is proposed to model dynamic subdomains, which continuously adjust seed samples during the process of domain adaptation to mitigate error accumulation. The dynamic subdomains are modeled by exploring Bayesian uncertainty statistics and global balance statistics, which alleviate the imbalance problem in uncertainty estimation. StaIA encourages the model to transfer comprehensive and genuine knowledge through the seed loss for inter-domain adaptation. Secondly, the structural intra-domain adaptation (StrIA) is proposed to align the intra-domain gap among dynamic subdomains by the structural priors. Specifically, the StrIA models structural priors by truncated conditional random field (TruCRF) loss within the neighborhood, which constrains intra-domain semantic consistency to reduce the intra-domain gap. Experimental results demonstrate the effectiveness of the proposed cross-domain scene segmentation approaches on two commonly-used unsupervised domain adaptation benchmarks. The code is available at https://github.com/ChicalH/SSDA. Pei He, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Ronghua Shang, Shuang Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | A Category-Aware Curriculum Learning for Data-Free Knowledge DistillationabstractConstructing effective proxy data is one of the core challenges in data-free knowledge distillation. The existing models ignore the influence of the category entanglement of the generated data on the distillation. To alleviate this issue, imitating the human learning process, a new category-aware curriculum learning mechanism is proposed in this paper to perform data-free knowledge distillation, called CCL-D. The main ideology of this category-aware curriculum learning mechanism is to provide a new learning mode for data generation and network training, which enables the model to realize the knowledge distillation process from easy to difficult through automated curriculum learning. In this novel learning mechanism, a category-aware monitoring module is proposed to constrain the category attribute of generated data. Based on this monitoring module, the curriculum learning process for data generation and network training is designed and applied. Initially, the generator is guided to obtain new data with clear category features. The utilization of data with apparent category features is easy for student network training, and it enables the student network to learn clear and significant category features at the early training stage. Subsequently, the generator is guided to generate data with category entanglement. Utilizing these new data with category entanglement problems can improve the recognition ability of the student network to interclass interference and enhance network robustness. The effectiveness of the CCL-D is verified on the six benchmark experimental datasets (MNIST, CIFAR-10, CIFAR-100, SVHN, Caltech-101, Tiny-Imagenet). Xiufang Li, Licheng Jiao, Qigong Sun, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Multi-Scale Contourlet Knowledge Guide Learning SegmentationabstractFor accurate segmentation, effective feature extraction has always been a challenging problem, since the variability of appearance and the fuzziness of object boundaries. Convolutional neural networks have recently gained recognition in feature representation learning. However, it is only conducted in the spatial domain, and lacks effective representation of directionality, singularity and regularity in the spectral domain for anomaly detection of images. This is the key to feature learning representation of high-order singularity. To solve this problem, a multi-scale contourlet knowledge guide learning network is proposed in this paper. It is novel in this sense that, different from the CNNs in the spatial domain, the proposed method learns the multi-scale contourlet sparse representation to obtain more effective and sparse features in multi-scales and multi-directions. Furthermore, the contourlet knowledge guide learning can enhance the representation of spectral domain features. It is shown that the proposed network can learn the multi-level discriminative features and capture the more accurate object boundaries. The segmentation ability in theoretical analysis and experiments on five polyp segmentation datasets (CVC-ColonDB, CVC-ClinicDB, Kvasir-SEG, ETIS-LaribPolypDB, EndoSceneStill) and two building datasets (Massachusetts, WHU) are compared with developed methods. It must be emphasized that there is potential in effective feature learning representation and the generalization capability of the proposed method in deep learning, recognition and interpretation. Mengkun Liu, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Shuang Wang 0001, Biao Hou |
IEEE Trans. Multim. | 3 |
| 2024 | Bio-Inspired Multi-Scale Contourlet Attention NetworksabstractInspired by the sparse and hierarchical features representation in the ventral stream of the human visual system, the biologically inspired multi-scale contourlet attention network (BMCAnet) is proposed to extract robust discriminative features. First, we constructed the multi-scale contourlet filter banks as a population of neurons in the primary visual cortex (V1), and extracted sparse features in a multi-scale and multi-direction way. It simulated a simple cell in V1 that responds to stimuli in a specific direction. Second, in order to refine contourlet features adaptively, the Shannon block attention module (SBAM) is introduced by integrating Shannon entropy as the third branch of the channel attention module (CAM), thus the weights of contourlet coefficients can be learned adaptively. Third, the responses of the spatial and spectral features are pooled by the proposed contourlet pooling layer to obtain the invariant structure features with the specified rules, which roughly stimulate the pooling process of complex cells in the V1 area. Last, the combination of global average pooling (GAP) and full connection (FC) is used for classification. The competitive results on eight databases demonstrate that the BMCAnet can effectively extract sparse and effective features for the classification tasks. Mengkun Liu, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Xiangrong Zhang |
IEEE Trans. Multim. | 3 |
| 2024 | Feature Distribution Representation Learning Based on Knowledge Transfer for Long-Tailed ClassificationabstractReal-world data typically follows a long-tailed distribution. When a small sample of tail classes does not cover the underlying distribution well, methods such as class re-balancing strategies and decoupled training are difficult to work, and additional knowledge needs to be introduced to recover the underlying distribution of the tail classes. In this work, we observe that the similarity between the variances of the feature distributions increases with the class similarity. Then, we also find that well-represented feature distributions typically contain multiple subcenters, which allows for denser samples at the edges of the distribution and promotes model learning to more robust decision bounds. Based on these observations, we propose to calibrate the feature distribution of the tail class by transferring the variance of the feature distribution of the head class, and then sample from the calibrated tail class distribution to generate augmented samples. To coordinate with the tail class calibration method, we also propose label-aware noise suppression (LANS) for reducing the generation of noisy samples and a three-stage training scheme for reshaping decision boundaries and compacting feature learning. Experimental results on iNaturalist2018, ImageNet-LT, CIFAR-10-LT, and CIFAR-100-LT show that our method achieves state-of-the-art performance in most metrics compared to similar approaches. Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006, Puhua Chen |
IEEE Trans. Multim. | 5 |
| 2024 | Multiresolution Interpretable Contourlet Graph Network for Image ClassificationabstractModeling contextual relationships in images as graph inference is an interesting and promising research topic. However, existing approaches only perform graph modeling of entities, ignoring the intrinsic geometric features of images. To overcome this problem, a novel multiresolution interpretable contourlet graph network (MICGNet) is proposed in this article. MICGNet delicately balances graph representation learning with the multiscale and multidirectional features of images, where contourlet is used to capture the hyperplanar directional singularities of images and multilevel sparse contourlet coefficients are encoded into graph for further graph representation learning. This process provides interpretable theoretical support for optimizing the model structure. Specifically, first, the superpixel-based region graph is constructed. Then, the region graph is applied to code the nonsubsampled contourlet transform (NSCT) coefficients of the image, which are considered as node features. Considering the statistical properties of the NSCT coefficients, we calculate the node similarity, i.e., the adjacency matrix, using Mahalanobis distance. Next, graph convolutional networks (GCNs) are employed to further learn more abstract multilevel NSCT-enhanced graph representations. Finally, the learnable graph assignment matrix is designed to get the geometric association representations, which accomplish the assignment of graph representations to grid feature maps. We conduct comparative experiments on six publicly available datasets, and the experimental analysis shows that MICGNet is significantly more effective and efficient than other algorithms of recent years. Jie Chen 0098, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Multiscale Dynamic Curvelet Scattering NetworkabstractThe feature representation learning process greatly determines the performance of networks in classification tasks. By combining multiscale geometric tools and networks, better representation and learning can be achieved. However, relatively fixed geometric features and multiscale structures are always used. In this article, we propose a more flexible framework called the multiscale dynamic curvelet scattering network (MSDCCN). This data-driven dynamic network is based on multiscale geometric prior knowledge. First, multiresolution scattering and multiscale curvelet features are efficiently aggregated in different levels. Then, these features can be reused in networks flexibly and dynamically, depending on the multiscale intervention flag. The initial value of this flag is based on the complexity assessment, and it is updated according to feature sparsity statistics on the pretrained model. With the multiscale dynamic reuse structure, the feature representation learning process can be improved in the following training process. Also, multistage fine-tuning can be performed to further improve the classification accuracy. Furthermore, a novel multiscale dynamic curvelet scattering module, which is more flexible, is developed to be further embedded into other networks. Extensive experimental results show that better classification accuracies can be achieved by MSDCCN. In addition, necessary evaluation experiments have been performed, including convergence analysis, insight analysis, and adaptability analysis. Jie Gao 0013, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | A Patch Diversity Transformer for Domain Generalized Semantic SegmentationabstractDomain generalization (DG) is one of the critical issues for deep learning in unknown domains. How to effectively represent domain-invariant context (DIC) is a difficult problem that DG needs to solve. Transformers have shown the potential to learn generalized features, since the powerful ability to learn global context. In this article, a novel method named patch diversity Transformer (PDTrans) is proposed to improve the DG for scene segmentation by learning global multidomain semantic relations. Specifically, patch photometric perturbation (PPP) is proposed to improve the representation of multidomain in the global context information, which helps the Transformer learn the relationship between multiple domains. Besides, patch statistics perturbation (PSP) is proposed to model the feature statistics of patches under different domain shifts, which enables the model to encode domain-invariant semantic features and improve generalization. PPP and PSP can help to diversify the source domain at the patch level and feature level. PDTrans learns context across diverse patches and takes advantage of self-attention to improve DG. Extensive experiments demonstrate the tremendous performance advantages of the PDTrans over state-of-the-art DG methods. Pei He, Licheng Jiao, Ronghua Shang, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Xiangrong Zhang, Shuang Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | An Adaptive Migration Collaborative Network for Multimodal Image ClassificationabstractThe multispectral (MS) and the panchromatic (PAN) images belong to different modalities with specific advantageous properties. Therefore, there is a large representation gap between them. Moreover, the features extracted independently by the two branches belong to different feature spaces, which is not conducive to the subsequent collaborative classification. At the same time, different layers also have different representation capabilities for objects with large size differences. In order to dynamically and adaptively transfer the dominant attributes, reduce the gap between them, find the best shared layer representation, and fuse the features of different representation capabilities, this article proposes an adaptive migration collaborative network (AMC-Net) for multimodal remote-sensing (RS) images classification. First, for the input of the network, we combine principal component analysis (PCA) and nonsubsampled contourlet transformation (NSCT) to migrate the advantageous attributes of the PAN and the MS images to each other. This not only improves the quality of images themselves, but also increases the similarity between the two images, thereby reducing the representational gap between them and the pressure on the subsequent classification network. Second, for the interaction on the feature migrate branch, we design a feature progressive migration fusion unit (FPMF-Unit) based on the adaptive cross-stitch unit of correlation coefficient analysis (CCA), which can make the network automatically learn the features that need to be shared and migrated, aiming to find the best shared-layer representation for multifeature learning. And we design an adaptive layer fusion mechanism module (ALFM-Module), which can adaptively fuse features of different layers, aiming to clearly model the dependencies among multiple layers for different sized objects. Finally, for the output of the network, we add the calculation of the correlation coefficient to the loss function, which can make the network converge to the global optimum as much as possible. The experimental results indicate that AMC-Net can achieve competitive performance. And the code for the network framework is available at: https://github.com/ru-willow/A-AFM-ResNet. Wenping Ma 0001, Mengru Ma, Licheng Jiao, Fang Liu 0001, Hao Zhu 0009, Xu Liu 0006, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Robust and Effective: A Deep Matrix Factorization Framework for ClassificationabstractFor complex data, high dimension and high noise are challenging problems, and deep matrix factorization shows great potential in data dimensionality reduction. In this article, a novel robust and effective deep matrix factorization framework is proposed. This method constructs a dual-angle feature for single-modal gene data to improve the effectiveness and robustness, which can solve the problem of high-dimensional tumor classification. The proposed framework consists of three parts, deep matrix factorization, double-angle decomposition, and feature purification. First, a robust deep matrix factorization (RDMF) model is proposed in the feature learning, to enhance the classification stability and obtain better feature when faced with noisy data. Second, a double-angle feature (RDMF-DA) is designed by cascading the RDMF features with sparse features, which contains the more comprehensive information in gene data. Third, to avoid the influence of redundant genes on the representation ability, a gene selection method is proposed to purify the features by RDMF-DA, based on the principle of sparse representation (SR) and gene coexpression. Finally, the proposed algorithm is applied to the gene expression profiling datasets, and the performance of the algorithm is fully verified. Chenxi Tian, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | A Fast Evolutionary Knowledge Transfer Search for Multiscale Deep Neural ArchitectureabstractThe emergence of neural architecture search (NAS) algorithms has removed the constraints on manually designed neural network architectures, so that neural network development no longer requires extensive professional knowledge, trial and error. However, the extremely high computational cost limits the development of NAS algorithms. In this article, in order to reduce computational costs and to improve the efficiency and effectiveness of evolutionary NAS (ENAS) is investigated. In this article, we present a fast ENAS framework for multiscale convolutional networks based on evolutionary knowledge transfer search (EKTS). This framework is novel, in that it combines global optimization methods with local optimization methods for search, and searches a multiscale network architecture. In this article, evolutionary computation is used as a global optimization algorithm with high robustness and wide applicability for searching neural architectures. At the same time, for fast search, we combine knowledge transfer and local fast learning to improve the search speed. In addition, we explore a multiscale gray-box structure. This gray box structure combines the Bandelet transform with convolution to improve network approximation, learning, and generalization. Finally, we compare the architectures with more than 40 different neural architectures, and the results confirmed its effectiveness. Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Curvature-Balanced Feature Manifold Learning for Long-Tailed ClassificationabstractTo address the challenges of long-tailed classification, researchers have proposed several approaches to reduce model bias, most of which assume that classes with few samples are weak classes. However, recent studies have shown that tail classes are not always hard to learn, and model bias has been observed on sample-balanced datasets, suggesting the existence of other factors that affect model bias. In this work, we systematically propose a series of geometric measurements for perceptual manifolds in deep neural networks, and then explore the effect of the geometric characteristics of perceptual manifolds on classification difficulty and how learning shapes the geometric characteristics of perceptual manifolds. An unanticipated finding is that the correlation between the class accuracy and the separation degree of perceptual manifolds gradually decreases during training, while the negative correlation with the curvature gradually increases, implying that curvature imbalance leads to model bias. Therefore, we propose curvature regularization to facilitate the model to learn curvature-balanced and flatter perceptual manifolds. Evaluations on multiple long-tailed and non-long-tailed datasets show the excellent performance and exciting generality of our approach, especially in achieving significant performance improvements based on current state-of-the-art techniques. Our work opens up a geometric analysis perspective on model bias and reminds researchers to pay attention to model bias on non-long-tailed and even sample-balanced datasets. The code and model will be made public. Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006, Lingling Li 0002 |
CVPR | 5 |
| 2023 | Delving into Semantic Scale Imbalance
Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006 |
ICLR | 6 |
| 2023 | Swin Resnetswin Transformers for Change Detection in Remote Sensing ImagesabstractThe change detection task of remote sensing images is a basic scientific problem, and has been further widely used in real life. Recently, transformer model has shown strong learning and representation abilities in visual interpretation. In this article, Inspired by the success of the Vision Transformer and its variants, we propose a novel change detection model for remote sensing images, named Swin ResNet Transformers (Swin ResNet). Different from other methods, the proposed Swin ResNet architecture uses a Swin transform encoder, which extracts feature representations of multiple resolutions through a shift window mechanism to calculate self-attention. On three datasets, the proposed model showed good performance, and demonstrate that the Swin transformer has a strong ability to learn long-term dependencies of multi-scale context representation. Xu Liu 0006, Yu Liu 0005, Licheng Jiao, Lingling Li 0002, Fang Liu 0001 |
IGARSS | 1 |
| 2023 | A Strong Vision Transformer Adapter with Adaptive Thresholding for fine-Grained Building ClassificationabstractFine-grained building classification provides a solid basis for the comparison of city morphologies and the investigation of urban planning. To this aim, the DFC23 establishes a large-scale and multi-modal benchmark for the classification of building roof types. However, the problems of long-tailed distribution, data insufficient, inter-class similarity, and intra-class difference severely inhibit the performance of the detector. In this work, we build a strong vision transformer adapter fine-tuned on the cropped building instances to enhance the capacity of feature extraction and design a cross-modal fusion (CMF) module to effectively aggregate features from RGB and SAR data. When transferring to building instance segmentation, we construct a robust training pipeline and a two-stage test-time results ensemble scheme. Furthermore, we introduce self-training with two key denoising techniques, global average filtering (GAF) and intra-class adaptive thresholding (IAT), to boost the generalization of the model. Experimental results show the effectiveness of our method, ranking 2nd in the test phase of the contest. Xiaoqiang Lu, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Yuting Yang 0008 |
IGARSS | 6 |
| 2023 | Trident Cooperation Network for Building Extraction and Height EstimationabstractBuilding extraction and height estimation provide solid fundamentals for reconstructing city morphologies and investigating urban planning. To this aim, the DFC23 establishes a large-scale and multi-modal benchmark for multi-task learning of building reconstruction. However, the problems of data limitation and fore-background confusion severely inhibit the performance of the model. In this work, we propose a novel trident cooperation network (TCNet) to perform end-to-end building extraction and height estimation using RGB and SAR data. Specifically, to enrich the feature representation and generalization of the shared backbone, we introduce a vision transformer adapter to inject vision-specific inductive biases and design a cross-modal fusion (CMF) module to effectively aggregate features from multi-modal data. For downstream visual tasks, we construct trident decoders including a detector, a lightweight MLP segmentation head, and a pixel-wise regression head. Moreover, to highlight the foreground object, we use the binary mask predicted by the MLP head to cooperate with the height estimation map predicted by the estimator. And the weighted sub-task losses are gathered to optimize our TCNet. Experimental results show the effectiveness of our method, ranking 2nd in the test phase of the contest. Xiaoqiang Lu, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Yuting Yang 0008 |
IGARSS | 6 |
| 2023 | Dense Cross-Scale Transformer with Channel Learning for Remote Sensing Scene ClassificationabstractThe recent explosive on Transformer has suggested its potential to become the mainstream model for feature representation and classification of remote sensing. Although Transformer has excellent global modeling capabilities, it lacks inductive bias. In contrast, CNNs show excellent performance in computer vision due to their strong inductive bias. To solve the above issue, and benefit from combining the Transformer model with CNNs, a dense cross-scale Transformer is proposed for remote sensing scene classification. Firstly, the attention aggregation based feature pyramid network (A2-FPN) is adopted to obtain the multi-scale features. And then, the multi-scale features are input into the multi-scale global features learning with channel attention (MSCA) module to obtain the global features. Besides, the multi-scale features are input into the dense cross-scale attention (DCSA) module to learn the multi-level cross-scale features. Finally, the outputs of these two modules are considered for computing the final class score. Experimental results obtained from the three public datasets indicate that the proposed method surpasses other remote sensing classification methods. Yuting Yang 0008, Xu Liu 0006, Wenping Ma 0001, Licheng Jiao |
IGARSS | 3 |
| 2023 | Orthogonal Uncertainty Representation of Data Manifold for Robust Long-Tailed LearningabstractIn scenarios with long-tailed distributions, the model's ability to identify tail classes is limited due to the under-representation of tail samples. Class rebalancing, information augmentation, and other techniques have been proposed to facilitate models to learn the potential distribution of tail classes. The disadvantage is that these methods generally pursue models with balanced class accuracy on the data manifold, while ignoring the ability of the model to resist interference. By constructing noisy data manifold, we found that the robustness of models trained on unbalanced data has a long-tail phenomenon. That is, even if the class accuracy is balanced on the data domain, it still has bias on the noisy data manifold. However, existing methods cannot effectively mitigate the above phenomenon, which makes the model vulnerable in long-tailed scenarios. In this work, we propose an Orthogonal Uncertainty Representation (hOUR) of feature embedding and an end-to-end training strategy to improve the long-tail phenomenon of model robustness. As a general enhancement tool, OUR has excellent compatibility with other methods and does not require additional data generation, ensuring fast and efficient training. Comprehensive evaluations on long-tailed datasets show that our method significantly improves the long-tail phenomenon of robustness, bringing consistent performance gains to other long-tailed learning methods. Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006, Lingling Li 0002 |
ACM Multimedia | 5 |
| 2023 | MinEnt: Minimum entropy for self-supervised representation learning
Shuo Li 0010, Fang Liu 0001, Zehua Hao, Licheng Jiao, Xu Liu 0006, Yuwei Guo 0001 |
Pattern Recognit. | 5 |
| 2023 | Knowledge transduction for cross-domain few-shot learning
Pengfang Li, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006 |
Pattern Recognit. | 6 |
| 2023 | Knowledge transfer evolutionary search for lightweight neural architecture with dynamic inference
Xiaoxue Qian, Fang Liu 0001, Licheng Jiao, Xiangrong Zhang, Shuo Li 0010, Puhua Chen, Xu Liu 0006 |
Pattern Recognit. | 8 |
| 2023 | Dual Wavelet Attention Networks for Image ClassificationabstractGlobal average pooling (GAP) plays an important role in traditional channel attention. However, there is the disadvantage of insufficient information to use the result of GAP as the channel scalar. At the same time, the existing spatial attention models focus on the areas of interest using average pooling or convolutional networks, but there is a loss of feature information and neglect of the structural feature. In this paper, dual wavelet attention is proposed, which can effectively alleviate the aforementioned problems and enhance the representation ability of CNNs. Firstly, the equivalence between the sum of the low-frequency subband coefficients of 2D DWT (Haar) and GAP is proved. On this basis, the statistical characteristics of low-frequency and high-frequency subbands are effectively combined to obtain the channel scalars, which can better measure the importance of each channel. In addition, 2D DWT can effectively capture the approximate and detailed structural features. Thus, wavelet spatial attention is proposed, which can effectively focus on the key spatial structural features. Different from traditional spatial attention, it can better curve the structural and spatial attention for different channels. The experiments are verified on four natural image data sets and three remote sensing scene classification data sets, which shows the effectiveness and versatility of the proposed methods. The code of this paper will be available athttps://github.com/yutinyang/DWAN. Yuting Yang 0008, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Lingling Li 0002, Puhua Chen, Xiufang Li, Zhongjian Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Learning Salient Feature for Salient Object Detection Without LabelsabstractSupervised salient object detection (SOD) methods achieve state-of-the-art performance by relying on human-annotated saliency maps, while unsupervised methods attempt to achieve SOD by not using any annotations. In unsupervised SOD, how to obtain saliency in a completely unsupervised manner is a huge challenge. Existing unsupervised methods usually gain saliency by introducing other handcrafted feature-based saliency methods. In general, the location information of salient objects is included in the feature maps. If the features belonging to salient objects are called salient features and the features that do not belong to salient objects, such as background, are called nonsalient features, by dividing the feature maps into salient features and nonsalient features in an unsupervised way, then the object at the location of the salient feature is the salient object. Based on the above motivation, a novel method called learning salient feature (LSF) is proposed, which achieves unsupervised SOD by LSF from the data itself. This method takes enhancing salient feature and suppressing nonsalient features as the objective. Furthermore, a salient object localization method is proposed to roughly locate objects where the salient feature is located, so as to obtain the salient activation map. Usually, the object in the salient activation map is incomplete and contains a lot of noise. To address this issue, a saliency map update strategy is introduced to gradually remove noise and strengthen boundaries. The visualization of images and their salient activation maps show that our method can effectively learn salient visual objects. Experiments show that we achieve superior unsupervised performance on a series of datasets. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Xu Liu 0006, Puhua Chen |
IEEE Trans. Cybern. | 4 |
| 2023 | Multisource Joint Representation Learning Fusion Classification for Remote Sensing ImagesabstractMultisource remote sensing images provide complementary multidimensional information for reliable and accurate classification. However, gaps in imaging mechanisms result in heterogeneity between multiple source images. During fusion, this heterogeneity causes the generated multisource representations may be redundant and ignore discriminative uni-source information, which significantly hampers the fusion classification performance. To address this challenge, we introduce a novel multisource joint representation learning method for remote sensing image fusion classification, termed Multisource Information Bottleneck Fusion Network (MIBF-Net). Based on the Information Bottleneck principle, MIBF-Net employs mutual information constraints to effectively integrate multisource information, generating a comprehensive and non-redundant multisource representation. Specifically, MIBF-Net first introduces an attribution-driven noise adaptation layer to dynamically balance the speed of feature learning across sources for extracting discriminative uni-source intrinsic information. Furthermore, a cross-source relationship encoding module is designed to fully explore cross-source complex dependencies for enhancing the richness of fused representations. Finally, we design an information bottleneck fusion module to fuse uni-source semantic information and cross-source information while reducing redundancy. In particular, we employ variational inference techniques to effectively address the mutual information optimization problem and provide theoretical derivations. Extensive experimental results on three heterogeneous multisource remote sensing data benchmarks show that the model significantly outperforms the state-of-the-art methods. Xueli Geng, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001, Xiangrong Zhang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Weak-to-Strong Consistency Learning for Semisupervised Image SegmentationabstractSupervised remote sensing (RS) image segmentation has achieved remarkable success with large amounts of manually labeled data, which may be difficult to acquire in some practical application scenarios. Semisupervised RS image segmentation can efficiently utilize the knowledge embedded in unlabeled data to improve recognition performance, which is of great significance for the generalization application of segmentation models. In this work, we propose an end-to-end semisupervised RS image segmentation method based on weak-to-strong consistency learning, denoted as WSCL. Specifically, a common strong data augmentation technique for image segmentation is introduced to provide powerful input perturbation to decouple self-biased cognition. By forcing weakly augmented, and strongly augmented perspectives from the same sample to be consistent, WSCL not only enables the model to steadily learn knowledge contained in unlabeled data but also alleviates overfitting. In addition, a novel sparse dual-view cross-sample image generation method is presented to generate new training samples, which helps provide a more comprehensive diversity of perturbations. Furthermore, an adaptive re-weighting strategy based on the entropy maps of the outputs of strongly perturbed samples is proposed to suppress noise, guiding the training process in a positive direction. Extensive experiments demonstrate the significant advantage of WSCL over other advanced methods, achieving new state-of-the-art under several evaluation metrics on DFC22, iSAID, MER, MSL, Vaihingen, and GID-15 datasets. The source code is open-sourced at https://github.com/xiaoqiang-lu/WSCL. Xiaoqiang Lu, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001, Zhixi Feng, Puhua Chen |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | A Spatial-Spectral Bilinear Representation Fusion Network for Multimodal ClassificationabstractThe complementary and heterogeneous properties fusion of multimodal data (such as hyperspectral, lidar, and synthetic aperture radar data) can significantly improve the accuracy of remote sensing (RS) images joint classification. Thus, we propose a spatial-spectral bilinear representation fusion network (S2BRFNet), which captures long-range dependencies cross-modality and within the same modality to achieve the final joint classification. Firstly, a cross-modal spatial-spectral representation module (S2RM) is designed, it utilizes spatial-spectral attention and self-attention between heterogeneous data to enhance the characterization capabilities of cross-modal complementary properties and spatial-spectral features of single-source data. Secondly, a semantic space-guided bilinear feature fusion module (S2BFM) is developed, which uses deep and shallow features to regain fine-grained features. It uses shallow location details to improve the semantic prediction of deep features. Furthermore, it uses the different representation capabilities of different layers for objects with obvious feature differences to enhance the feature advantages. Therefore, rich global context information is obtained. Finally, the semantic space re-weight strategy is used to guide the outer product fusion of heterogeneous features, which enhances the ability of the network to identify similar features. Classification experiments are carried out on four common datasets of different modality combinations (HS-SAR-DSM Augsburg, Berlin, Trento, and Muufl), and this can prove the superiority of the S2BRFNet. Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Which Target to Focus on: Class-Perception for Semantic Segmentation of Remote SensingabstractDeep Learning-based (DL) methods have dominated the task of semantic segmentation of remote sensing images. However, the sizes of different objects vary widely, and there is a great deal of label-noise due to the inevitable shadows. Therefore, there is an urgent need for a method that can precisely handle complex ground data. In this paper, we propose an Inter-Class Enhanced Network (ICEN) for representing features of varying sizes. It comprises two branches: Sparse Representation Network (SPN) and Feature Extraction Network (FEN). Then, a Class-Perception Block is inserted between the two branches to instruct the SPN’s low-level semantic features to be merged into the deeper network. Such a block can reduce label-noise in remote sensing image segmentation. In addition, the proposed EIRI provides a more precise classification process for target edges containing many misclassified points without requiring excessive computational overhead. The experimental results of our proposed Class-Perception Network (C-PNet) achieve competitive performance on the Vaihingen, Potsdam, LoveDA, and UAVid datasets. Lingling Li 0002, Yilin Shao, Licheng Jiao, Xu Liu 0006, Puhua Chen, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | SDCDNet: A Semi-Dual Change Detection Network Framework With Super-Weak Label for Remote Sensing ImageabstractMost current change detection methods require a large amount of labeled data to train huge parameters. To break this limitation, this paper proposes a novel semi-supervised learning framework for remote sensing change detection, named a semi-dual change detection network (SDCDNet). The SDCDNet consists of a dual shared network and dual branching networks. The dual shared network is designed to exploit the full potential of the data, and the dual branching network is proposed to differentiate the kinds of annotated data and eliminate the disturbance between different types of data. In addition, the adaptive weighting module (AWM) enhances the features of weak branching, and the mask constraint module (MCM) is proposed to increase the ability of the network to extract foreground features. To solve the complex problem of data labeling, a patch-based weak label construction method is proposed to build super-weak labels. Experiments show that the proposed SDCDNet achieves excellent results on two remote sensing image change detection datasets. Jiahao Wang 0002, Fang Liu 0001, Hao Wang 0211, Xu Liu 0006, Licheng Jiao, Lingling Li 0002, Puhua Chen |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | CSLT: Contourlet-Based Siamese Learning Tracker for Dim and Small Targets in Satellite VideosabstractMost popular visual trackers for natural scenarios always adopt handcraft features or deep features to track the target in a video. However, they face with difficulties in discriminative feature representation and usually suffer from severe model drift for satellite videos, especially when encountering challenges of dim and small targets, low contrast or similar target interference. To overcome these difficulties, we propose a Contourlet-based Siamese Learning Tracker (CSLT), which mainly aims at tracking dim and small objects in satellite videos. In contrast to conventional methods, the contourlet transform enriches directional multi-resolution information which is crucial to discriminative feature representation for dim and small targets in satellite video frames that lack distinguishable appearance features. We jointly use multi-resolution features with deep features by spatial-attention fusion strategy and then track the targets by a Siamese structure network. To further improve the accuracy and robustness, a model drift alarm and calibration module, including translation drifting penalty and rotation drifting penalty, is employed during tracking. We conduct extensive comparisons with 16 popular state-of-the-art trackers on three satellite video datasets. The experimental results validate the effectiveness of the proposed tracker. Yinan Wu 0001, Licheng Jiao, Fang Liu 0001, Zhaoliang Pi, Xu Liu 0006, Lingling Li 0002, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | An Explainable Spatial-Frequency Multiscale Transformer for Remote Sensing Scene ClassificationabstractDeep convolutional neural networks (CNNs) are significant in remote sensing. Due to the strong local representation learning ability, CNNs have excellent performance in remote sensing scene classification. However, CNNs focus on location-sensitive representations in the spatial domain and lack contextual information mining capabilities. Meanwhile, remote sensing scene classification still faces challenges, such as complex scenes and significant differences in target sizes. To address the problems and challenges above, more robust feature representation learning networks are necessary. In this paper, a novel and explainable spatial-frequency multi-scale Transformer framework, SF-MSFormer, is proposed for remote sensing scene classification. It mainly comprises spatial-domain and frequency-domain multi-scale Transformer branches, which consider the spatial-frequency global multi-scale representation features. Besides, the texture-enhanced encoder is designed in the frequency-domain multi-scale Transformer branch, which is adaptive to capture the global texture features. In addition, an adaptive feature aggregation module is designed to integrate the spatial-frequency multi-scale feature for final recognition. The experimental results verify the effectiveness of SF-MSFormer and show better convergence. It achieves state-of-the-art results (98.72%, 98.6%, 99.72%, and 94.83% overall accuracies, respectively) on the AID, UCM, WHU-RS19, and NWPU-RESISC45 datasets. Besides, the feature visualizations evaluate the explainability of the texture-enhanced encoder. The code implementation of this article will be available at https://github.com/yutinyang/SF-MSFormer. Yuting Yang 0008, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Boundary-Aware Multiscale Learning Perception for Remote Sensing Image SegmentationabstractFor remote sensing image segmentation, the boundaries of objects are difficult to distinguish, which is ignored by most methods. Therefore, it is challenging how to excavate and recover the boundaries of objects accurately. In this article, we propose a boundary-aware multi-scale network (BMNet) to solve this problem. The key components of BMNet include the scale attention module (SA-module) and boundary guidance module (BG-module). Specifically, SA-module is proposed to guide the refinement of multi-scale features in a context-aware way. It enhances the discriminability of multi-scale features by establishing contextual dependencies, which enables the refinement of the prediction of objects. Then, BG-module is proposed to enable networks to distinguish the boundary of objects. It utilizes manifold information of features to generate boundary guidance maps and forces the network to focus more on the boundary of objects. The effectiveness of the proposed BMNet is demonstrated on two public remote sensing datasets: ISPRS 2-D semantic labeling Potsdam dataset and Vaihingen dataset, where BMNet achieves better segmentation than prevalent methods. Finally, the experimental results indicate that BMNet can produce sharper boundaries of objects to reconstruct more detailed segmentation results. Chao You, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | A Spatial Hierarchical Reasoning Network for Remote Sensing Visual Question AnsweringabstractFor visual question answering on remote sensing (RSVQA), current methods scarcely consider geospatial objects typically with large-scale differences and positional sensitive properties. Besides, modeling and reasoning the relationships between entities have rarely been explored, which leads to one-sided and inaccurate answer predictions. In this article, a novel method called spatial hierarchical reasoning network (SHRNet) is proposed, which endows a remote sensing (RS) visual question answering (VQA) system with enhanced visual–spatial reasoning capability. Specifically, a hash-based spatial multiscale visual representation module is first designed to encode multiscale visual features embedded with spatial positional information. Then, spatial hierarchical reasoning is conducted to learn the high-order inner group object relations across multiple scales under the guidance of linguistic cues. Finally, a visual-question (VQ) interaction module is employed to learn an effective image–text joint embedding for the final answer predicting. Experimental results on three public RS VQA datasets confirm the effectiveness and superiority of our model SHRNet. Zixiao Zhang, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Fang Liu 0001, Yuxuan Li 0004, Zhicheng Guo |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Curvelet Adversarial Augmented Neural Network for SAR Image ClassificationabstractConvolutional neural networks (CNNs) have superior feature learning capabilities with large numbers of labeled samples. The reality is that labeling these samples is costly in terms of human labor. Existing data augmentation methods alleviate the scarcity of labeled samples. However, these methods are not suitable for synthetic aperture radar (SAR) images, owing to special imaging mechanisms and observational objects. The generative SAR images by existing augmented methods show structure distortion. To address this issue, we introduce a curvelet adversarial augmented neural network (CA2NN) for SAR image classification. Specifically, an$\text{A}^{2}$NN is established, which consists of two generative streams and one discriminative stream. In the generative stream, through the mutual transformation between the whole and partial images, more new samples with structural consistency are generated to augment the limited labeled data. In the discriminative stream, these generated samples show certain appearance variations after adversarial training based on the novel joint discriminant criterion. Simultaneously, given the multiscale and multidirectional nature of SAR images, we construct discretized curvelet in 2-D space, aiming to extract the singularity features and avoid overfitting. By integrating curvelet kernels into$\text{A}^{2}$NN, CA2NN can automatically generate more representative features adapting to complex terrain, while greatly reducing the complexity of the network. Experiments are conducted on the SAR images with large-scale and complex scenes, suggesting that the proposed approach significantly improves the classification performance with few labeled samples. Yake Zhang, Fang Liu 0001, Licheng Jiao, Shuyuan Yang 0001, Lingling Li 0002, Meijuan Yang, Jianlong Wang, Xu Liu 0006 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2023 | GeoFormer: A Geometric Representation Transformer for Change DetectionabstractDeep representation learning has improved automatic remote change detection (RSCD) in recent years. Existing methods emphasize primarily convolutional neural networks (CNNs) or Transformer-based networks. However, most of them neither effectively combine CNNs and Transformer nor use prior geometric information to refine regions. In this paper, a novel geometric representation Transformer (GeoFormer) is proposed for high-resolution RSCD. GeoFormer utilizes convolutional information to guide the Transformer by employing geometric prior knowledge. Specifically, the proposed GeoFormer consists of three carefully designed components: the geometric-based Swin Transformer (Geo-Swin Transformer) encoder, the Laplace attention fusion (LAFusion) module, and the UNet++CD decoder. Firstly, Geo-Swin Transformer is a novel designed non-local Siamese encoder that combines geometric convolution with Transformer to provide local geometric representation information for remote contextual features. Then, a LAFusion module is proposed to achieve robust bi-temporal feature fusion, which is founded on attention mechanism and edge information. Finally, UNet++CD decodes fine-grained information from the fused features by dense multiscale upsampling process. Experimental results demonstrate that the proposed GeoFormer performs better than benchmark methods on four change detection datasets (LEVIR-CD, WHU-CD, DSIFN-CD, and CDD) and is able to detect the edges of change regions more precisely. Our code is available at https://github.com/Jiaxzhao/GeoFormer. Jiaxuan Zhao, Licheng Jiao, Chao Wang 0099, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | EPT-Net: Edge Perception Transformer for 3D Medical Image SegmentationabstractThe convolutional neural network has achieved remarkable results in most medical image seg- mentation applications. However, the intrinsic locality of convolution operation has limitations in modeling the long-range dependency. Although the Transformer designed for sequence-to-sequence global prediction was born to solve this problem, it may lead to limited positioning capability due to insufficient low-level detail features. Moreover, low-level features have rich fine-grained information, which greatly impacts edge segmentation decisions of different organs. However, a simple CNN module is difficult to capture the edge information in fine-grained features, and the computational power and memory consumed in processing high-resolution 3D features are costly. This paper proposes an encoder-decoder network that effectively combines edge perception and Transformer structure to segment medical images accurately, called EPT-Net. Under this framework, this paper proposes a Dual Position Transformer to enhance the 3D spatial positioning ability effectively. In addition, as low-level features contain detailed information, we conduct an Edge Weight Guidance module to extract edge information by minimizing the edge information function without adding network parameters. Furthermore, we verified the effectiveness of the proposed method on three datasets, including SegTHOR 2019, Multi-Atlas Labeling Beyond the Cranial Vault and the re-labeled KiTS19 dataset called KiTS19-M by us. The experimental results show that EPT-Net has significantly improved compared with the state-of-the-art medical image segmentation method. Licheng Jiao, Ronghua Shang, Xu Liu 0006, Longchang Xu |
IEEE Trans. Medical Imaging | 4 |
| 2023 | A Universal Quaternion Hypergraph Network for Multimodal Video Question AnsweringabstractFusion and interaction of multimodal features are essential for video question answering. Structural information composed of the relationships between different objects in videos is very complex, which restricts understanding and reasoning. In this paper, we propose a quaternion hypergraph network (QHGN) for multimodal video question answering, to simultaneously involve multimodal features and structural information. Since quaternion operations are suitable for multimodal interactions, four components of the quaternion vectors are applied to represent the multimodal features. Furthermore, we construct a hypergraph based on the visual objects detected in the video. Most importantly, the quaternion hypergraph convolution operator is theoretically derived to realize multimodal and relational reasoning. Question and candidate answers are embedded in quaternion space, and a Q&A reasoning module is creatively designed for selecting the answer accurately. Moreover, the unified framework can be extended to other video-text tasks with different quaternion decoders. Experimental evaluations on the TVQA dataset and DramaQA dataset show that our method achieves state-of-the-art performance. Zhicheng Guo, Jiaxuan Zhao, Licheng Jiao, Xu Liu 0006, Fang Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | D³K: Dynastic Data-Free Knowledge DistillationabstractData-free knowledge distillation further broadens the applications of the distillation model. Nevertheless, the problem of providing diverse data with rich expression patterns needs to be further explored. In this paper, a novel dynastic data-free knowledge distillation ($D^{3}K$) model is proposed to alleviate this problem. In this model, a dynastic supernet generator (D-SG) with a flexible network structure is proposed to generate diverse data. The D-SG can adaptively alter architectural configurations and activate different subnet generators in different sequential iteration spaces. The variable network structure increases the complexity and capacity of the generator, and strengthens its ability to generate diversified data. In addition, a novel additive constraint based on the differentiable dhash (D-Dhash) is designed to guide the structure parameter selection of the D-SG. This constraint forces the D-SG to constantly jump out of the fixed generation mode and generate diverse data in semantics and instance. The effectiveness of the proposed model is verified on the experimental benchmark datasets (MNIST, CIFAR-10, CIFAR-100, and SVHN). Xiufang Li, Qigong Sun, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Yi Zuo 0003 |
IEEE Trans. Multim. | 5 |
| 2023 | Transformer Based Conditional GAN for Multimodal Image FusionabstractMultimodal Image fusion is becoming urgent in multi-sensor information utilization. However, existing end-to-end image fusion frameworks ignore a priori knowledge integration and long-distance dependencies across domains, which brings challenges to the network convergence and global image perception in complex scenes. In this paper, a conditional generative adversarial network with transformer (TCGAN) is proposed for multimodal image fusion. The generator is to generate a fused image with the source images content. The discriminators are adopted to distinguish the differences between the fused image and the source images. Adversarial training makes the final fused image to maintain the structural and textural details in the cross-modal images simultaneously. In particular, a wavelet fusion module makes the inputs contain image content from different domains as much as possible. The extracted convolutional features interact in the multiscale cross-modal transformer fusion module to fully complement the associated information. It makes the generator to focus on both local and global context. TCGAN fully considers the training efficiency of the adversarial process and the integrated retention of redundant information. Various experimental results of TCGAN have highlighted targets, rich details, and fast convergence properties on public datasets. Jun Zhang 0045, Licheng Jiao, Wenping Ma 0001, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Multiscale Curvelet Scattering NetworkabstractFeature representation has received more and more attention in image classification. Existing methods always directly extract features via convolutional neural networks (CNNs). Recent studies have shown the potential of CNNs when dealing with images' edges and textures, and some methods have been explored to further improve the representation process of CNNs. In this article, we propose a novel classification framework called the multiscale curvelet scattering network (MSCCN). Using the multiscale curvelet-scattering module (CCM), image features can be effectively represented. There are two parts in MSCCN, which are the multiresolution scattering process and the multiscale curvelet module. According to multiscale geometric analysis, curvelet features are utilized to improve the scattering process with more effective multiscale directional information. Specifically, the scattering process and curvelet features are effectively formulated into a unified optimization structure, with features from different scale levels being efficiently aggregated and learned. Furthermore, a one-level CCM, which can essentially improve the quality of feature representation, is constructed to be embedded into other existing networks. Extensive experimental results illustrate that MSCCN achieves better classification accuracy when compared with state-of-the-art techniques. Eventually, the convergence, insight, and adaptability are evaluated by calculating the trend of loss function's values, visualizing some feature maps, and performing generalization analysis. Jie Gao 0013, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou, Xu Liu 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Learning Social Spatio-Temporal Relation Graph in the Wild and a Video BenchmarkabstractSocial relations are ubiquitous and form the basis of social structure in our daily life. However, existing studies mainly focus on recognizing social relations from still images and movie clips, which are different from real-world scenarios. For example, movie-based datasets define the task as the video classification, only recognizing one relation in the scene. In this article, we aim to study the problem of social relation recognition in an open environment. To close the gap, we provide the first video dataset collected from real-life scenarios, named social relation in the wild (SRIW), where the number of people can be huge and vary, and each pair of relations needs to be recognized. To overcome new challenges, we propose a spatio-temporal relation graph convolutional network (STRGCN) architecture, utilizing correlative visual features to recognize social relations intuitively. Our method decouples the task into two classification tasks: person-level and pair-level relation recognition. Specifically, we propose a person behavior and character module to encode moving and static features in two explicit ways. Then we take them as node features to build a relation graph with meaningful edges in a scene. Based on the relation graph, we introduce the graph convolutional network (GCN) and local GCN to encode social relation features which are used for both recognitions. Experimental results demonstrate the effectiveness of the proposed framework, achieving 83.1% and 40.8% mAP in person-level and pair-level classification. Moreover, the study also contributes to the practicality in this field. Haoran Wang 0008, Licheng Jiao, Fang Liu 0001, Lingling Li 0002, Xu Liu 0006, Deyi Ji, Weihao Gan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | RDLNet: A Regularized Descriptor Learning NetworkabstractLocal image descriptor learning has been instrumental in various computer vision tasks. Recent innovations lie with similarity measurement of descriptor vectors with metric learning for randomly selected Siamese or triplet patches. Local image descriptor learning focuses more on hard samples since easy samples do not contribute much to optimization. However, few studies focus on hard samples of image patches from the perspective of loss functions and design appropriate learning algorithms to obtain a more compact descriptor representation. This article proposes a regularized descriptor learning network (RDLNet) that makes the network focus on the learning of hard samples and compact descriptor with triplet networks. A novel hard sample mining strategy is designed to select the hardest negative samples in mini-batch. Then batch margin loss concerned with hard samples is adopted to optimize the distance of extreme cases. Finally, for a more stable network and preventing network collapsing, orthogonal regularization is designed to constrain convolutional kernels and obtain rich deep features. RDLNet provides a compact discriminative low-dimensional representation and can be embedded in other pipelines easily. This article gives extensive experimental results for large benchmarks in multiple scenarios and generalization in matching applications with significant improvements. Jun Zhang 0045, Licheng Jiao, Wenping Ma 0001, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Hao Zhu 0009 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Hierarchical Scene Normality-Binding Modeling for Anomaly Detection in Surveillance VideosabstractAnomaly detection in surveillance videos is an important topic in the multimedia community, which requires efficient scene context extraction and the capture of temporal information as a basis for decision. From the perspective of hierarchical modeling, we parse the surveillance scene from global to local and propose a Hierarchical Scene Normality-Binding Modeling framework (HSNBM) to handle anomaly detection. For the static background hierarchy, we design a Region Clustering-driven Multi-task Memory Autoencoder (RCM-MemAE), which can simultaneously perform region segmentation and scene reconstruction. The normal prototypes of each local region are stored, and the frame reconstruction error is subsequently amplified by global memory augmentation. For the dynamic foreground object hierarchy, we employ a Scene-Object Binding Frame Prediction module (SOB-FP) to bind all foreground objects in the frame with the prototypes stored in the background hierarchy according their positions, thus fully exploit the normality relationship between foreground and background. The bound features are then fed into the decoder to predict the future movement of the objects. With the binding mechanism between foreground and background, HSNBM effectively integrates the "reconstruction" and "prediction" tasks and builds a semantic bridge between the two hierarchies. Finally, HSNBM fuses the anomaly scores of the two hierarchies to make a comprehensive decision. Extensive empirical studies on three standard video anomaly detection datasets demonstrate the effectiveness of the proposed HSNBM framework. Qianyue Bao, Fang Liu 0001, Yang Liu 0349, Licheng Jiao, Xu Liu 0006, Lingling Li 0002 |
ACM Multimedia | 5 |
| 2022 | Learning Visible Surface Area Estimation for Irregular ObjectsabstractVisible surface area estimation for irregular objects, one of the most fundamental and challenging topics in mathematics, supports a wide range of applications. The existing techniques usually estimate the visible surface area via mathematical modeling from 3D point clouds. However, the 3D scanner is expensive, and the corresponding evaluation method is too complex. In this paper, we propose a novel problem setting, deep learning for visible surface area estimation, which is the first trial to estimate the visible surface area for irregular objects from monocular images. Technically, we first build a novel visible surface area estimation dataset including 9099 real annotations. Then, we design a learning-based architecture to predict the visible surface area, including two core modules (i.e., the classification module and the area-bins module). The classification module is presented to predict the visible surface area distribution interval and assist network training for more accurate visible surface area estimation. Meanwhile, the area-bins module using the transformer encoder is proposed to distinguish the difference in visible surface area between irregular objects of the same category. The experimental results demonstrate that our approach can effectively estimate the visible surface area for irregular objects with various categories and sizes. We hope that this work will attract further research into this newly identified, yet crucial research direction. Our source code and data are available at \textcolormagenta \urlhttps://github.com/liuxu0303/VSAnet . Xu Liu 0006, Jianing Li 0001, Xianqi Zhang, Xiaopeng Fan 0001, Yonghong Tian 0001 |
ACM Multimedia | 1 |
| 2022 | Augmentative contrastive learning for one-shot object detection
Yaoyang Du, Fang Liu 0001, Licheng Jiao, Zehua Hao, Shuo Li 0010, Xu Liu 0006, Jing Liu 0006 |
Neurocomputing | 6 |
| 2022 | Region NMS-based deep network for gigapixel level pedestrian detection with two-step cropping
Lingling Li 0002, Xiaohui Guo, Jingjing Ma 0001, Licheng Jiao, Fang Liu 0001, Xu Liu 0006 |
Neurocomputing | 7 |
| 2022 | Entire Deformable ConvNets for semantic segmentation
Bingqi Yu, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Xu Tang 0004 |
Knowl. Based Syst. | 3 |
| 2022 | CDANet: Common-and-Differential Attention Network for Object Detection and Instance Segmentation
Xiaohui Guo, Licheng Jiao, Xu Liu 0006 |
Pattern Recognit. Lett. | 5 |
| 2022 | Deep Multiview Union Learning Network for Multisource Image ClassificationabstractWith the development of the imaging technology of various sensors, multisource image classification has become a key challenge in the field of image interpretation. In this article, a novel classification method, called the deep multiview union learning network (DMULN), is proposed to classify multisensor data. First, an associated feature extractor is designed to process the multisource data by canonical correlation analysis (CCA) in the head of the network. Second, an improved deep learning architecture with two branches is presented to extract high-level view features from the associated features. Third, a novel pooling, called view union pooling, is proposed to fuse the multiview feature from the deep model. Finally, the fused feature is fed into the classifier. The proposed framework is easy to optimize since it is an end-to-end network. Extensive experiments and analysis on the datasets IEEE_grss_dfc_2017 and IEEE_grss_dfc_2018 show that the proposed method achieves comparable results. Our results demonstrate that abundant multisource information can improve the classification performance. Xu Liu 0006, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Cybern. | 1 |
| 2022 | GAFnet: Group Attention Fusion Network for PAN and MS Image High-Resolution ClassificationabstractPanchromatic (PAN) and multispectral (MS) images have coordinated and paired spatial spectral information, which can complement each other and make up for their shortcomings for image interpretation. In this article, a novel classification method called the deep group spatial-spectral attention fusion network is proposed for PAN and MS images. First, the MS image is processed by unpooling to obtain the same resolution as that of the PAN image. Second, the group spatial attention and group spectral attention modules are proposed to extract image features. The PAN and the processed MS images are regarded as the input of the two modules, respectively. Third, the features from the previous step are fused by the attention fusion module, which aims to fully fuse multilevel features, take into account both the low-level features and the high-level features, and maintain the global abstract and local detailed information of the pixels. Finally, the fusion feature is fed into the classifier and the resulting map is obtained by pixel level. Extensive experiments and analysis on four datasets show that the proposed method achieves comparable results. Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Biao Hou, Shuyuan Yang 0001, Licheng Jiao |
IEEE Trans. Cybern. | 1 |
| 2022 | Automatic Graph Learning Convolutional Networks for Hyperspectral Image ClassificationabstractThe excellent performance of graph convolutional networks (GCNs) on non-Euclidean data has drawn widespread attention from the hyperspectral image classification (HSIC) community, where the predefined graph (including node modeling and adjacency matrix calculation) plays a key role. However, existing GCN-based methods rely on manual efforts in constructing and updating graphs, and the superpixel-based node features lack high-level semantics. In this article, we propose an automatic graph learning convolutional network (Auto-GCN), which unifies the graph learning and HSIC in a “network-in-network” manner. Specifically, the graph is employed to model the interaction of the high-order tensors. Considering the powerful learning and representation capabilities of convolutional neural networks (CNNs), the semisupervised Siamese network (SiamNet) is embedded into GCNs and HSIC networks to accomplish the automatic learning and dynamic updating of the graph. GCNs further encode and infer the dynamic graph, and then, the learnable graph reprojection matrix is designed to assign graph representations to pixels. The dynamic graph serves the HSIC task during forward propagation, while the HSIC task continuously corrects the graph during backward propagation. Therefore, the “automatic” of the proposed Auto-GCN is not only reflected in the fact that the graph representation is designed and updated by an end-to-end network but is also HSIC task-oriented. The experimental results show that the proposed Auto-GCN outperforms other state-of-the-art methods on four publicly available hyperspectral datasets. Jie Chen 0098, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Deep Shearlet Network for Change Detection in SAR ImagesabstractConvolutional neural networks (CNN) can extract shift-invariant features, and have been widely applied in change detection task. However, common CNN lacks noise robustness and needs supervised data, to alleviate these problems, in this paper, we propose a novel deep shearlet network (ShearNet) for change detection in SAR images. In the network, a shearlet denoising layer (SDL) is designed to enhance the representation ability of common CNN. In SDL, feature maps are decomposed into subband coefficients by shearlet transform (ST). Due to optimal sparse representation property and highly direction sensitivity of ST, the network can capture important geometric information. Then, hard-threshold shrinkage is applied to high frequency subbands to drop small coefficients that are most likely to be noise, so that reduce the effect of noise. Finally, ShearNet is trained by introducing a noise-robust loss with noisy labels. The noisy labels are obtained by deep clustering that shows more robustness than existing preclassification methods. This fine-tuning process novelly follows the paradigm of learning from noisy labels to aside the difficulty of precisely labeling samples. Our experimental results on multiple real SAR datasets show that ShearNet can boost accuracy, and have better applicability for change detection in SAR images. The source code is available at https://github.com/yizhilanmaodhh/ShearNet. Huihui Dong, Licheng Jiao, Wenping Ma 0001, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | MANet: Multi-Scale Aware-Relation Network for Semantic Segmentation in Aerial ScenesabstractSemantic segmentation is an important yet unsolved problem in aerial scenes understanding. One of the major challenges is the intense variations of scenes and object scales. In this paper, we propose a novel multi-scale aware-relation network (MANet) to tackle this problem in remote sensing. Inspired by the process of human perception of multi-scale information, we explore discriminative and diverse multi-scale representations. For discriminative multi-scale representations, we propose an inter-class and intra-class region refinement method (IIRR) to reduce feature redundancy caused by fusion. IIRR utilizes the refinement maps with intra- and inter-class scale variation to guide multi-scale fine-grained features. Then, we propose multi-scale collaborative learning (MCL) to enhance the diversity of multi-scale feature representations. The MCL constrains the diversity of multi-scale feature network parameters to obtain diverse information. And the segmentation results are rectified according to the dispersion of the multi-level network predictions. In this way, MANet can learn multi-scale features by collaboratively exploiting the correlation among different scales. Extensive experiments on image and video datasets which have large scale variations have demonstrated the effectiveness of our proposed MANet. Pei He, Licheng Jiao, Ronghua Shang, Shuang Wang 0001, Xu Liu 0006, Dou Quan, Dong Zhao 0007 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Simple and Efficient: A Semisupervised Learning Framework for Remote Sensing Image Semantic SegmentationabstractSemantic segmentation based on deep learning has achieved impressive results in recent years, but these results are supported by a large amount of labeled data which requires intensive annotation at the pixel level, particularly for high-resolution remote sensing (RS) images. In this work, we propose a simple yet efficient semisupervised learning framework based on linear sampling self-training, named LSST, to improve the performance of RS image semantic segmentation. Specifically, the classical pseudo-labeling-based self-training paradigm is enhanced by injecting strong data augmentations (SDA) applicable to RS images, based on which a powerful baseline is constructed. Nevertheless, the problem of insufficient data training to generate pseudo-labels with a high level of noise persists, and the noisy pseudo-labels will continue to accumulate and impede model improvement during the re-training phase. Previous works commonly employ a pre-defined threshold to remove noise, but it will lead to overfitting the model to easily identified classes. To address it, a method using linear sampling (LS) is presented for assigning thresholds to different classes in an adaptive manner, which provides noiseless regions for re-training. Experiments prove that the proposed pixel-wise selection is more available for segmentation than image-level selection in RS images. Finally, LSST achieves state-of-the-art on several datasets and different evaluation metrics. The source code of the this paper is available at https://github.com/xiaoqiang-lu/LSST. Xiaoqiang Lu, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006, Zhixi Feng, Lingling Li 0002, Puhua Chen |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Transfer Representation Learning Meets Multimodal Fusion Classification for Remote Sensing ImagesabstractTo maximize the complementary advantages of synergistic multimodal, a transfer representation learning fusion network (TRLF-Net) is proposed for multisource remote sensing images collaborative classification in this article. First, with respect to the feature encoding, we design a dual-branch attention sparse transfer module (DAST-Module), which combines the spatial and channel attention (CA) masks to migrate the advantage attributes of the panchromatic (PAN) and the MS images mutually. This not only enhances their respective image advantages but also facilitates the sparse fusion of low-level features. Second, for the separation of multiscale information, a deep dual-scale decomposition module (DDSD-Module) is designed, which allows the decompose of high-frequency and low-frequency components. Then it uses the decomposed information to make the essential difference as small as possible, and the surrounding contour difference is as large as possible of the complementary multimodal image through the design of the loss function. Finally, to address the problem of large intraclass and small interclass differences, we develop a representation fusion of the global and local features’ module (RFGAL-Module). It mainly adopts global features to sort local features within classes, and then outputs them in a cascade. Thus, the characterization ability of features is improved, and the global and local features are used in a coordinated manner to accomplish the sample classification tasks. In particular, the experimental results demonstrate that TRLF-Net can obtain much improved accuracy and efficiency. The code is accessible in:https://github.com/ru-willow/SRLF-Net. Mengru Ma, Wenping Ma 0001, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Very Low-Resolution Moving Vehicle Detection in Satellite VideosabstractThis paper proposes a practical end-to-end neural network framework to detect tiny moving vehicles in satellite videos with low imaging quality. Some instability factors such as illumination changes, motion blurs, and low contrast to the cluttered background make it difficult to distinguish true objects from noise and other point-shaped distractors. Moving vehicle detection in satellite videos can be carried out based on background subtraction or frame differencing. However, these methods are prone to produce lots of false alarms and miss many positive targets. Appearance-based detection can be an alternative but is not well-suited since classifier models are of weak discriminative power for the vehicles in top view at such low resolution. This article addresses these issues by integrating motion information from adjacent frames to facilitate the extraction of semantic features and incorporating the Transformer to refine the features for key points estimation and scale prediction. Our proposed model can well identify the actual moving targets and suppress interference from stationary targets or background. The experiments and evaluations using satellite videos show that the proposed approach can accurately locate the targets under weak feature attributes and improve the detection performance in complex scenarios. Zhaoliang Pi, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Biao Hou, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | A Joint Siamese Attention-Aware Network for Vehicle Object Tracking in Satellite VideosabstractRemote sensing object tracking is a novel and challenging problem due to the negative effects of weak features and background noise. In this paper, from the perspective of attention-focus deep learning, we propose a Joint Siamese Attention-Aware Network (JSANet) for efficient remote sensing tracking which contains both self-attention and cross-attention modules. First, the self-attention modules we propose emphasize the interdependent channel-wise coefficient via channel attention and conduct corresponding space transformation of spatial domain information with spatial attention. Second, the cross-attention is designed to aggregate rich contextual interdependencies between the siamese branches via channel attention and excavate association produces reliable correspondence with spatial attention. In addition, a composite feature combine strategy is designed to fuse multiple attention features. Experimental results on the Jilin-1 satellite video datasets demonstrate that the proposed JSANet achieves state-of-the-art performance in terms of precision and success rate, demonstrate the effectiveness of the proposed methods. Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | MBLT: Learning Motion and Background for Vehicle Tracking in Satellite VideosabstractRecently, satellite videos provide a new way to dynamically monitor the Earth’s surface. The interpretation of satellite videos has attracted more and more attentions. In this article, we focus on the problem of the vehicle tracking in satellite videos. Satellite videos usually own a lower resolution, which leads to the following phenomena: 1) the size of a vehicle target usually includes a few pixels and 2) vehicles are usually with similar appearance which easily results in the wrong tracking within the observing region. General popular tracking methods usually focus on the representation of the target and recognize it from background which are limited in this problem. As a consequence, in this article, we propose to learn motion and background of the target in order to help the trackers recognize the target with higher accuracy. A prediction network is proposed to predict the location probability of the target in each pixel in next frame based on fully convolutional network (FCN) which is learned from previous results. In addition, a segmentation method is introduced to generate the feasible region for target in each frame and assign high probability for such a region. For quantitative comparison, we manually annotate 20 representative vehicle targets from nine satellite videos taken by JiLin-1. In addition, we also selected two public satellite video datasets for experiments. Numerous experimental results demonstrate the superior of the proposed method. Licheng Jiao, Fang Liu 0001, Lingling Li 0002, Xu Liu 0006, Jia Liu 0020 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | MFNet: A Novel GNN-Based Multi-Level Feature Network With Superpixel PriorsabstractSince the superpixel segmentation method aggregates pixels based on similarity, the boundaries of some superpixels indicate the outline of the object and the superpixels provide prerequisites for learning structural-aware features. It is worthwhile to research how to utilize these superpixel priors effectively. In this work, by constructing the graph within superpixel and the graph among superpixels, we propose a novel Multi-level Feature Network (MFNet) based on graph neural network with the above superpixel priors. In our MFNet, we learn three-level features in a hierarchical way: from pixel-level feature to superpixel-level feature, and then to image-level feature. To solve the problem that the existing methods cannot represent superpixels well, we propose a superpixel representation method based on graph neural network, which takes the graph constructed by a single superpixel as input to extract the feature of the superpixel. To reflect the versatility of our MFNet, we apply it to an image-level prediction task and a pixel-level prediction task by designing different prediction modules. An attention linear classifier prediction module is proposed for image-level prediction tasks, such as image classification. An FC-based superpixel prediction module and a Decoder-based pixel prediction module are proposed for pixel-level prediction tasks, such as salient object detection. Our MFNet achieves competitive results on a number of datasets when compared with related methods. The visualization shows that the object boundaries and outline of the saliency maps predicted by our proposed MFNet are more refined and pay more attention to details. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Puhua Chen, Xu Liu 0006, Lingling Li 0002 |
IEEE Trans. Image Process. | 5 |
| 2021 | Sparse Flow Adversarial Model For Robust Image CompressionabstractExisting learned-based image compression methods have shown impressive performance. However, most of them rely on the consistency of distribution between training images and test images, which limits the robustness of the trained model. In this paper, we propose a novel compression method called sparse flow adversarial model (SFAM). SFAM employs a deep generative framework to learn a reversible and stable mapping between image distributions, thus it can work in varied scenes for robust compression. Moreover, a sparse adversarial map is introduced into SFAM, to constrain the SFAM to generate more sparser features for efficient compression. Extensive experiments are conducted on different datasets, in which the effectiveness and robustness of the proposed method is verified. Meanwhile, SFAM is trained only once and it can work well on three different datasets, which also proves the robustness of the proposed SFAM. Shihui Zhao, Shuyuan Yang 0001, Zhi Liu 0010, Zhixi Feng, Xu Liu 0006 |
ICASSP | 5 |
| 2021 | DO-UNet, DO-LinkNet: UNet, D-LinkNet with DO-Conv for the Detection of Settlements without Electricity ChallengeabstractIn this paper, two semantic segmentation models, DO-UNet and DO-LinkNet, are presented for the detection of human settlements, and a threshold-based model is proposed to detect areas with electricity. In DO-UNet and DO-LinkNet, the conventional convolutional layer is replaced with depthwise over-parameterized convolutional layer. Also, an extra pooling operation is carried out in the last layer since the size of the input images is different from that of the labels. Depthwise over-parameterized convolutional layer enhances the convolutional layer with an additional depthwise convolution. Pooling operation can accelerate training speed, increase the receptive field in feature extraction, and reduce the requirement of network complexity. In the detection of settlements without electricity challenge track, our best F1-score on the validation set and the test set are 0.8820 and 0.8798, respectively. Ruoxian Feng, Xuanming Zhang, Jun Zhang 0045, Licheng Jiao, Xu Liu 0006, Fang Liu 0001 |
IGARSS | 6 |
| 2021 | Deep multi-level fusion network for multi-source image pixel-wise classification
Xu Liu 0006, Licheng Jiao, Lingling Li 0002, Xu Tang 0004, Yuwei Guo 0001 |
Knowl. Based Syst. | 1 |
| 2021 | Ridgelet-Nets With Speckle Reduction Regularization for SAR Image Scene ClassificationabstractWith powerful feature representations, convolutional neural networks (CNNs) have produced tremendous achievements in image classification tasks and, typically, entail millions of labeled samples to train massive parameters. However, the sample labeling of synthetic aperture radar (SAR) images is extremely difficult, especially pixelwise labels, and has, sometimes, required field trips to accomplish labeling. Moreover, the inherent speckle noise may weaken the ability of networks to extract effective features from SAR images. In this article, we address these issues by labeling a few patchwise samples and propose Ridgelet-Nets with speckle reduction regularization for SAR image scene classification by combining deep learning with multiscale geometric analysis and statistical modeling of SAR images. First, we design Ridgelet-Nets with convolutional kernels constructed by ridgelet filters to reduce the training parameters and learn more discriminative features. Then, we embed speckle reduction regularization in the Ridgelet-Nets to restrain the influence of speckle noise and smooth the classification maps, in which the prior information of SAR image statistical modeling is introduced. Finally, we propose an adaptive SAR image scene classification framework based on an extended hierarchical visual semantic model, considering the differences in the structures and spatial relationships of different regions in the SAR images, particularly large-scale and complex scenes. Experimental results on real SAR images demonstrate that the proposed framework can achieve preferable classification performance using very limited labeled samples. Xiaoxue Qian, Fang Liu 0001, Licheng Jiao, Xiangrong Zhang, Yuwei Guo 0001, Xu Liu 0006, Yuanhao Cui |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | IPGN: Interactiveness Proposal Graph Network for Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) Detection is an important task to understand how humans interact with objects. Most of the existing works treat this task as an exhaustive triplet 〈 human, verb, object 〉 classification problem. In this paper, we decompose it and propose a novel two-stage graph model to learn the knowledge of interactiveness and interaction in one network, namely, Interactiveness Proposal Graph Network (IPGN). In the first stage, we design a fully connected graph for learning the interactiveness, which distinguishes whether a pair of human and object is interactive or not. Concretely, it generates the interactiveness features to encode high-level semantic interactiveness knowledge for each pair. The class-agnostic interactiveness is a more general and simpler objective, which can be used to provide reasonable proposals for the graph construction in the second stage. In the second stage, a sparsely connected graph is constructed with all interactive pairs selected by the first stage. Specifically, we use the interactiveness knowledge to guide the message passing. By contrast with the feature similarity, it explicitly represents the connections between the nodes. Benefiting from the valid graph reasoning, the node features are well encoded for interaction learning. Experiments show that the proposed method achieves state-of-the-art performance on both V-COCO and HICO-DET datasets. Haoran Wang 0008, Licheng Jiao, Fang Liu 0001, Lingling Li 0002, Xu Liu 0006, Deyi Ji, Weihao Gan |
IEEE Trans. Image Process. | 5 |
| 2021 | C-CNN: Contourlet Convolutional Neural NetworksabstractExtracting effective features is always a challenging problem for texture classification because of the uncertainty of scales and the clutter of textural patterns. For texture classification, spectral analysis is traditionally employed in the frequency domain. Recent studies have shown the potential of convolutional neural networks (CNNs) when dealing with the texture classification task in the spatial domain. In this article, we try combining both approaches in different domains for more abundant information and proposed a novel network architecture named contourlet CNN (C-CNN). The network aims to learn sparse and effective feature representations for images. First, the contourlet transform is applied to get the spectral features from an image. Second, the spatial-spectral feature fusion strategy is designed to incorporate the spectral features into CNN architecture. Third, the statistical features are integrated into the network by the statistical feature fusion. Finally, the results are obtained by classifying the fusion features. We also investigated the behavior of the parameters in contourlet decomposition. Experiments on the widely used three texture data sets (kth-tips2-b, DTD, and CUReT) and five remote sensing data sets (UCM, WHU-RS, AID, RSSCN7, and NWPU-RESISC45) demonstrate that the proposed approach outperforms several well-known classification methods in terms of classification accuracy with fewer trainable parameters. Mengkun Liu, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Deep Adaptive Proposal Network in Optical Remote Sensing Images Objective DetectionabstractIt is difficult to distinguish complicated distribution characteristics of objects, which limits the performance of two-stage detectors in the field of optical remote sensing images object detection. In this paper, we propose a deep adaptive proposal network (DAPNet), which is a new category prior network (CPN) on the basis of the existing faster region convolutional neural network (Faster RCNN) architecture. The adaptive candidate boxes for each image is obtained by combining the candidate regions and the object number, which are generated by the fine-region proposal network (F-RPN) and the CPN respectively. These adaptive candidate boxes can satisfy the detection tasks in sparse and dense scenes. A set of experimental results verify the superiority of the proposed approach. Lingling Li 0002, Xiaohui Guo, Xu Liu 0006, Licheng Jiao, Fang Liu 0001 |
IGARSS | 4 |
| 2020 | Feature Correlation Analysis of Two-Branch Convolutional Networks for Multi-Source Image ClassificationabstractWith the development of multi-sensor imaging technology, constructing a multi-branch network model has become an important requirement for fusion decision. In the literature, many two-branch networks are proposed to interpret multisensor data and also get the satisfactory results. In this paper, we divide these models into two types and study them by mining changes in data and features. The main method used is analysis the correlation between the features of the same layer from the first branch and the second branch. The task of dual-source image classification serves as a means of experimentation. When the classification network is trained, the features of each layer are extracted for experimental analysis. Extensive experiments and analysis on the dataset IEEE_grss_dfc_2017 show that the analysis is meaningful. It is found that the correlation of the features from the low level to the high level is more and more consistent, a quantitative analysis is given in this paper. Xu Liu 0006, Licheng Jiao, Fang Liu 0001 |
IGARSS | 1 |
| 2019 | Polsar Image Classification Based on Polarimetric Scattering Coding and Sparse Support Matrix MachineabstractPOLSAR image has an advantage over optical image because it can be acquired independently of cloud cover and solar illumination. PolSAR image classification is a hot and valuable topic for the interpretation of POLSAR image. In this paper, a novel POLSAR image classification method is proposed based on polarimetric scattering coding and sparse support matrix machine. First, we transform the original POLSAR data to get a real value matrix by the polarimetric scattering coding, which is called polarimetric scattering matrix and is a sparse matrix. Second, the sparse support matrix machine is used to classify the sparse polarimetric scattering matrix and get the classification map. The combination of these two steps takes full account of the characteristics of POLSAR. The experimental results show that the proposed method can get better results and is an effective classification method. Xu Liu 0006, Licheng Jiao, Fang Liu 0001 |
IGARSS | 1 |
| 2019 | Semi-Supervised Complex-Valued GAN for Polarimetric SAR Image ClassificationabstractPolarimetric synthetic aperture radar (PolSAR) images are widely used in disaster detection and military reconnaissance and so on. However, their interpretation faces some challenges, e.g., deficiency of labeled data, inadequate utilization of data information and so on. In this paper, a complex-valued generative adversarial network (GAN) is proposed for the first time to address these issues. The complex number form of model complies with the physical mechanism of PolSAR data and in favor of utilizing and retaining amplitude and phase information of PolSAR data. GAN architecture and semi-supervised learning are combined to handle deficiency of la-beled data. GAN expands training data and semi-supervised learning is used to train network with generated, labeled and unlabeled data. Experimental results on two benchmark data sets show that our model outperforms existing state-of-the-art models, especially for conditions with fewer labeled data. Qigong Sun, Xiufang Li, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Licheng Jiao |
IGARSS | 4 |
| 2019 | Semi-Supervised Pyramid Feature Co-Training Network for Lidar Data ClassificationabstractIn the light detection and ranging (LiDAR) data classification, there are two problems affecting the LiDAR data classification performance. One is limited labeled samples in LiDAR dataset, another is confusing categories with extremely similar appearances. To address these challenges, this paper proposes a novel semi-supervised extended label algorithm (SSELA) for the data expansion, based on a novel pyramid feature co-training network (PFCTN). The proposed approach characterizes each level of feature pyramids without increasing the computational burden and time consumption, which has a benefit for learning complementary features of confusing categories. Experimental results demonstrate that the PFCTN alleviates the decreasing accuracy phenomenon at different scales of similar scenes, and achieves the promising results with limited labeled LiDAR dataset. Zexin Wang 0003, Haoran Wang 0008, Licheng Jiao, Xu Liu 0006 |
IGARSS | 4 |
| 2019 | Multi-Scale Feature Fusion Network for Object Detection in VHR Optical Remote Sensing ImagesabstractIn this paper, we propose a multi-scale feature fusion network (MS-FF Net) based on convolutional neural network (CNN) to deal with object detection in VHR images. In CNN, the low-level layers contain rich detail information and the high-level layers contain rich semantic information. Inspired by the idea of feature fusion, we propose an additional multi-scale feature fusion layer (MFL) to fuse the information between detail and semantic features. Then both large and small objects are considered by this network. Moreover, the network architecture and training strategies are designed to improve performance. Experiments on NWPU VHR-10 dataset demonstrate that the method with MFLs achieves significant improvement and outperforms compared methods in terms of mean average precision. Specially, the detection precision of airplane, baseball diamond, basketball court, ground track field and harbor categories exceeds 90% which is much higher than that of compared methods. Licheng Jiao, Xu Liu 0006, Jia Liu 0020 |
IGARSS | 3 |
| 2019 | Adaptive Multiscale Deep Fusion Residual Network for Remote Sensing Image ClassificationabstractWith the development of remote sensing imaging technology, remote sensing images with high-resolution and complex structure can be acquired easily. The classification of remote sensing images is always a hot and challenging problem. In order to improve the performance of remote sensing image classification, we propose an adaptive multiscale deep fusion residual network (AMDF-ResNet). The AMDF-ResNet consists of a backbone network and a fusion network. The backbone network including several residual blocks generates multiscale hierarchy features, which contain semantic information from low to high levels. In the fusion network, the adaptive feature fusion module proposed can emphasize useful information and suppress useless information by learning the weights, which represent the importance of the features. The AMDF-ResNet can make full use of the multiscale hierarchy features and the extracted feature is discriminative. In addition, we propose a samples selection method named important samples selection strategy (ISSS). Based on superpixels segmentation result, gradient information and spatial distribution are used as two references to determine the selection numbers and select samples. Compared with the random selection strategy, training samples selected by ISSS are more representative and diverse. The experimental results on four data sets demonstrate that the AMDF-ResNet and ISSS are effective. Lingling Li 0002, Hao Zhu 0009, Xu Liu 0006, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | Polarimetric Convolutional Network for PolSAR Image ClassificationabstractThe approaches for analyzing the polarimetric scattering matrix of polarimetric synthetic aperture radar (PolSAR) data have always been the focus of PolSAR image classification. Generally, the polarization coherent matrix and the covariance matrix obtained by the polarimetric scattering matrix are used as the main research object to extract features. In this paper, we focus on the original polarimetric scattering matrix and propose a polarimetric scattering coding way to deal with polarimetric scattering matrix and obtain a close complete feature. This encoding mode can also maintain polarimetric information of scattering matrix completely. At the same time, in view of this encoding way, we design a corresponding classification algorithm based on the convolution network to combine this feature. Based on the polarimetric scattering coding and convolution neural network, the polarimetric convolutional network is proposed to classify PolSAR images by making full use of polarimetric information. We perform the experiments on the PolSAR images acquired by AIRSAR and RADARSAT-2 to verify the proposed method. The experimental results demonstrate that the proposed method get better results and has huge potential for PolSAR data classification. Source code for polarimetric scattering coding is available at https://github.com/liuxuvip/Polarimetric-Scattering-Coding. Xu Liu 0006, Licheng Jiao, Xu Tang 0004, Qigong Sun |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Rate-mixed HEVC Tile based 360 Video Streaming SystemabstractRecently, tile-based viewport adaptation is a popular method for 360 video streaming. Our demonstration adopts a rate-mixed transmission approach utilizing a VR adaptation agent at the server end for viewport-based streaming, which is client-compatible and can be scalable to different users. The FOV prediction is applied to improve the viewing experience. The feasibility of our system to head-mounted displays is verified, which can reduce bandwidth consumption by up to 36%. Xu Liu 0006, Rong Xie 0004, Li Song 0001 |
VCIP | 2 |
| 2018 | Deep Multiple Instance Learning-Based Spatial-Spectral Classification for PAN and MS ImageryabstractPanchromatic (PAN) and multispectral (MS) imagery classification is one of the hottest topics in the field of remote sensing. In recent years, deep learning techniques have been widely applied in many areas of image processing. In this paper, an end-to-end learning framework based on deep multiple instance learning (DMIL) is proposed for MS and PAN images’ classification using the joint spectral and spatial information based on feature fusion. There are two instances in the proposed framework: one instance is used to capture the spatial information of PAN and the other is used to describe the spectral information of MS. The features obtained by the two instances are concatenated directly, which can be treated as simple fusion features. To fully fuse the spatial–spectral information for further classification, the simple fusion features are fed into a fusion network with three fully connected layers to learn the high-level fusion features. Classification experiments carried out on four different airborne MS and PAN images indicate that the classifier provides feasible and efficient solution. It demonstrates that DMIL performs better than using a convolutional neural network and a stacked autoencoder network separately. In addition, this paper shows that the DMIL model can learn and fuse spectral and spatial information effectively, and has huge potential for MS and PAN imagery classification. Xu Liu 0006, Licheng Jiao, Jiaqi Zhao 0001, Jin Zhao 0002, Fang Liu 0001, Shuyuan Yang 0001, Xu Tang 0004 |
IEEE Trans. Geosci. Remote. Sens. | 1 |