EDBT 2026 Demo / reviewers in the wild / expert
Puhua Chen
dblp:179/2779
· DBLP profile ↗
88ranked-venue papers
2as first author
81since 2021 · last 2026
0000-0001-5472-1426ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 1 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 1 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 23 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evolving Semantic Propagation for Aerial Semantic 3D Gaussian SplattingabstractSemantic understanding of large-scale aerial scenes represents a critical challenge in 3D computer vision, hindered by the prohibitive cost of dense annotation. This paper introduces EvoPropGS, a novel approach for the semantic segmentation of 3D Gaussian Splatting models that requires only minimal supervision. Our core insight is to leverage the inherent structural repetitions within aerial environments to propagate semantic information from a sparse set of annotations across the entire 3D scene. Our approach constructs a prompt library by pairing SAM-generated mask candidates with DINOv2 feature embeddings from annotated views. For unannotated regions, we generate pseudo-labels by matching region proposals with these featured prompts via cosine similarity. We then formulate optimal prompt selection as a discrete optimization problem solved via evolutionary search, guided by our novel fitness function that evaluates both 3D consistency and 2D semantic coherence. Extensive experiments demonstrate that EvoPropGS achieves accurate segmentation with only 2 percent annotated pixels. Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Licheng Jiao, Puhua Chen, Wenping Ma 0001, Shuyuan Yang 0001 |
AAAI | 6 |
| 2026 | HTTrack: Learning to Perceive Targets via Historical Trajectories in Satellite Video TrackingabstractIn recent years, the rapid progress of deep learning has driven notable advancements in satellite video tracking, a critical task for applications such as environmental monitoring, disaster management, and defense. Despite these strides, existing approaches remain constrained by their inability to handle dynamic challenges, such as target appearance variations, complex motion patterns, and occlusions. Traditional methods often suffer from static template matching or overly complex update mechanisms, compromising their robustness and practicality in real-world scenarios. To address these limitations, we propose a paradigm shift in satellite video tracking by integrating historical trajectory knowledge with visual features. This fusion enhances the tracker's perceptual understanding of targets over time, enabling more adaptive and resilient tracking. By aligning spatial, temporal, and cross-modal information, our approach effectively bridges the gap between fragmented observations and coherent tracking performance, even under challenging conditions like small target detection and cluttered backgrounds. Extensive experiments conducted on multiple satellite video tracking benchmarks demonstrate the superiority of our method, with HTTrack achieving success rates of 51.5% on SV248S, 52.9% on SatSOT, and 32.6% on VISO, significantly outperforming state-of-the-art trackers and marking a step forward in achieving robust, accurate, and scalable satellite video tracking. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
AAAI | 8 |
| 2026 | Semantic Feature Purification for Adversarially-Aware RGB-T TrackingabstractRGB-T tracking is increasingly deployed in safety-critical applications such as autonomous driving, surveillance, and rescue robotics, where tracking reliability is essential under adverse conditions. Although the fusion of RGB and thermal infrared (TIR) modalities offers improved robustness in low-light and occluded scenes, recent findings show that RGB-T trackers remain highly susceptible to subtle input perturbations, human-imperceptible modifications that exploit cross-modal inconsistencies to mislead tracking outputs. In real-world scenarios, such perturbations can arise from sensor spoofing, infrared camouflage, or physical-world attacks, posing serious risks to operational safety. To address this, we propose SFPT, a Semantic Feature Purification framework that enhances RGB-T tracking at the representation level. Rather than filtering corrupted inputs at the pixel level, SFPT introduces task-specific semantic anchors into the feature space to reinforce perturbation-invariant cues. These anchors are derived from descriptive language, interact with visual features to purify representations. To further suppress modality-specific interference, we design an Adaptive Perturbation-Guided Cross-Modal Fusion (APG-CMF) module, which leverages language and visual signals to estimate reliability and dynamically reweight cross-modal features, ensuring robust fusion under perturbation conditions. Extensive experiments under diverse perturbation conditions validate the effectiveness of our approach. Notably, SFPT maintains performance comparable to clean settings even when subjected to perturbations of strength 1/255 and 4/255, demonstrating strong resilience to real-world interference. Jiahao Wang 0002, Fang Liu 0001, Hao Wang 0211, Shuo Li 0010, Puhua Chen |
AAAI | 6 |
| 2026 | Enhancing few-shot segmentation via mask combination learning
Shuo Li 0010, Fang Liu 0034, Licheng Jiao, Xuejian Gou, Lingling Li 0002, Xu Liu 0006, Puhua Chen |
Neurocomputing | 7 |
| 2026 | Privacy-preserving video anomaly detection via federated learning
Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Yanbiao Ma, Qianyue Bao, Lingling Li 0002, Puhua Chen |
Knowl. Based Syst. | 11 |
| 2026 | Causality-inspired learning semantic segmentation in unseen domain
Pei He, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Ronghua Shang, Yuwei Guo 0001, Puhua Chen, Shuyuan Yang 0001 |
Pattern Recognit. | 8 |
| 2026 | Concept-Aware Learning for Weakly Supervised Video Anomaly Detection
Shuo Li 0010, Fang Liu 0034, Licheng Jiao, Jiahao Wang 0002, Xu Liu 0006, Lingling Li 0002, Puhua Chen |
Pattern Recognit. | 8 |
| 2026 | VCGPrompt: Visual Concept Graph-Aware Prompt Learning for Vision-Language Models
Mengjia Wang, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
Pattern Recognit. | 6 |
| 2026 | Vision-by-prompt: Context-aware dual prompts for composed video retrieval
Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
Pattern Recognit. | 7 |
| 2026 | TFBTrack: Target-Aware Foreground-Background Modeling for vision-language tracking
Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
Pattern Recognit. | 7 |
| 2026 | Language-guided modulation-update for semi-supervised semantic segmentation
Libo Yan, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Jiahao Wang 0002, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Xuejian Gou |
Pattern Recognit. | 7 |
| 2026 | ERFC: Energy-Aware Reinforcement Feedback Calibration for Zero-Shot CaptioningabstractZero-shot captioning aims to generate descriptive captions for unseen image and video data by leveraging the potential of visual language models (VLMs) and language models (LMs) without requiring task-specific training. It has emerged as a critical task, but its performance is often hindered by the inherent gap between the training distribution and unseen test data. The fundamental challenge lies in the model’s strong dependence on the marginal distribution of the training data, which leads to biased predictions when handling test samples. To address this issue, we propose an Energy-aware Reinforcement Feedback Calibration (ERFC) framework to calibrate the distribution and predictions of caption models from a novel energy perspective. The calibration process of ERFC is divided into two key components: 1) We first construct an Energy Stabilizer (ES) based on the caption model, where energy is considered a measure of the affinity between the input sample and the model’s learned distribution. ES iteratively adjusts the embedding features of the input sample using Langevin Dynamics, reducing its energy to implicitly align the model’s distribution with the unseen target domain. 2) We deploy a Reinforcement Calibrator (RC) to refine and calibrate the generated captions through a reward-feedback mechanism. RC leverages the expert CLIP model as a reward signal to assess the quality of the generated captions and employs the policy gradient algorithm to reward or penalize the model, thereby improving its performance. By iteratively combining energy-based optimization and reward-driven calibration, ERFC achieves superior zero-shot generalization capabilities, as demonstrated on image benchmarks such as MSCOCO, Flickr30K, and NoCaps, as well as video benchmarks such as MSR-VTT and MSVD. Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | Image Singularity Scattering Representation Learning ClassificationabstractThe multi-scale geometric analysis is a great representation tool. It can be used to improve the feature representation and learning process of deep networks. In addition to extracting features, the multi-scale geometric prior knowledge can also be used for the structure improvement of deep networks. In this paper, we propose a multi-scale scattering representation learning network, abbreviated as MSRLN, for image classification tasks. The exploration of structure improvement can be made with multi-scale scattering operations. In this way, the better singularity representation learning process for networks can be achieved. Firstly, the filter banks and multi-scale scattering operator are introduced for non-linear and singularity representation. Secondly, the novel multi-scale scattering representation learning network structure is designed. The scaling- wise scattering process is deployed in the shallow layer as a non-linear layer. This structure essentially supplements deep networks with geometric prior knowledge. It can further improve the non-linear activation and singularity representation process. Thirdly, we put forward the multi-stage scattering representation strategy and the prior knowledge weakening mechanism. With flexible scaling factors and learning rates, the stepwise approximation and learning process of networks can be achieved. In sum, MSRLN is a kind of structural innovative, and the scattering singularity representation structure can be extended to other backbones or tasks. Extensive experimental results show that MSRLN can achieve better image classification accuracy. Finally, necessary convergence, insight, and adaptability analyses are provided in evaluation experiments. Jie Gao 0013, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Puhua Chen, Yuwei Guo 0001, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2026 | Learning to Prompt With Refining Text Knowledge for Zero-Shot Video Action RecognitionabstractFoundational vision-language models (VLMs) like CLIP are redefining the vision domain with their exceptional generalization capabilities. Prompt-based learning methods adapt pre-trained VLMs to video action recognition tasks using task-specific learnable text tokens. However, these tokens often struggle to generalize to unseen categories, as they tend to forget general textual knowledge. To address this, we construct knowledge prompts composed of handcrafted and descriptive prompts and introduce a novel knowledge-guided context mapping to enhance the generalization of learnable prompts to unseen categories. This approach mitigates the forgetting of fundamental knowledge by reducing the discrepancy between learnable prompts and knowledge prompts while simultaneously allowing the prompts to extract rich contextual knowledge from LLM data. Then, incorporating the knowledge-guided context mapping into the contrastive loss enables zero-shot transfer of prompts to new categories and data, providing discriminative prompts for both seen and unseen tasks. In addition, we propose an advanced temporal aggregation method that refines uniform mean pooling by incorporating frame-level textual relevance scoring. Extensive evaluations on multiple benchmarks demonstrate that learning to prompt with refining text knowledge is an effective quick-tuning method, achieving superior sample generalization performance without increasing training parameters. Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
IEEE Trans. Multim. | 7 |
| 2026 | Adaptive Multi-Modal Visual Tracking With Dynamic Semantic PromptsabstractRGB-based object tracking is a fundamental task in computer vision, aiming to identify, locate, and continuously track objects of interest across sequential video frames. Despite the significant advancements in the performance of traditional RGB trackers, they still face challenges in maintaining accuracy and robustness in the presence of complex backgrounds, occlusions, and rapid movements. To tackle these challenges, combining visual auxiliary modalities has gained significant attention. Beyond this, integrating natural language information offers additional advantages by providing high-level semantic context, enhancing robustness, and clarifying target priorities, further elevating tracker performance. This work proposes theAdaptiveMulti-modalVisual Tracking with Dynamic Semantic Prompts (AMVTrack) tracker, which efficiently incorporates image descriptions and avoids text dependency during tracking to improve flexibility and adaptability. AMVTrack significantly reduces computational resource consumption by freezing the parameters of the image encoder, text encoder, and Box Head and only optimizing a few learnable prompt parameters. Additionally, we introduce the Adaptive Dynamic Semantic Prompt Generator (ADSPG), which dynamically generates semantic prompts based on visual features, and theVisual-LanguageFusionAdaptation (V-L FA) method, which integrates multi-modal features to ensure consistency and complementarity of information. Additionally, we partition the Image Encoder to conduct an in-depth investigation into the relationship between the importance of features across different depth and width regions. Experimental results demonstrate that AMVTrack achieves significant performance improvements on multiple benchmark datasets, proving its effectiveness and robustness in complex scenarios. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
IEEE Trans. Multim. | 7 |
| 2026 | Adaptive Visual Prompting for Effective Satellite Video TrackingabstractSatellite video tracking presents significant challenges due to unpredictable target variations, environmental disturbances, and occlusions. Existing approaches either rely on auxiliary modalities or require full fine-tuning of foundation models, resulting in excessive parameter sensitivity and poor generalization. Meanwhile, conventional prompt-based tuning only updates parameters at a single location, limiting its ability to adapt to complex appearance changes. To address these limitations, we propose Adaptive Visual Prompting for Effective Satellite Video Tracking (AVPTrack). Unlike conventional prompts, introduced Super Prompts dynamically refine the original template at multiple distinct positions. This multi-location adaptation allows for fine-grained representation learning, enabling the tracker to better capture target variations and resist environmental disturbances. Additionally, Dynamic Templates are introduced to mitigate tracking failures in highly challenging scenarios, such as occlusions and background clutter, ensuring robust target localization. Furthermore, the Template Selection Adapter (TSA) selects the most relevant templates in real-time, enhancing tracking efficiency. These components are optimized during training while keeping other parameters frozen, ensuring parameter efficiency. We also investigate the relationship between fine-tuning proportions and learning rates to optimize model performance. Extensive evaluations on the SV248S, SatSOT, and VISO datasets demonstrate the superior adaptability and robustness of AVPTrack compared to existing methods. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Yanbiao Ma, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Mengjia Wang |
IEEE Trans. Multim. | 8 |
| 2026 | PromptVAD: Abnormal Prompt via Vision-Language ModelabstractWeakly supervised video anomaly detection (WSVAD) aims at predicting frame-level anomaly scores by modeling training videos with video-level annotations. The category names of abnormal events contain high-level knowledge abstracted by humans about abnormalities, which is of great help in identifying abnormal events. To utilize the knowledge implicit in category names, based on the visual-language pretraining model, we introduce a learnable abnormal prompt from three aspects: learnable domain prompt, learnable category prompt, and nonlearnable category definition prompt. Based on the learnable abnormal prompt, we propose a novel fine-grained WSVAD method: PromptVAD, which exploits a learnable abnormal prompt to reduce the semantic gap between visual images and anomaly categories. Through a similarity measure and our proposed coarse-grained two-class prompt module, our PromptVAD jointly learns coarse-grained and fine-grained VAD. Extensive experimental results on the ShanghaiTech, University of Central Florida (UCF)-Crime, and XD-Violence datasets show that our method achieves state-of-the-art performance. Specifically, our method achieves an area under the curve (AUC) of 88.62% on the UCF-Crime dataset. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Zehua Hao, Jiahao Wang 0002, Lingling Li 0002, Xu Liu 0006, Puhua Chen |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2025 | Logits DeConfusion with CLIP for Few-Shot LearningabstractWith its powerful visual-language alignment capability, CLIP performs well in zero-shot and few-shot learning tasks. However, we found in experiments that CLIP’s logits suffer from serious inter-class confusion problems in down-stream tasks, and the ambiguity between categories seriously affects the accuracy. To address this challenge, we propose a novel method called Logits DeConfusion, which effectively learns and eliminates inter-class confusion in logits by combining our Multi-level Adapter Fusion (MAF) module with our Inter-Class Deconfusion (ICD) module. Our MAF extracts features from different levels and fuses them uniformly to enhance feature representation. Our ICD learnably eliminates inter-class confusion in logits with a residual structure. Experimental results show that our method can significantly improve the classification performance and alleviate the inter-class confusion problem. The code is available at https://github.com/LiShuo1001/LDC. Shuo Li 0010, Fang Liu 0001, Zehua Hao, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
CVPR | 7 |
| 2025 | Knowledge-Guided Part Segmentation
Xuejian Gou, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, Hao Wang 0211, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
ICCV | 8 |
| 2025 | Hierarchical Variational Test-Time Prompt Generation for Zero-Shot Generalization
Zhaoyang Wu, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, LiXu Liu, Puhua Chen, Wenping Ma 0001 |
ICCV | 7 |
| 2025 | FA3T: Feature-Aware Adversarial Attacks for Multi-modal TrackingabstractMulti-modal visual tracking leverages complementary sensor information to enhance robustness under challenging conditions. However, the security of multi-modal tracking systems remains largely unexplored. Existing attacks primarily target single-modal trackers or independently disrupt each modality, failing to exploit the inherent feature interactions and fusion mechanisms that define multi-modal tracking. As a result, these methods exhibit limited attack effectiveness and fail to assess multi-modal tracking systems' vulnerabilities accurately. Understanding these security risks is crucial, as adversarial threats could lead to severe failures in safety-critical applications. To address these challenges, a feature-aware adversarial attack, termed FA3T is proposed. It is designed to explicitly disrupt feature extraction and cross-modal alignment, thereby weakening the fusion process that multi-modal trackers rely on. To achieve this, a Frequency-Spatial Feature Separation (FSFS) module is constructed to perturb feature representations at multiple levels, weakening the modality-complementary advantages of multi-modal tracking. Furthermore, a Target Confusion Attack (TCA) module is devised to manipulate the target-background-template relationships, making it increasingly difficult for the tracker to distinguish the true target, significantly impairing tracking performance. Extensive experiments on five benchmark datasets (i.e., LasHeR, RGBT234, DepthTrack, VOT-RGBD2022, VisEvent) across three different modalities (RGB-T, RGB-D, and RGB-E) demonstrate that our attack substantially degrades state-of-the-art multi-modal trackers, exposing their susceptibility to adversarial threats. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
ACM Multimedia | 7 |
| 2025 | Imagining Vision From Language for Few-Shot Class-Incremental Learning
Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Yanbiao Ma, Puhua Chen, Lingling Li 0002, Xu Liu 0006, Xuejian Gou |
ACM Multimedia | 8 |
| 2025 | Preserving text space integrity for robust compositional zero-shot learning via mixture of pretrained experts
Zehua Hao, Fang Liu 0001, Licheng Jiao, Yaoyang Du, Shuo Li 0010, Hao Wang 0211, Pengfang Li, Xu Liu 0006, Puhua Chen |
Neurocomputing | 9 |
| 2025 | LLM Knowledge-Driven Target Prototype Learning for Few-Shot Segmentation
Pengfang Li, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Xu Liu 0006, Puhua Chen, Lingling Li 0002, Zehua Hao |
Knowl. Based Syst. | 6 |
| 2025 | Visual-Language Scene-Relation-Aware Zero-Shot CaptionerabstractZero-shot image captioning can harness the knowledge of pre-trained visual language models (VLMs) and language models (LMs) to generate captions for target domain images without paired sample training. Existing methods attempt to establish high-quality connections between visual and textual modalities in text-only pre-training tasks. These methods can be divided into two perspectives: sentence-level and entity-level. Although they achieve effective performance on some metrics, they suffer from hallucinations due to biased associations during training. In this paper, we propose a scene-relation-level pre-training task by considering relations as more valuable modal connection bridges. Based on this, we construct a novel Visual-Language Scene Relation Aware Captioner (SRACap), which expands the ability to predict scene relations while generating captions for images. In addition, SRACap possesses excellent cross-domain zero-shot generalization capability, which is driven by a well-designed scene reinforcement switching pipeline. We introduce a scene policy network to dynamically crop salient regions from images and feed them into a language model to generate captions. We integrate multiple expert CLIP models to form a mixture-of-rewards module (MoR) as a reward source, and deeply optimized SRACap through the policy gradient algorithm in the zero-shot inference stage. With the iteration of scene reinforcement switching, SRACap can gradually refine the generated caption details while maintaining high semantic consistency across visual-linguistic modalities. We conduct extensive experiments on multiple standard image captioning benchmarks, showing that SRACap can accurately understand scene structures and generate high-quality text, significantly outperforming other zero-shot inference methods. Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | Unveiling and Mitigating Generalized Biases of DNNs Through the Intrinsic Dimensions of Perceptual ManifoldsabstractBuilding fair deep neural networks (DNNs) is a crucial step towards achieving trustworthy artificial intelligence. Delving into deeper factors that affect the fairness of DNNs is paramount and serves as the foundation for mitigating model biases. However, current methods are limited in accurately predicting DNN biases, relying solely on the number of training samples and lacking more precise measurement tools. Here, we establish a geometric perspective for analyzing the fairness of DNNs, comprehensively exploring how DNNs internally shape the intrinsic geometric characteristics of datasets-the intrinsic dimensions (IDs) of perceptual manifolds, and the impact of IDs on the fairness of DNNs. Based on multiple findings, we propose Intrinsic Dimension Regularization (IDR), which enhances the fairness and performance of models by promoting the learning of concise and ID-balanced class perceptual manifolds. In various image recognition benchmark tests, IDR significantly mitigates model bias while improving its performance. Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Lingling Li 0002, Wenping Ma 0001, Shuyuan Yang 0001, Xu Liu 0006, Puhua Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | Predicting and Enhancing the Fairness of DNNs With the Curvature of Perceptual ManifoldsabstractTo address the challenges of long-tailed classification, researchers have proposed several approaches to reduce model bias, most of which assume that classes with few samples are weak classes. However, recent studies have shown that tail classes are not always hard to learn, and model bias has been observed on sample-balanced datasets, suggesting the existence of other factors that affect model bias. In this work, we first establish a geometric perspective for analyzing model fairness and then systematically propose a series of geometric measurements for perceptual manifolds in deep neural networks. Subsequently, we comprehensively explore the effect of the geometric characteristics of perceptual manifolds on classification difficulty and how learning shapes the geometric characteristics of perceptual manifolds. An unanticipated finding is that the correlation between the class accuracy and the separation degree of perceptual manifolds gradually decreases during training, while the negative correlation with the curvature gradually increases, implying that curvature imbalance leads to model bias. We thoroughly validate this finding across multiple networks and datasets, providing a solid experimental foundation for future research. We also investigate the convergence consistency between the loss function and curvature imbalance, demonstrating the lack of curvature constraints in existing optimization objectives. Building upon these observations, we propose curvature regularization to facilitate the model to learn curvature-balanced and flatter perceptual manifolds. Evaluations on multiple long-tailed and non-long-tailed datasets show the excellent performance and exciting generality of our approach, especially in achieving significant performance improvements based on current state-of-the-art techniques. Our work opens up a geometric analysis perspective on model bias and reminds researchers to pay attention to model bias on non-long-tailed and even sample-balanced datasets. Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Maoji Wen, Lingling Li 0002, Wenping Ma 0001, Shuyuan Yang 0001, Xu Liu 0006, Puhua Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2025 | VLPA-CLIP: Video Language Prompting and Adapting CLIP for efficient video action recognition
Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
Pattern Recognit. | 7 |
| 2025 | Heterogeneous Dual-Branch Emotional Consistency Network for Facial Expression RecognitionabstractDue to labeling subjectivity, label noises have become a critical issue that is addressed in facial expression recognition. From the view of human visual perception, the facial exhibited emotion characteristic should be unaltered corresponding to its truth expression, rather than the noise label, whereas most methods ignore the emotion consistency during FER, especially from different networks. Based on this, we propose a new FER method based heterogeneous dual-branch emotional consistency constrains, to prevent the model from memorizing noise samples based on features associated with noisy labels. In the proposed method, the emotion consistency from spatial transformation and heterogeneous networks are simultaneously considered to guide the model to perceive the overall visual features of expressions. Meanwhile, the confidence of the given label is evaluated based on emotional attention maps of original and transformed images, which effectively enhances the classification reliability of two branches to alleviate the negative effect of noisy labels in the learning process. Additionally, the weighted ensemble strategy is used to unify two branches. Experimental results illustrate that the proposed method achieves better performance than the state-of-the-art methods for 10%, 20% and 30% label noises. Shasha Mao, Puhua Chen |
IEEE Signal Process. Lett. | 4 |
| 2025 | Prompt-Based Concept Learning for Few-Shot Class-Incremental LearningabstractFew-Shot Class-Incremental Learning (FSCIL) faces a huge stability-plasticity challenge due to continuously learning knowledge from new classes with a small number of training samples without forgetting the knowledge of previously seen old classes. To alleviate this challenge, we propose a novel method called Prompt-based Concept Learning (PCL) for FSCIL, which generalizes conceptual knowledge learned from old classes to new classes by simulating human learning capabilities. In our PCL, in the base session, we simultaneously learn common basic concepts from the training data and the class-concept weight of each class in a prompt learning manner, and in each incremental session, class-concept weights between new classes and previously learned basic concepts are learned to achieve incremental learning. Furthermore, in order to avoid catastrophic forgetting, we propose a distribution estimation module to retain feature distributions of previously seen classes and a data replay module to randomly sample features of previously seen classes in incremental sessions. We verify the effectiveness of our PCL on widely used benchmarks, such as miniImageNet, CIFAR-100, and CUB-200. Experimental results show that our PCL achieves competitive results compared with other state-of-the-art methods, especially we achieve an average accuracy of 94.02% across all sessions on the miniImageNet benchmark. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Change Knowledge-Guided Vision-Language Remote Sensing Change DetectionabstractRemote sensing image change detection plays a critical role in applications like video surveillance and geographic information systems. However, existing binary and semantic change detection methods often rely solely on visual information, neglecting language information, which limits interpretability and the ability to provide specific change details. This work proposes the Change Knowledge-Guided Vision-Language Remote Sensing Change Detection (CKCD) method to address these limitations. By introducing change knowledge as language information, CKCD enhances semantic understanding and change detail representation. A Cross-Modal Affinity (CMA) module is designed to effectively fuse visual and textual features, improving information complementarity and fusion coherence. CKCD further enhances data utilization efficiency by merging change area detection and change category information into a single output through endto- end learning. This design reduces redundant data representations and simplifies the detection process, leading to a more compact and efficient use of the input data without requiring additional branches or multiple output heads. Experimental results demonstrate consistent performance improvements over traditional methods across multiple change detection datasets. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Temporal-Feedback Self-Training for Semi-Supervised Object Detection in Remote Sensing ImagesabstractAlthough modern Remote Sensing Object Detection (RSOD) methods have achieved advanced performance, they heavily rely on a large amount of annotated data. This paper explores semi-supervised RSOD to mitigate annotation costs, leveraging recent extensive research in generic Semi-Supervised Object Detection (SSOD) based on the self-training paradigm. Current SSOD methods encounter challenges in adapting to remote sensing images due to the complexity and variability of RSIs. Two key issues remain underexplored: the noise in pseudo-labels caused by model instability and the difficulty in distinguishing similar categories. This paper introduces the Temporal-Feedback Self-Training (TST) framework, a novel approach to tackle these challenges in semi-supervised RSOD. TST consists of two components: Temporal Consistency Based Pseudo-labels Certainty Estimation (TCE) and Temporal Self-Feedback Feature Refinement (TSF). TCE addresses pseudo-label noise during training by evaluating the stability of pseudo-label classification and localization over time series to assess the quality of pseudo-labels. On the other hand, TSF enhances pseudo-label quality by dynamically identifying the models confusing categories as feedback for feature refinement. Both components facilitate the progression of the self-training-based RSOD during training. We conducted extensive experiments on two challenging public datasets, DOTA and DIOR. The results demonstrate that the proposed TST and TCE components significantly improve the baseline models performance, surpassing the state-of-the-art generic SSOD method. This suggests that our approach is more effective than generic SSOD methods in addressing the challenges posed by remote sensing images. Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Xu Tang 0004, Puhua Chen, Huiyu Zhou 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | LGSNet: Local-Global Semantics Learning Object DetectionabstractSelf-attention learns capturing the long-range dependencies between embeddings (e.g., image pixels). However, the memory overhead and computation cost are prohibitive due to being quadratic in term of the spatial resolution. The structure analysis reveals two crucial roles in the attention: the correlation-based dependency structure and feature normalization. In this work, an efficacious Local-Global Semantics (LGS) module is proposed to alleviate the above issues by modeling the local semantic aggregation and global semantic interaction. Our LGS module contains a group convolution and an Efficient Global Semantic Attention (EGSA). Firstly, the group convolution aggregates local semantics. Secondly, considering a feature map as a sequence of 2-D channel representations, EGSA formulates a general model for the global semantic interaction. The linear correlation is computed between global semantics. LGS has the linear memory overhead and computation cost in term of the spatial resolution. The LGS module can be smoothly incorporated into object detection frameworks. The experiment results verify its effectiveness on two popular detection datasets: the MS COCO and PASCAL VOC. Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Puhua Chen |
IEEE Trans. Multim. | 6 |
| 2025 | Semantic-Aware Wavelet Transformer for Pyramid Learning Object DetectionabstractTransformer displays the impressive capabilities on vision tasks. The built-in self-attention retains the quadratic computation burden in respect of the spatial resolution of image features. The traditional downsampling (e.g., average pooling) can reduce the resolution. Nonetheless, it may suffer from the dropping of detailed information. In this work, we propose an Efficient Wavelet Attention (EWA), which injects the wavelet transform and a Mean GELU (MGELU) function. Firstly, the wavelet transform enables the detailed information to participate in the efficient interaction modeling. Secondly, MGELU regards the statistical mean as reference and loosely passes the high relative responses. Building upon EWA, we present an effective Semantic-aware Wavelet Transformer (SWFormer), which is then employed for pyramid learning, including CNN feature hierarchy or Region of Interest (RoI) features. For the feature hierarchy, a Pyramid SWFormer (PSWFormer) incorporates SWFormer at each level to fit the bidirectional features. For RoIs, a Recognition-Localization SWFormer (RLSWFormer) is inserted into the head to fit their features from all levels. The effectiveness of our SWFormer is displayed experimentally on the MS COCO detection dataset and the Pascal VOC dataset. When exploiting Swin-small backbone, our SWFormer-based method acquires AP of 52.1 in the single-scale evaluation on the COCO test-dev set. This work will have the codes athttps://github.com/TimeIsFuture/Dt2_SWFormer. Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Puhua Chen |
IEEE Trans. Multim. | 6 |
| 2025 | Tracking Like Human: Dynamic Scene Learning Reasoning Tracker in Satellite VideosabstractIn satellite video object tracking, the individual frame analysis method is usually used for target localization, ignoring informative cues of the dynamic scene. Temporal information could contribute to identifying the target from distractors. In this work, a novel dynamic scene learning reasoning tracker is proposed for satellite videos, which reasons over temporal dynamic information to derive the target location. It is inspired by the tracking pattern through human perception and reasoning. First, static-dynamic united analysis is designed to construct dynamic scenes by concatenating the static searching results along the temporal dimension. Second, the information of each response object is aggregated by wavelet transforms. Meanwhile, these scenes are projected into low-frequency and high-frequency subspaces, which could imitate different levels of perceptions of humans for scenes. Third, an object-aware reasoning transformer is proposed to utilize the temporal dynamics of input response objects. In each subspace, it models the mutual interactions between dynamic objects and further learns the intrinsic property of each object for target reasoning. Finally, to obtain the current reasoning result, inverse wavelet transforms are utilized to integrate the results of low-frequency and high-frequency subspaces. The effectiveness of the proposed method is validated on three public satellite video datasets, including SV248S, SkySat, and VISO. Qualitative and quantitative experimental results show that the proposed tracker outperforms 22 popular approaches in seven challenging tracking satellite scenarios. Licheng Jiao, Yangyang Li 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Heterogeneous Riemannian Few-Shot Learning NetworkabstractHow to learn and accurately distinguish new concepts from few samples, as humans do, is a long-standing concern in artificial intelligence (AI). Studies in brain science and neuroscience have shown that human brain perception is based on nonlinear manifolds, and high-dimensional manifolds can facilitate concept learning in neural circuits. Based on this inspiration, in this paper, we propose a heterogeneous Riemannian few-shot learning network (HRFL-Net), which is the first few-shot learning method to perform end-to-end deep learning on heterogeneous Riemannian manifolds. Specifically, to enhance the geometric invariance of the image representation, the image features are projected into three heterogeneous Riemannian manifold spaces. Then, the implicit Riemannian kernel function maps the manifolds to the separable high-dimensional reproducing Hilbert space. It is assumed that the embedded kernel features of the complementary manifolds are mapped to the same common subspace. Thus, a novel neural network-based Riemannian metric learning method is designed to solve the subspace feature vectors by imposing orthogonal normalized projection, which overcomes the data extension limitation of the Riemannian metric. Finally, with the optimization objective of increasing the interclass distance and decreasing the intraclass distance in Hilbert space, the HRFL-Net is trained with end-to-end stochastic optimization, and the optimal aggregation subspace is learned during the gradient descent process. Thus, the proposed HRFL-Net can be easily generalized to challenging nonconvex data. The evaluation of four public datasets shows that the proposed HRFL-Net has significant superiority and also achieves competitive results compared with the state-of-the-art methods. Jie Chen 0098, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Yuwei Guo 0001, Puhua Chen, Wenping Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | A Spatial-Spectral Relation-Guided Fusion Network for Multisource Optical RS Image ClassificationabstractMultisource optical remote sensing (RS) image classification has obtained extensive research interest with demonstrated superiority. Existing approaches mainly improve classification performance by exploiting complementary information from multisource data. However, these approaches are insufficient in effectively extracting data features and utilizing correlations of multisource optical RS images. For this purpose, this article proposes a generalized spatial-spectral relation-guided fusion network (S2RGF-Net) for multisource optical RS image classification. First, we elaborate on spatial- and spectral-domain-specific feature encoders based on data characteristics to explore the rich feature information of optical RS data deeply. Subsequently, two relation-guided fusion strategies are proposed at the dual-level (intradomain and interdomain) to integrate multisource image information effectively. In the intradomain feature fusion, an adaptive de-redundancy fusion module (ADRF) is introduced to eliminate redundancy so that the spatial and spectral features are complete and compact, respectively. In interdomain feature fusion, we construct a spatial-spectral joint attention module (SSJA) based on interdomain relationships to sufficiently enhance the complementary features, so as to facilitate later fusion. Experiments on various multisource optical RS datasets demonstrate that S2RGF-Net outperforms other state-of-the-art (SOTA) methods. Xueli Geng, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Chain-of-Situation Aware Progressive Inference LearningabstractThe grounded situation recognition (GSR) task aims to recognize the structured semantics of an image to achieve "human-like" event understanding. Most previous studies primarily focus on the visual features of the situation, overlooking the step-by-step cognitive reasoning process that humans employ in complex task settings. Recently, the emergence of multimodal large language models (MLLMs) has provided novel directions for addressing complex problems. However, directly deploying MLLMs on the GSR task is suboptimal due to their tendency to exhibit "hallucination" issues. Additionally, fine-tuning MLLMs for the GSR task incurs high training costs. To address these challenges, inspired by human cognitive theory and the chain-of-thought (CoT) strategy, we propose the chain-of-situation progressive inference learning (CoS-PIL) framework, a lightweight approach that progressively completes verb prediction, noun prediction, and role grounding. The prediction of each step depends on the historical information of the previous step. Specifically, we first design situation prompts tailored to the GSR task and utilize MLLMs to analyze the input image and language prompts, generating heuristic response text for the current situation in the image. Instead of fine-tuning the MLLM, we activate the reasoning capabilities of the frozen MLLM and adapt its generated responses into three lightweight modules: CoS-Verb, CoS-Noun, and CoS-Ground. Considering that MLLMs may generate redundant content, we carefully design the chain-of-interest predictor (CoI-Predictor) to extract key information from the extensive response text and inject it into the model as prompts to enhance the performance. Extensive experiments on the challenging SWiG benchmark demonstrate that CoS-PIL outperforms other state-of-the-art methods. The code is publically available at https://github.com/XDLiuyyy/CoS-PIL. Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationabstractPre-trained vision-language(V-L) models such as CLIP have demonstrated impressive Zero-Shot performance in many downstream tasks. Since adopting contrastive video-text pairs methods like CLIP to video tasks is limited by its high cost and scale, recent approaches focus on efficiently transferring the image-based CLIP to the video domain. A major finding is that fine-tuning the pre-trained model to achieve strong fully supervised performance leads to low zero shot, few shot, and base to novel generalization. Instead, freezing the backbone network to maintain generalization ability weakens fully supervised performance. Otherwise, no single prompt tuning branch consistently performs optimally. In this work, we proposed a multimodal prompt learning scheme that balances supervised and generalized performance. Our prompting approach contains three sections: 1) Independent prompt on both the vision and text branches to learn the language and visual contexts. 2) Inter-modal prompt mapping to ensure mutual synergy. 3) Reducing the discrepancy between the hand-crafted prompt (a video of a person doing [CLS]) and the learnable prompt, to alleviate the forgetting about essential video scenarios. Extensive validation of fully supervised, zero-shot, few-shot, base-to-novel generalization settings for video recognition indicates that the proposed approach achieves competitive performance with less commute cost. Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Zehua Hao, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
AAAI | 8 |
| 2024 | Multiplane Prior Guided Few-Shot Aerial Scene RenderingabstractNeural Radiance Fields (NeRF) have been successfully applied in various aerial scenes, yet they face challenges with sparse views due to limited supervision. The acquisition of dense aerial views is often prohibitive, as unmanned aerial vehicles (UAVs) may encounter constraints in perspective range and energy constraints. In this work, we introduce Multiplane Prior guided NeRF (MPNeRF), a novel approach tailored for few-shot aerial scene rendering-marking a pioneering effort in this domain. Our key insight is that the intrinsic geometric regularities specific to aerial imagery could be leveraged to enhance NeRF in sparse aerial scenes. By investigating NeRF's and Multiplane Image (MPI)'s behavior, we propose to guide the training process of NeRF with a Multiplane Prior. The proposed Multiplane Prior draws upon MPI's benefits and incorporates advanced image comprehension through a Swin V2 Transformer, pre-trained via SimMIM. Our extensive experiments demonstrate that MPN-eRF outperforms existing state-of-the-art methods applied in non-aerial contexts, by tripling the performance in SSIM and LPIPS even with three views available. We hope our work offers insights into the development of NeRF-based applications in aerial scenes with limited data. Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Puhua Chen, Yuwei Guo 0001 |
CVPR | 6 |
| 2024 | Geometric Prior Guided Feature Representation Learning for Long-Tailed Classification
Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006, Puhua Chen |
Int. J. Comput. Vis. | 6 |
| 2024 | Dual-Branch Residual Disentangled Adversarial Learning Network for Facial Expression RecognitionabstractThe facial expression recognition is very important for human-computer interaction. Therefore, a large number of researchers are focusing on this topic research and have acquired many valuable research achievements. However, there still exist many problems that need to be solved for practical applications, such as the impact of identity and appearance differences, posture change etc. In this work, a dual-branch residual disentangled adversarial learning network is proposed to learn more accurate expression features by disentangling the non-expression features from basic features through a novel combinatorial loss function. In the proposed method, dual-branch network structure is designed, one branch with a D-Net module is utilized to explore non-expression features and another branch just uses subtraction operation to obtain expression features. Based on the above network structure, a novel loss function is constructed to guide the two branches to learn different type features, which contains expression recognition loss, adversarial loss and cosine similarity loss. The main highlight of this work is that the proposed method could achieve the disentanglement of expression features and non-expression features just based on a low-complexity network and expression datasets without other auxiliary data. Finally, abundant experimental results on multiple expression datasets have confirmed the proposed method could obtain better expression recognition results than other state-ofthe-art methods. Puhua Chen, Shasha Mao, Xinyue Hui, Ning Huyan |
IEEE Signal Process. Lett. | 1 |
| 2024 | Satellite Video Object Tracking Based on Location PromptsabstractObject Tracking in satellite videos is a challenging task due to the small target size, low spatial resolution, limited appearance and texture information, and the potential for background confusion. While current state-of-the-art tracking methods perform well on natural images, they often produce unsatisfactory results when applied to satellite videos. In this paper, we address these challenges by leveraging location prompts and refining the feature extractor and bounding box refinement module. Furthermore, we integrate motion features to effectively handle illumination variations that frequently arise in satellite videos, thereby enhancing the overall robustness of the tracker. Our proposed approach, abbreviated as SVLPNet, has been thoroughly evaluated through extensive experiments conducted on two authentic satellite video datasets. The obtained results unequivocally showcase the promising potential of SVLPNet in facilitating object tracking on satellite videos. The source code and raw results will be released at https://github.com/Wprofessor/SVLPNet. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Yingjia Gao, Hao Wang 0211, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Shuo Li 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | A Graph Association Motion-Aware Tracker for Tiny Object in Satellite VideosabstractSatellite video object tracking involves tracking a specified tiny object within a wide scene. The insufficient appearance features of these tiny objects pose significant challenges to appearance-based object trackers, particularly in situations involving occlusion, target blur, and similar interferences. In this paper, a novel Graph Association MOtion-aware tracker (GAMO) is proposed for tiny object in satellite videos, which integrates motion and spatial relationship information. First, a Gaussian motion estimator is proposed that decouples motion into velocity and direction, rather than using traditional x-y movement modeling. This estimator predicts the object’s position and estimates motion uncertainty with a directional motion probability map. Furthermore, the estimated motion serves as a prior to guide the proposal sampling. A probabilistic proposal sampling module is designed that samples candidate bounding boxes according to the directional motion probability map, focusing on the region where the target is most likely to appear. Additionally, we implement a graph association module to model and propagate the spatial relationships between the target and neighboring objects over time. This relationship information assists the appearance features in distinguishing the target from similar interferences. Experiments on the Skysat-1, SV248S, and VISO datasets demonstrate the superiority of the proposed tracker. GAMO leverages motion and surrounding information, resulting in significant improvements with minimal computational overhead. The code and results will be publicly available inhttps://github.com/Midkey/GAMO. Zhongjian Huang, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Xiangrong Zhang, Lingling Li 0002, Puhua Chen |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2024 | Visual and Language Collaborative Learning for RGBT Object TrackingabstractDespite the extensive research on RGBT object tracking, there are still several challenges and issues in practical applications, such as modality differences, lighting variations and disappearance of the target, and changes in viewpoint. Existing methods mostly address these issues by fusing image features, while neglecting a significant amount of target label information. To address these challenges, this paper introduces text to drive the alignment of visible and infrared image features, transforming features from different modalities into the same feature space and fully using complementary features between different modalities. Furthermore, inspired by the success of prompt learning in various tasks, we utilize prior boxes and language as prompts to further guide the model in tracking the target. Extensive experiments demonstrate that the proposed VLCTrack tracker has excellent potential in RGBT object tracking. Compared to previous methods developed for this purpose, our approach achieves state-of-the-art performance on three benchmark datasets. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Yingjia Gao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2024 | Mask-Guided Correlation Learning for Few-Shot Segmentation in Remote Sensing ImageryabstractFew-shot segmentation aims to segment specific objects in a query image based on a few densely annotated images and has been extensively studied in recent years. In remote sensing, image segmentation faces challenges such as less training data, large intraclass diversity, and low foreground-background contrast. In this work, we propose a novel few-shot segmentation method in remote sensing imagery based on mask-guided correlation learning (MGCL) to alleviate the above challenges. In our MGCL, a novel mask-guided feature enhancement (MGFE) module is proposed, which makes features have intramask consistency by leveraging oversegmented masks. In order to enhance the contrast between foreground and background, a novel foreground-background correlation (FBC) module is proposed, which enhances background correlation representation by learning foreground correlation and background correlation separately. Furthermore, a novel mask-guided correlation decoder (MGCD) module is proposed to guide the decoder to focus on the consistency within the mask, thereby learning how to segment complete objects and improving segmentation accuracy. Sufficient experiments on the iSAID-$5^{i}$and DLRSD-$5^{i}$datasets show that our MGCL outperforms all comparative methods. In particular, in the one-shot setting of the iSAID-$5^{i}$dataset, we achieve an mIoU of 39.92 based on ResNet50, which is an improvement of 4.25 over the state-of-the-art (SOAT) method. The visualization of features before and after the MGFE module further concretely demonstrates the motivation and advantages of our MGCL. The code is available athttps://github.com/LiShuo1001/MGCL. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Xu Liu 0006, Puhua Chen, Lingling Li 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Intertemporal Interaction and Symmetric Difference Learning for Remote Sensing Image Change CaptioningabstractRemote sensing image change captioning (RSICC) is more challenging than remote sensing change detection task, which requires extracting occurred changes in similar remote sensing image (RSI) pairs while generating change caption. However, few works have been investigated on RSICC, the main challenges come from how to learn abundant change clues and face the modality gap. To handle these problems, we rethink this task from the perspective of obtaining and aligning symmetrical change features for temporal RSIs. In this work, the proposed intertemporal interaction and symmetric difference learning network are cascaded through several multitemporal integration units to model differences from coarse to fine representations. Specifically, we design a cross-temporal attention (CTA) mechanism to probe direct interaction between bi-temporal RSIs for motivating information coupling between intralevel representations and suppressing irrelevant interferences. To learn robust change features, a symmetric difference transformer (SDT) module is devised to guarantee temporal symmetry between the “before-to-after” and “after-to-before” change representations. Besides, the bi-directional triplet ranking loss is adopted to guide the network to learn strongly discriminative and temporal-symmetric change representation. Extensive experimental results on Dubai-CC and LEVIR-CC datasets demonstrate that our framework with the proposed components can achieve excellent performance and surpass recent state-of-the-art methods.https://github.com/romanticLYP/TISDNet Yunpeng Li 0010, Xiangrong Zhang, Xina Cheng, Puhua Chen, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | High-Order Relation Learning Transformer for Satellite Video Object TrackingabstractSurrounding contexts are generally perceived as interfering with object tracking in satellite videos, leading to model drift. From another perspective, they can also be seen as reference objects of the tracked target, the dynamic interactions between them could provide essential information. In this article, a high-order relation learning transformer (HRLT) is proposed for satellite video object tracking, which not only models the high-order interactions of different target-context pairs but also reasons the associations between these high-order relations across multiple frames. First, a spatial high-order relation reasoning (SHR2) module is designed to model the high-order interactions between the target and scene contexts. Second, a temporal high-order relation reasoning (THR2) module is proposed to associate and reason these spatial high-order relations across multiple frames. Third, historical high-order relations are collected to provide more reasoning bases for the current frame prediction. Finally, qualitative and quantitative evaluations are performed on the SV248S, SkySat, and VISO datasets. The results show that HRLT outperforms 20 popular methods in different challenging scenarios. Licheng Jiao, Yangyang Li 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | LGLFormer: Local-Global Lifting Transformer for Remote Sensing Scene ParsingabstractIn deep learning, convolutional neural networks (CNNs) and transformers have gained excellent achievements in remote sensing scene parsing. Strong feature representation ability is still a challenge for them. Besides, the complex scenes are still essential challenges for deep learning in remote sensing scene parsing. In this article, an efficient local–global lifting transformer (LGLFormer) framework is proposed to ease the challenges above. It effectively combines CNNs, transformer, and wavelet transform to build a strong local–global (LG) feature representation network. Besides, global feature learning driven by LG adaptive features is proposed based on the 2-D LG adaptive feature extractor (LGAFE) and refined global feature attention module. The 2-D LG lifting feature extractor is inspired by the lifting scheme, which introduces local and global dependency. Furthermore, two LG lifting schemes are proposed, including the series and parallel modes, which can effectively learn LG relations between pixels. Finally, experiments are validated on three remote sensing benchmark datasets. The proposed LGLFormer achieves the state-of-the-art with 99.02%, 99.2%, and 99.48% overall accuracy (OA) on AID, WHU-RS19, and UCM datasets, respectively. In addition, LGLFormer shows good convergence with competitive parameters. The experimental code will be available athttps://github.com/yutinyang/LGLFormer. Yuting Yang 0008, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Relation Learning Reasoning Meets Tiny Object Tracking in Satellite VideosabstractTiny objects in satellite videos are usually not independent individuals, there exist rich semantic and temporal relations with each other. Thus, modeling and reasoning the variation of such intrinsic relationships can be beneficial for tiny object tracking. In this paper, a relation learning reasoning method is proposed for tiny object tracking in satellite videos. The core of the proposed is the relation reasoning network that consists of a key context module, a global semantic module, and a relation reasoning module sequentially. First, the key context module exploits global key contexts which explicitly or implicitly contribute to the target object, modeling the intrinsic relations with the target. Second, to reason the contribution, the global semantic module analyses the interaction between them in the same frame. Third, the relation reasoning module deduces the target based on the variation of the semantic relations among different frames. Such a relation learning reasoning approach which takes the target as the core is aligned with the satellite tiny object tracking task, significantly improves the identification performance in dense similarity scenes and the retrieval ability after completely occluded. Furthermore, the proposed method is shown to report improved qualitative and quantitative results on Jilin-1 and SkySat satellite video datasets. Licheng Jiao, Yangyang Li 0001, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | High-Resolution Remote Sensing Image Segmentation With Global-Guided Normalization and Local Affinity DistillationabstractIn recent years, high-resolution (HR) remote sensing images (RSIs) segmentation has received growing attention. The huge number of pixels poses a challenge to the semantic segmentation algorithm, which is limited by the storage of GPUs, so the current methods for processing HR RSIs are categorized into two main categories, i.e., global methods and local methods. The former downsamples the original image and loses a lot of feature details. The latter crops the original image and fails to obtain global contextual information. Both types of methods lead to limited segmentation accuracy. In this article, we propose an end-to-end framework, called global injection network (GINet), which explores two levels of feature distribution and feature relationship to achieve tradeoff between global context and local details. In concrete terms, we propose the global-guided normalization (GGN) module, which injects global context information into local branch and modulates local features using global features to enhance the global perception of local branch. In addition, to constrain the spatial consistency of two branches, inspired by the knowledge distillation technique, we propose local affinity distillation (LAD) loss, which distills the relations in local features into global features to keep the similarity of the relationships corresponding to patches in the two branches. The comprehensive experimental results on three large-scale land-cover classification datasets, DeepGlobe ($2448 \times 2448$), Inria Aerial ($5000 \times 5000$), and GID-15 ($7200 \times 6800$), confirm the effectiveness and superiority of our method in HR semantic segmentation tasks. Peng Zhu 0004, Xiangrong Zhang, Xiao Han 0012, Puhua Chen, Xu Tang 0004, Xina Cheng, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Hierarchical Dynamic Graph Clustering NetworkabstractConnections between visual components are ubiquitous. Graphs, as a highly flexible data structure, not only allow imposing relational induction bias on data, but can provide a completely distinct learning perspective for regular image data. In this paper, we propose a hierarchical dynamic graph clustering network (HDGCN) for visual feature learning. We construct hierarchical graph representations in graph domain in an adaptive, data-adaptive and task-adaptive manner. First, the initial graph is constructed in high-dimensional feature domain of images. To mine the hierarchical geometric features in latent graph space, adaptive clustering network (ClusterNet) is performed to learn discriminative clusters and generates cluster-based coarse graph. Then, graph convolutional networks (GCNs) are used to diffuse, transform and aggregate information among clusters. So, the intra-class and inter-class information is fully explored to increase the discriminativity of graph representations. Next, coarsened graph representations are mapped to grid based on its affinity with linear projection features. To further improve the task adaptation of clusters and hierarchical graph representations, ClusterNet and GCNs are fused in the same framework for end-to-end training and clusters is updated dynamically. We have conducted extensive experiments on classification and segmentation tasks. The experimental results fully validate the robustness of the proposed algorithm. Jie Chen 0098, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Puhua Chen, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2024 | A Category-Aware Curriculum Learning for Data-Free Knowledge DistillationabstractConstructing effective proxy data is one of the core challenges in data-free knowledge distillation. The existing models ignore the influence of the category entanglement of the generated data on the distillation. To alleviate this issue, imitating the human learning process, a new category-aware curriculum learning mechanism is proposed in this paper to perform data-free knowledge distillation, called CCL-D. The main ideology of this category-aware curriculum learning mechanism is to provide a new learning mode for data generation and network training, which enables the model to realize the knowledge distillation process from easy to difficult through automated curriculum learning. In this novel learning mechanism, a category-aware monitoring module is proposed to constrain the category attribute of generated data. Based on this monitoring module, the curriculum learning process for data generation and network training is designed and applied. Initially, the generator is guided to obtain new data with clear category features. The utilization of data with apparent category features is easy for student network training, and it enables the student network to learn clear and significant category features at the early training stage. Subsequently, the generator is guided to generate data with category entanglement. Utilizing these new data with category entanglement problems can improve the recognition ability of the student network to interclass interference and enhance network robustness. The effectiveness of the CCL-D is verified on the six benchmark experimental datasets (MNIST, CIFAR-10, CIFAR-100, SVHN, Caltech-101, Tiny-Imagenet). Xiufang Li, Licheng Jiao, Qigong Sun, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | A Knowledge-Based Hierarchical Causal Inference Network for Video Action RecognitionabstractCurrently, existing action recognition methods mainly use a data-driven method to extract spatio-temporal representations of actions for recognition. However, this method may face performance bottlenecks. At the same time, existing action recognition methods are easily affected by the bias of scene information and object information in videos. In order to explore the essential causal relationship between factors and remove bias in action recognition, we introduce the theory of causal inference into the field of action recognition and propose a Knowledge-based Hierarchical Causal Inference Network (KHCIN) to help us step toward a new direction of inference in action recognition. First, we construct a Knowledge-based Hierarchical Causal Graph (KHCG) to structurally represent the scene, object and motion knowledge of a video. Then, in the model inference stage, we perform factual causal inference on a video on the constructed KHCG, and then deploy counterfactual inference on the Direct Content Hierarchy (DCH) and Indirect Interaction Hierarchy (IIH) in the KHCG. For DCH, we intervene in the model at the decision level to highlight bias errors in the model predictions. For the IIH, we focus on intervening in the feature modelling process. The biased interactions are revealed by interrupting the information communication in the feature space. By comparing the results of factual and counterfactual inference, we can easily expose the biased information in the original representations and eliminate them. Driven by counterfactual causal inference, our approach can significantly improve the performance of action recognition while improving model explainability. Extensive experiments demonstrate the effectiveness of this method. We hope that KHCIN can provide some new ideas for better introduction of causal inference theory in the action recognition community in the future. Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Lingling Li 0002, Yuwei Guo 0001, Puhua Chen |
IEEE Trans. Multim. | 7 |
| 2024 | Feature Distribution Representation Learning Based on Knowledge Transfer for Long-Tailed ClassificationabstractReal-world data typically follows a long-tailed distribution. When a small sample of tail classes does not cover the underlying distribution well, methods such as class re-balancing strategies and decoupled training are difficult to work, and additional knowledge needs to be introduced to recover the underlying distribution of the tail classes. In this work, we observe that the similarity between the variances of the feature distributions increases with the class similarity. Then, we also find that well-represented feature distributions typically contain multiple subcenters, which allows for denser samples at the edges of the distribution and promotes model learning to more robust decision bounds. Based on these observations, we propose to calibrate the feature distribution of the tail class by transferring the variance of the feature distribution of the head class, and then sample from the calibrated tail class distribution to generate augmented samples. To coordinate with the tail class calibration method, we also propose label-aware noise suppression (LANS) for reducing the generation of noisy samples and a three-stage training scheme for reshaping decision boundaries and compacting feature learning. Experimental results on iNaturalist2018, ImageNet-LT, CIFAR-10-LT, and CIFAR-100-LT show that our method achieves state-of-the-art performance in most metrics compared to similar approaches. Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006, Puhua Chen |
IEEE Trans. Multim. | 6 |
| 2024 | Multiscale Dynamic Curvelet Scattering NetworkabstractThe feature representation learning process greatly determines the performance of networks in classification tasks. By combining multiscale geometric tools and networks, better representation and learning can be achieved. However, relatively fixed geometric features and multiscale structures are always used. In this article, we propose a more flexible framework called the multiscale dynamic curvelet scattering network (MSDCCN). This data-driven dynamic network is based on multiscale geometric prior knowledge. First, multiresolution scattering and multiscale curvelet features are efficiently aggregated in different levels. Then, these features can be reused in networks flexibly and dynamically, depending on the multiscale intervention flag. The initial value of this flag is based on the complexity assessment, and it is updated according to feature sparsity statistics on the pretrained model. With the multiscale dynamic reuse structure, the feature representation learning process can be improved in the following training process. Also, multistage fine-tuning can be performed to further improve the classification accuracy. Furthermore, a novel multiscale dynamic curvelet scattering module, which is more flexible, is developed to be further embedded into other networks. Extensive experimental results show that better classification accuracies can be achieved by MSDCCN. In addition, necessary evaluation experiments have been performed, including convergence analysis, insight analysis, and adaptability analysis. Jie Gao 0013, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | A Complex-Former Tracker With Dynamic Polar Spatio-Temporal EncodingabstractRecently, the excellent performance of transformer has attracted the attention of the visual community. Visual transformer models usually reshape images into sequence format and encode them sequentially. However, it is difficult to explicitly represent the relative relationship in distance and direction of visual data with typical 2-D spatial structures. Also, the temporal motion properties of consecutive frames are hardly exploited when it comes to dynamic video tasks like tracking. Therefore, we propose a novel dynamic polar spatio-temporal encoding for video scenes. We use spiral functions in polar space to fully exploit the spatial dependences of distance and direction in real scenes. We then design a dynamic relative encoding mode for continuous frames to capture the continuous spatio-temporal motion characteristics among video frames. Finally, we construct a complex-former framework with the proposed encoding applied to video-tracking tasks, where the complex fusion mode (CFM) realizes the effective fusion of scenes and positions for consecutive frames. The theoretical analysis demonstrates the feasibility and effectiveness of our proposed method. The experimental results on multiple datasets validate that our method can improve tracker performance in various video scenarios. Licheng Jiao, Hao Zhu 0009, Zhongjian Huang, Fang Liu 0001, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Self-Supervised Self-Organizing Clustering Network: A Novel Unsupervised Representation Learning MethodabstractDeep learning-based clustering methods usually regard feature extraction and feature clustering as two independent steps. In this way, the features of all images need to be extracted before feature clustering, which consumes a lot of calculation. Inspired by the self-organizing map network, a self-supervised self-organizing clustering network ( [Formula: see text]OCNet) is proposed to jointly learn feature extraction and feature clustering, thus realizing a single-stage clustering method. In order to achieve joint learning, we propose a self-organizing clustering header (SOCH), which takes the weight of the self-organizing layer as the cluster centers, and the output of the self-organizing layer as the similarities between the feature and the cluster centers. In order to optimize our network, we first convert the similarities into probabilities which represents a soft cluster assignment, and then we obtain a target for self-supervised learning by transforming the soft cluster assignment into a hard cluster assignment, and finally we jointly optimize backbone and SOCH. By setting different feature dimensions, a Multilayer SOCHs strategy is further proposed by cascading SOCHs. This strategy achieves clustering features in multiple clustering spaces. [Formula: see text]OCNet is evaluated on widely used image classification benchmarks such as Canadian Institute For Advanced Research (CIFAR)-10, CIFAR-100, Self-Taught Learning (STL)-10, and Tiny ImageNet. Experimental results show that our method significant improvement over other related methods. The visualization of features and images shows that our method can achieve good clustering results. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Puhua Chen, Lingling Li 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | CTACL:Hyperspectral Image Change Detection Based on Adaptive Contrastive LearningabstractHyperspectral image change detection (HSI-CD) can accurately identify changing regions by capturing subtle spectral differences and has become a research hotspot in the field of remote sensing (RS). Convolutional neural networks (CNNs) have excellent local context modeling capabilities and have been proven to be powerful feature extractors in HSI-CD. However, due to its inherent network structure limitation, CNN cannot well mine and represent the sequential properties of spectral features, especially the medium and long-term dependencies. In contrast, transformer-based network architecture shows a strong ability to model long-distance dependencies, which can fully mine and extract global features, but exhibits weak performance in extracting local information. To this end, we propose HSI-CD network based on adaptive contrastive learning (CTACL). Specifically, we first propose a parallel network of CNNs and transformers to mine local and global temporal-spatial-spectral features of HSI, respectively. Second, we propose adaptive contrastive learning to pre-train the network to learn the latent features of a large amount of unlabeled data and better mine and utilize local and global information. Experimental results on the farmland dataset show that the proposed method performs well. Shunli Tian, Xiangrong Zhang, Guanchun Wang, Xiao Han 0012, Puhua Chen, Xina Cheng |
IGARSS | 5 |
| 2023 | Global-Local Representation Coupling Network for Remote Sensing Image Change DetectionabstractChange detection is one of the important tasks in remote sensing image processing, and the powerful feature extraction ability of convolutional networks has achieved some success in change detection. However, the problem of the limited field size of pure convolutional networks makes the change detection accuracy of high-resolution remote sensing images limited. The introduction of transformers can link the concept of long-distance in space and time. Therefore, in order to maximize the respective advantages of transformers and CNNs, we propose a new parallel architecture. To accomplish the above goal, we propose a new network, which consists of a local detail branch and transformer global spatial-temporal feature branch and a feature fusion module. The experimental results reach the current sota level. Fanghan Yang, Xiangrong Zhang, Peng Zhu 0004, Zhenhang Weng, Puhua Chen |
IGARSS | 5 |
| 2023 | Task context transformer and GCN for few-shot learning of cross-domain
Pengfang Li, Fang Liu 0001, Licheng Jiao, Lingling Li 0002, Puhua Chen, Shuo Li 0010 |
Neurocomputing | 5 |
| 2023 | Knowledge transfer evolutionary search for lightweight neural architecture with dynamic inference
Xiaoxue Qian, Fang Liu 0001, Licheng Jiao, Xiangrong Zhang, Shuo Li 0010, Puhua Chen, Xu Liu 0006 |
Pattern Recognit. | 7 |
| 2023 | Dual Wavelet Attention Networks for Image ClassificationabstractGlobal average pooling (GAP) plays an important role in traditional channel attention. However, there is the disadvantage of insufficient information to use the result of GAP as the channel scalar. At the same time, the existing spatial attention models focus on the areas of interest using average pooling or convolutional networks, but there is a loss of feature information and neglect of the structural feature. In this paper, dual wavelet attention is proposed, which can effectively alleviate the aforementioned problems and enhance the representation ability of CNNs. Firstly, the equivalence between the sum of the low-frequency subband coefficients of 2D DWT (Haar) and GAP is proved. On this basis, the statistical characteristics of low-frequency and high-frequency subbands are effectively combined to obtain the channel scalars, which can better measure the importance of each channel. In addition, 2D DWT can effectively capture the approximate and detailed structural features. Thus, wavelet spatial attention is proposed, which can effectively focus on the key spatial structural features. Different from traditional spatial attention, it can better curve the structural and spatial attention for different channels. The experiments are verified on four natural image data sets and three remote sensing scene classification data sets, which shows the effectiveness and versatility of the proposed methods. The code of this paper will be available athttps://github.com/yutinyang/DWAN. Yuting Yang 0008, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Lingling Li 0002, Puhua Chen, Xiufang Li, Zhongjian Huang |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Learning Salient Feature for Salient Object Detection Without LabelsabstractSupervised salient object detection (SOD) methods achieve state-of-the-art performance by relying on human-annotated saliency maps, while unsupervised methods attempt to achieve SOD by not using any annotations. In unsupervised SOD, how to obtain saliency in a completely unsupervised manner is a huge challenge. Existing unsupervised methods usually gain saliency by introducing other handcrafted feature-based saliency methods. In general, the location information of salient objects is included in the feature maps. If the features belonging to salient objects are called salient features and the features that do not belong to salient objects, such as background, are called nonsalient features, by dividing the feature maps into salient features and nonsalient features in an unsupervised way, then the object at the location of the salient feature is the salient object. Based on the above motivation, a novel method called learning salient feature (LSF) is proposed, which achieves unsupervised SOD by LSF from the data itself. This method takes enhancing salient feature and suppressing nonsalient features as the objective. Furthermore, a salient object localization method is proposed to roughly locate objects where the salient feature is located, so as to obtain the salient activation map. Usually, the object in the salient activation map is incomplete and contains a lot of noise. To address this issue, a saliency map update strategy is introduced to gradually remove noise and strengthen boundaries. The visualization of images and their salient activation maps show that our method can effectively learn salient visual objects. Experiments show that we achieve superior unsupervised performance on a series of datasets. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Xu Liu 0006, Puhua Chen |
IEEE Trans. Cybern. | 5 |
| 2023 | Weak-to-Strong Consistency Learning for Semisupervised Image SegmentationabstractSupervised remote sensing (RS) image segmentation has achieved remarkable success with large amounts of manually labeled data, which may be difficult to acquire in some practical application scenarios. Semisupervised RS image segmentation can efficiently utilize the knowledge embedded in unlabeled data to improve recognition performance, which is of great significance for the generalization application of segmentation models. In this work, we propose an end-to-end semisupervised RS image segmentation method based on weak-to-strong consistency learning, denoted as WSCL. Specifically, a common strong data augmentation technique for image segmentation is introduced to provide powerful input perturbation to decouple self-biased cognition. By forcing weakly augmented, and strongly augmented perspectives from the same sample to be consistent, WSCL not only enables the model to steadily learn knowledge contained in unlabeled data but also alleviates overfitting. In addition, a novel sparse dual-view cross-sample image generation method is presented to generate new training samples, which helps provide a more comprehensive diversity of perturbations. Furthermore, an adaptive re-weighting strategy based on the entropy maps of the outputs of strongly perturbed samples is proposed to suppress noise, guiding the training process in a positive direction. Extensive experiments demonstrate the significant advantage of WSCL over other advanced methods, achieving new state-of-the-art under several evaluation metrics on DFC22, iSAID, MER, MSL, Vaihingen, and GID-15 datasets. The source code is open-sourced at https://github.com/xiaoqiang-lu/WSCL. Xiaoqiang Lu, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001, Zhixi Feng, Puhua Chen |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2023 | Which Target to Focus on: Class-Perception for Semantic Segmentation of Remote SensingabstractDeep Learning-based (DL) methods have dominated the task of semantic segmentation of remote sensing images. However, the sizes of different objects vary widely, and there is a great deal of label-noise due to the inevitable shadows. Therefore, there is an urgent need for a method that can precisely handle complex ground data. In this paper, we propose an Inter-Class Enhanced Network (ICEN) for representing features of varying sizes. It comprises two branches: Sparse Representation Network (SPN) and Feature Extraction Network (FEN). Then, a Class-Perception Block is inserted between the two branches to instruct the SPN’s low-level semantic features to be merged into the deeper network. Such a block can reduce label-noise in remote sensing image segmentation. In addition, the proposed EIRI provides a more precise classification process for target edges containing many misclassified points without requiring excessive computational overhead. The experimental results of our proposed Class-Perception Network (C-PNet) achieve competitive performance on the Vaihingen, Potsdam, LoveDA, and UAVid datasets. Lingling Li 0002, Yilin Shao, Licheng Jiao, Xu Liu 0006, Puhua Chen, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | SDCDNet: A Semi-Dual Change Detection Network Framework With Super-Weak Label for Remote Sensing ImageabstractMost current change detection methods require a large amount of labeled data to train huge parameters. To break this limitation, this paper proposes a novel semi-supervised learning framework for remote sensing change detection, named a semi-dual change detection network (SDCDNet). The SDCDNet consists of a dual shared network and dual branching networks. The dual shared network is designed to exploit the full potential of the data, and the dual branching network is proposed to differentiate the kinds of annotated data and eliminate the disturbance between different types of data. In addition, the adaptive weighting module (AWM) enhances the features of weak branching, and the mask constraint module (MCM) is proposed to increase the ability of the network to extract foreground features. To solve the complex problem of data labeling, a patch-based weak label construction method is proposed to build super-weak labels. Experiments show that the proposed SDCDNet achieves excellent results on two remote sensing image change detection datasets. Jiahao Wang 0002, Fang Liu 0001, Hao Wang 0211, Xu Liu 0006, Licheng Jiao, Lingling Li 0002, Puhua Chen |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2023 | High-Quality Angle Prediction for Oriented Object Detection in Remote Sensing ImagesabstractOriented object detection is a challenging task in remote sensing, where the detected objects can be represented by oriented bounding boxes (OBBs). Angle prediction in oriented object detection has been widely studied, due to its crucial role in object detection. However, the precision of angle prediction is severely limited by misalignments in most of the existing methods, including representation-, evaluation-, and optimization-based misalignments. To alleviate these misalignments, this paper presents a novel angle prediction method, called Angle Quality Estimation (AQE). Specifically, our proposed AQE transforms the angle prediction task into a distribution estimation task to address the representation misalignment problem and implicitly measure the quality of the predicted angles. Based on the estimated angle quality, we then propose a new metric to comprehensively evaluate the quality of OBBs. Then we propose an object aspect ratio based loss function to optimize angle prediction for addressing the optimization misalignment. Our proposed AQE is a plug-and-play method, which can be embedded on any existing oriented object detector. Experimental results on three public benchmarks, including DOTA, HRSC2016, and ICDAR2015 datasets, show that our method achieves better performance than the other state-of-the-art. Guanchun Wang, Xiangrong Zhang, Peng Zhu 0004, Xu Tang 0004, Puhua Chen, Licheng Jiao, Huiyu Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | An Explainable Spatial-Frequency Multiscale Transformer for Remote Sensing Scene ClassificationabstractDeep convolutional neural networks (CNNs) are significant in remote sensing. Due to the strong local representation learning ability, CNNs have excellent performance in remote sensing scene classification. However, CNNs focus on location-sensitive representations in the spatial domain and lack contextual information mining capabilities. Meanwhile, remote sensing scene classification still faces challenges, such as complex scenes and significant differences in target sizes. To address the problems and challenges above, more robust feature representation learning networks are necessary. In this paper, a novel and explainable spatial-frequency multi-scale Transformer framework, SF-MSFormer, is proposed for remote sensing scene classification. It mainly comprises spatial-domain and frequency-domain multi-scale Transformer branches, which consider the spatial-frequency global multi-scale representation features. Besides, the texture-enhanced encoder is designed in the frequency-domain multi-scale Transformer branch, which is adaptive to capture the global texture features. In addition, an adaptive feature aggregation module is designed to integrate the spatial-frequency multi-scale feature for final recognition. The experimental results verify the effectiveness of SF-MSFormer and show better convergence. It achieves state-of-the-art results (98.72%, 98.6%, 99.72%, and 94.83% overall accuracies, respectively) on the AID, UCM, WHU-RS19, and NWPU-RESISC45 datasets. Besides, the feature visualizations evaluate the explainability of the texture-enhanced encoder. The code implementation of this article will be available at https://github.com/yutinyang/SF-MSFormer. Yuting Yang 0008, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | MFGNet: Multibranch Feature Generation Networks for Few-Shot Remote Sensing Scene ClassificationabstractFew-shot remote sensing scene classification aims to identify unseen classes using only a small number of labeled samples. Considering the large intra-class variances and inter-class similarity of remote sensing scenes, most existing methods focus on feature extraction, ignoring the overfitting problem caused by insufficient samples. To this end, we propose a novel few-shot learning framework, called multibranch feature generation networks (MFGNets), which solves the few-shot scene classification from the source by online sample generation at the representation space. Specifically, we first build a feature generation net to transform the few-shot classification into a regular classification problem, in which the generated samples are achieved by combining the class-specific features with the sampled intra-class features. Then, to ensure the quality of the generated samples, we introduce two novel regularization terms: the intra-class diversity loss (ID-Loss) and the inter-class consistency loss (IC-Loss), which aid the model in generating more diverse samples. Furthermore, we introduce a scale-angle aware self-supervised pretext to learn scale-invariant and rotation-invariant features, improving the model’s feature representation capability in remote sensing scenes. We evaluate the proposed method on three publicly available datasets, namely UC_Merced, NWPU-RESISC45, and AID. Our approach has achieved state-of-the-art performance, with an improvement of more than 3.31%, 2.64%, and 6.86% on the most challenging 1-shot tasks, respectively. Xiangrong Zhang, Xiyu Fan, Guanchun Wang, Puhua Chen, Xu Tang 0004, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | A Spatial Hierarchical Reasoning Network for Remote Sensing Visual Question AnsweringabstractFor visual question answering on remote sensing (RSVQA), current methods scarcely consider geospatial objects typically with large-scale differences and positional sensitive properties. Besides, modeling and reasoning the relationships between entities have rarely been explored, which leads to one-sided and inaccurate answer predictions. In this article, a novel method called spatial hierarchical reasoning network (SHRNet) is proposed, which endows a remote sensing (RS) visual question answering (VQA) system with enhanced visual–spatial reasoning capability. Specifically, a hash-based spatial multiscale visual representation module is first designed to encode multiscale visual features embedded with spatial positional information. Then, spatial hierarchical reasoning is conducted to learn the high-order inner group object relations across multiple scales under the guidance of linguistic cues. Finally, a visual-question (VQ) interaction module is employed to learn an effective image–text joint embedding for the final answer predicting. Experimental results on three public RS VQA datasets confirm the effectiveness and superiority of our model SHRNet. Zixiao Zhang, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Fang Liu 0001, Yuxuan Li 0004, Zhicheng Guo |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Spectral-Spatial Distribution Consistent Network Based on Meta-Learning for Cross-Domain Hyperspectral Image ClassificationabstractCross-domain networks can solve the problem of insufficient labeled samples, especially for hyperspectral images (HSIs) where obtaining labeled samples is time-consuming and laborious. Most of the current methods rely on the spatial information to achieve domain alignment, without considering the rich spectral information of HSIs. Furthermore, the methods based on convolutional neural network (CNN) cannot get the spatial information of irregular image regions, resulting in poor classification results of object edges. Therefore, we design a spectral-spatial distribution consistent network (SSDC) based on meta-learning. Firstly, to improve the feature extraction ability of the cross-domain classification model, we introduce a feature pre-extraction module, which uses the spectral attention mechanism and the alternating meta-learning method to obtain the general features of the source domain and the discriminative features of the target domain, so as to obtain the spectral weight matrix for subsequent processing. Secondly, we propose a spectral consistent module based on singular value decomposition, which increases the difference between different classes of features by penalizing the singular values of the feature matrix to achieve data distribution alignment in the spectral dimension. Finally, aiming at the low classification accuracy of irregular image regions, we propose a spatial consistent module to obtain non-local spatial topological information through stacked cross modules and graph sample and aggregate networks, which can reduce domain shift. The experiments of SSDC on four classical HSI datasets show that the proposed method can obtain competitive results with other methods based on CNN and cross-domain. Xiangrong Zhang, Qi Zhen, Xiao Han 0012, Puhua Chen, Xu Tang 0004, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Semantics and Contour Based Interactive Learning Network for Building Footprint ExtractionabstractBuilding footprint extraction plays an important role in the analysis of remote sensing images and has an extensive range of applications. Obtaining precise boundaries of buildings remains a challenge in existing building extraction methods. Some previous works have made notable efforts to address this concern. However, most of these methods require cumbersome and expensive post-processing steps. Moreover, they ignored the correlation between building semantics and contours, which we believe is crucial for building footprint extraction. To mitigate this issue, our paper presents an intuitive and effective framework that explores semantic and contour cues of buildings and fully excavates their correlation. Specifically, we construct an interactive dual-stream decoder. The Intermediate connections within this decoder interactively transmit features between branches, contributing to learning correlations between semantics and contours. We propose the Semantic Collaboration Module (SCM) to strengthen the connection between the two branches. To further boost performance, we build the Multi-Scale Semantic Context Fusion Module (MSCF) to fuse semantic information from the higher and lower layers of the network, allowing the network to obtain superior feature representations. The experimental results on the WHU, INRIA, and Massachusetts building datasets demonstrate the superior performance of our method. Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Xu Tang 0004, Puhua Chen, Huiyu Zhou 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | D³K: Dynastic Data-Free Knowledge DistillationabstractData-free knowledge distillation further broadens the applications of the distillation model. Nevertheless, the problem of providing diverse data with rich expression patterns needs to be further explored. In this paper, a novel dynastic data-free knowledge distillation ($D^{3}K$) model is proposed to alleviate this problem. In this model, a dynastic supernet generator (D-SG) with a flexible network structure is proposed to generate diverse data. The D-SG can adaptively alter architectural configurations and activate different subnet generators in different sequential iteration spaces. The variable network structure increases the complexity and capacity of the generator, and strengthens its ability to generate diversified data. In addition, a novel additive constraint based on the differentiable dhash (D-Dhash) is designed to guide the structure parameter selection of the D-SG. This constraint forces the D-SG to constantly jump out of the fixed generation mode and generate diverse data in semantics and instance. The effectiveness of the proposed model is verified on the experimental benchmark datasets (MNIST, CIFAR-10, CIFAR-100, and SVHN). Xiufang Li, Qigong Sun, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Yi Zuo 0003 |
IEEE Trans. Multim. | 7 |
| 2023 | Transformer Based Conditional GAN for Multimodal Image FusionabstractMultimodal Image fusion is becoming urgent in multi-sensor information utilization. However, existing end-to-end image fusion frameworks ignore a priori knowledge integration and long-distance dependencies across domains, which brings challenges to the network convergence and global image perception in complex scenes. In this paper, a conditional generative adversarial network with transformer (TCGAN) is proposed for multimodal image fusion. The generator is to generate a fused image with the source images content. The discriminators are adopted to distinguish the differences between the fused image and the source images. Adversarial training makes the final fused image to maintain the structural and textural details in the cross-modal images simultaneously. In particular, a wavelet fusion module makes the inputs contain image content from different domains as much as possible. The extracted convolutional features interact in the multiscale cross-modal transformer fusion module to fully complement the associated information. It makes the generator to focus on both local and global context. TCGAN fully considers the training efficiency of the adversarial process and the integrated retention of redundant information. Various experimental results of TCGAN have highlighted targets, rich details, and fast convergence properties on public datasets. Jun Zhang 0045, Licheng Jiao, Wenping Ma 0001, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Multim. | 7 |
| 2023 | Deep Learning in Visual Tracking: A ReviewabstractDeep learning (DL) has made breakthroughs in many computer vision tasks and also in visual tracking. From the beginning of the research on the automatic acquisition of high abstract feature representation, DL has gone deep into all aspects of tracking to date, to name a few, similarity metric, data association, and bounding box estimation. Also, pure DL-based trackers have obtained the state-of-the-art performance after the community's constant research. We believe that it is time to comprehensively review the development of DL research in visual tracking. In this article, we overview the critical improvements brought to the field by DL: deep feature representations, network architecture, and four crucial issues in visual tracking (spatiotemporal information integration, target-specific classification, target information update, and bounding box estimation). The scope of the survey of DL-based tracking covers two primary subtasks for the first time, single-object tracking and multiple-object tracking. Also, we analyze the performance of DL-based approaches and give meaningful conclusions. Finally, we provide several promising directions and tasks in visual tracking and relevant fields. Licheng Jiao, Yidong Bai, Puhua Chen, Fang Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Simple and Efficient: A Semisupervised Learning Framework for Remote Sensing Image Semantic SegmentationabstractSemantic segmentation based on deep learning has achieved impressive results in recent years, but these results are supported by a large amount of labeled data which requires intensive annotation at the pixel level, particularly for high-resolution remote sensing (RS) images. In this work, we propose a simple yet efficient semisupervised learning framework based on linear sampling self-training, named LSST, to improve the performance of RS image semantic segmentation. Specifically, the classical pseudo-labeling-based self-training paradigm is enhanced by injecting strong data augmentations (SDA) applicable to RS images, based on which a powerful baseline is constructed. Nevertheless, the problem of insufficient data training to generate pseudo-labels with a high level of noise persists, and the noisy pseudo-labels will continue to accumulate and impede model improvement during the re-training phase. Previous works commonly employ a pre-defined threshold to remove noise, but it will lead to overfitting the model to easily identified classes. To address it, a method using linear sampling (LS) is presented for assigning thresholds to different classes in an adaptive manner, which provides noiseless regions for re-training. Experiments prove that the proposed pixel-wise selection is more available for segmentation than image-level selection in RS images. Finally, LSST achieves state-of-the-art on several datasets and different evaluation metrics. The source code of the this paper is available at https://github.com/xiaoqiang-lu/LSST. Xiaoqiang Lu, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006, Zhixi Feng, Lingling Li 0002, Puhua Chen |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2022 | A Hybrid Network With Structural Constraints for SAR Image Scene ClassificationabstractData-based image classification methods, such as convolutional neural networks (CNNs), have achieved state-of-the-art performance. They usually leverage thousands of labeled samples to train the networks but ignore some prior knowledge. However, labeled samples are difficult to be obtained for synthetic aperture radar (SAR) images. Model-based methods are adept at utilizing the prior information of data, while they have to introduce some restrictions or assumptions during the realization of models. Consequently, to develop the advantages of both methods and improve their disadvantages, we propose a hybrid network by coupling the data-based with model-based methods for SAR image scene classification in this article. First, to fully use the prior information of SAR images and large amounts of unlabeled samples, we improve the$G^{0}$-based variational Bayesian inference model (GVBI) and construct a$G^{0}$-based convolutional variational auto-encoder (GCVAE) for unsupervised learning of the distributional characteristics of SAR images. After that, we further extend the GCVAE by combining it with CNN, resulting in a stronger hybrid network to classify SAR images with a few labeled samples. In addition, considering the abundant structural information is crucial for SAR image classification, we design a sketch fitter and two structural constraints on both pixel and sketch spaces to assist the hybrid network to improve its classification performance. Finally, we evaluate the performance of our method on real-SAR images, and the experimental results demonstrate that the proposed framework outperforms related methods on classification while reducing the manual annotation substantially. Xiaoxue Qian, Fang Liu 0001, Licheng Jiao, Xiangrong Zhang, Puhua Chen, Lingling Li 0002, Yuanhao Cui |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Remote Sensing Image Super-Resolution via Dual-Resolution Network Based on Connected Attention MechanismabstractLimited by hardware conditions and complex degradation processes, the obtained remote sensing images (RSIs) are often low-resolution (LR) data with insufficient high-frequency information. Image super-resolution (SR) aims to improve the spatial resolution of images and add reasonable detailed information. Although existing convolutional neural network (CNN)-based methods achieve good performance by adding residual structure and attention mechanism to the network, simply stacking the residual structure and embedding the attention module directly on the residual branch lead to localized use of features and information loss. To address the above problems, we propose a dual-resolution connected attention network (DRCAN). Specifically, a high-resolution (HR) learning branch is constructed to complement the mapping learning between LR images and HR images, and a connected attention module with residual learning is introduced to make full use of the different levels of intermediate layer features. Besides, we collect data at different resolutions from Google Earth to form a dataset named XD IPIU for RSIs SR. Extensive experiments demonstrate the effectiveness of the proposed model and DRCAN shows the state-of-the-art performance in terms of quantitative evaluation and visual quality. Xiangrong Zhang, Tianyang Zhang 0002, Fengsheng Liu, Xu Tang 0004, Puhua Chen, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Foreground Refinement Network for Rotated Object Detection in Remote Sensing ImagesabstractObject detection has been a fundamental task in the field of remote sensing and has made considerable progress in recent years. However, the high background complexity in remote sensing images (RSIs) remains challenging. In this article, we propose a refined rotation detector, namely, the Foreground Refinement Network (FoRDet), to alleviate the above problem by leveraging the information of foreground regions from the perspectives of feature and optimization. Specifically, we propose a foreground relation module (FRL) that aggregates the foreground-contextual representations from the coarse stage and improves the discrimination of foreground regions on feature maps in the refined stage. Besides, considering the risk of the potential foreground anchors being overwhelmed in the training phase, we design a foreground anchor reweighting (FRW) loss that integrates the classification confidence and localization accuracy of each foreground anchor from the coarse stage to dynamically regulate their contributions in the refined stage, which highlights the potential foreground anchors. The comprehensive experimental results on three public datasets for rotated object detection DOTA, HRSC2016, and UCAS-AOD demonstrate the effectiveness of our proposed method. Tianyang Zhang 0002, Xiangrong Zhang, Peng Zhu 0004, Puhua Chen, Xu Tang 0004, Chen Li 0011, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | MFNet: A Novel GNN-Based Multi-Level Feature Network With Superpixel PriorsabstractSince the superpixel segmentation method aggregates pixels based on similarity, the boundaries of some superpixels indicate the outline of the object and the superpixels provide prerequisites for learning structural-aware features. It is worthwhile to research how to utilize these superpixel priors effectively. In this work, by constructing the graph within superpixel and the graph among superpixels, we propose a novel Multi-level Feature Network (MFNet) based on graph neural network with the above superpixel priors. In our MFNet, we learn three-level features in a hierarchical way: from pixel-level feature to superpixel-level feature, and then to image-level feature. To solve the problem that the existing methods cannot represent superpixels well, we propose a superpixel representation method based on graph neural network, which takes the graph constructed by a single superpixel as input to extract the feature of the superpixel. To reflect the versatility of our MFNet, we apply it to an image-level prediction task and a pixel-level prediction task by designing different prediction modules. An attention linear classifier prediction module is proposed for image-level prediction tasks, such as image classification. An FC-based superpixel prediction module and a Decoder-based pixel prediction module are proposed for pixel-level prediction tasks, such as salient object detection. Our MFNet achieves competitive results on a number of datasets when compared with related methods. The visualization shows that the object boundaries and outline of the saliency maps predicted by our proposed MFNet are more refined and pay more attention to details. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Puhua Chen, Xu Liu 0006, Lingling Li 0002 |
IEEE Trans. Image Process. | 4 |
| 2020 | Multi-Feature Weighted Sparse Graph for SAR Image AnalysisabstractSparse representation (SR) method has the advantages of good category distinguishing performance, noise robustness, and data adaptiveness. In this article, a multi-feature weighted sparse graph (MWSG) is presented for synthetic aperture radar (SAR) image analysis. First, multiple types of features are extracted to fully describe the characteristics of SAR image. Then, multiple SRs of samples in multiple feature spaces are obtained by solving a weighted joint SR model, in which the weight is the Gaussian kernel distance among samples. Moreover, a new fusion mechanism is given to integrate multiple weighted SRs, which aims to eliminate the negative influence of the singular data, so the MWSG is obtained. Afterward, the brief steps of the SAR image segmentation and semisupervised classification based on MWSG are stated. A series of experiments on the simulated and real SAR images shows that the MWSG has better performance than other existing relevant methods. Licheng Jiao, Fang Liu 0001, Xiangrong Zhang, Xu Tang 0004, Puhua Chen |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2020 | Sketch-Based Region Adaptive Sparse Unmixing Applied to Hyperspectral ImageabstractHyperspectral image (HSI) unmixing is an important issue of research due to its effect on the subsequent processing of HSIs. Recently, the sparse regression method with spatial information has been successfully applied in hyperspectral unmixing (HU). However, most sparse regression methods ignore the difference in spatial structure handling with only one sparse constraint. In fact, the pixels in detail regions are more likely to be severely mixed with more endmembers participated, and the sparsity degree of its corresponding abundances is relatively low. Considering the sparsity difference of abundances, a sketchbased region adaptive sparse unmixing applied to HSI is proposed in this article. Inspired by the vision computing theory, we use the region generation algorithm based on a sketch map to differentiate the homogeneous regions and detail regions. Then, the abundances of these two kind regions in HSIs are separately constrained by sparse regularizers of L1/2and L1with a proposed manifold constraint. Our method not only makes full use of the spatial information in HSIs but also exploits the latent structure of data. The encouraging experimental results on three data sets validate the effectiveness of our method for HU. Xiangrong Zhang, Xu Tang 0004, Puhua Chen, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | Random subspace based ensemble sparse representation
Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Rongfang Wang, Puhua Chen, Yuanhao Cui, Junhu Xie, Yake Zhang |
Pattern Recognit. | 6 |
| 2017 | Fast Classification for Large Polarimetric SAR Data Based on Refined Spatial-Anchor GraphabstractThe graph model-based semisupervised machine learning is well established. However, its computational complexity is still high in terms of the time consumption especially for large data. In this letter, we propose a fast semisupervised classification algorithm using the recently presented spatial-anchor graph for a large polarimetric synthetic aperture radar (Pol-SAR) data, named as Fast Spatial-Anchor Graph (FSAG) based algorithm. Based on an initial superpixel segmentation on the PolSAR image, the homogenous regions are obtained. The border pixels are reassigned to the most similar superpixel according to majority voting and distance measurement. Then, feature vectors are weighted within local homogenous regions. The refined spatial-anchor graph is constructed with these regions, and the semisupervised classification is conducted. Experimental results on synthesized and real PolSAR data indicate that the proposed FSAG greatly reduces time consumption and maintains the accuracy for terrain classifications compared with state-of-the-art graph-based approaches. Hongying Liu 0001, Shuyuan Yang 0001, Shuiping Gou, Puhua Chen, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2017 | Semi-supervised double sparse graphs based discriminant analysis for dimensionality reduction
Puhua Chen, Licheng Jiao, Fang Liu 0001, Jiaqi Zhao 0001, Shuai Liu 0016 |
Pattern Recognit. | 1 |
| 2016 | Locality-constraint discriminant feature learning for high-resolution SAR image classification
Licheng Jiao, Biao Hou, Shuang Wang 0001, Jiaqi Zhao 0001, Puhua Chen |
Neurocomputing | 6 |
| 2016 | Semisupervised Discriminant Feature Learning for SAR Image Category via Sparse EnsembleabstractTerrain scene classification plays an important role in various synthetic aperture radar (SAR) image understanding and interpretation. This paper presents a novel approach to characterize SAR image content by addressing category with a limited number of labeled samples. In the proposed approach, each SAR image patch is characterize by a discriminant feature which is generated in a semisupervised manner by utilizing a spare ensemble learning procedure. In particular, a nonnegative sparse coding procedure is applied on the given SAR image patch set to generate the feature descriptors first. The set is combined with a limited number of labeled SAR image patches and an abundant number of unlabeled ones. Then, a semisupervised sampling approach is proposed to construct a set of weak learners, in which each one is modeled by a logistic regression procedure. The discriminant information can be introduced by projecting SAR image patch on each weak learner. Finally, the features of SAR image patches are produced by a sparse ensemble procedure which can reduce the redundancy of multiple weak learners. Experimental results show that the proposed discriminant feature learning approach can achieve a higher classification accuracy than several state-of-the-art approaches. Licheng Jiao, Fang Liu 0001, Jiaqi Zhao 0001, Puhua Chen |
IEEE Trans. Geosci. Remote. Sens. | 5 |