Xi Yang 0011

dblp:13/1520-11 · DBLP profile ↗
← Back
122ranked-venue papers
65as first author
97since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 62 · 36 first-author · 52 since 2021Artificial intelligence and machine learning · 38 · 16 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 14 first-author · 20 since 2021Security and privacy · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Out-of-Context Misinformation Detection via Variational Domain-Invariant Learning with Test-Time Training
abstract
Out-of-context misinformation (OOC) is a low-cost form of misinformation in news reports, which refers to place authentic images into out-of-context or fabricated image-text pairings. This problem has attracted significant attention from researchers in recent years. Current methods focus on assessing image-text consistency or generating explanations. However, these approaches assume that the training and test data are drawn from the same distribution. When encountering novel news domains, models tend to perform poorly due to the lack of prior knowledge. To address this challenge, we propose Variational Domain-Invariant Learning with Test-Time Training (VDT) framework to enhance the domain adaptation capability for OOC misinformation detection. Domain-Invariant Variational Align module is employed to jointly encodes source and target domain data to learn a separable distributional space and domain-invariant features. For preserving semantic integrity, we utilize domain consistency constraint module to reconstruct the source and target domain latent distribution. During testing phase, we adopt the test-time training strategy and confidence-variance filtering module to dynamically updating the VAE encoder and classifier, facilitating the model's adaptation to the target domain distribution. Extensive experiments conducted on the benchmark dataset NewsCLIPpings demonstrate that our method outperforms state-of-the-art baselines under most domain adaptation settings.
Xi Yang 0011, Zhijian Lin, Yibiao Hu
AAAI1
2026 An Enhanced Adaptive Confidence Margin for Semi-Supervised Facial Expression Recognition
abstract
Semi-supervised learning (SSL) provides a practical framework for leveraging massive unlabeled samples, especially when labels are expensive for facial expression recognition (FER). Typical SSL methods like FixMatch select unlabeled samples with confidence scores above a fixed threshold for training. However, these methods face two primary limitations: failing to consider the varying confidence across facial expression categories and failing to utilize unlabeled facial expression samples efficiently. To address these challenges, we propose an Enhanced Adaptive Confidence Margin (EACM), consisting of dynamic thresholds for different categories, to fully learn unlabeled samples. Specifically, we employ the predictions on labeled samples at each training iteration to learn an EACM. It then partitions unlabeled samples into two subsets: (1) subset I, including samples whose confidence scores are no less than the margin; (2) subset II, including samples whose confidence scores are less than the margin. For samples in subset I, we constrain their predictions on strongly-augmented versions to match the pseudo-labels derived from the predictions on weakly-augmented versions. Meanwhile, we introduce a feature-level contrastive objective to enhance the similarity between two weakly-augmented features of a sample in subset II. We extensively evaluate EACM on image-based and video-based facial expression datasets, showing that our method achieves superior performance, significantly surpassing fully-supervised baselines in a semi-supervised manner. Additionally, our EACM is promising to leverage cross-dataset unlabeled samples for practical training to boost fully-supervised performance.
Hangyu Li 0001, Nannan Wang 0001, Xi Yang 0011, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Condense loss: Exploiting vector magnitude during person Re-identification training process
Xi Yang 0011, Wenjiao Dong, Yingzhi Tang, Gu Zheng, Nannan Wang 0001, Xinbo Gao 0001
Pattern Recognit.1
2026 Probabilistic Distribution Alignment for Text-Based Person Retrieval
Xi Yang 0011, Chenghuan Qi, Nannan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 Nearest Neighbor Sample Constraint and ODE Guided Feature Reconstruction for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification aims to retrieve a given pedestrian image from unlabeled data. The method of clustering and assigning pseudo-labels has become mainstream, but there are still some problems that will reduce recognition accuracy. On the one hand, in the process of clustering, poor classification of hard samples between neighboring classes leads to inadequate clustering accuracy, which affects the quality of pseudo-labels. On the other hand, the representational capacity of features extracted by the backbone network is also crucial for the model’s performance. To this end, this paper proposes an unsupervised person re-identification method based on nearest neighbor sample constraint and ordinary differential equation guided feature reconstruction (NNSC-FR) to improve the clustering accuracy and pseudo-label quality while enhancing the representation of features. Specifically, we present a novel nearest neighbor sample constraint (NNSC) after neighbor sample mining for each instance sample to recognize the hard samples’ fine classification between classes. To further improve clustering accuracy, an inter-class balance loss (CB loss) is introduced to better identify the hard samples between the nearest neighbor classes. In addition, guided by the third-order adam solution of the Ordinary Differential Equation, we design a Feature Reconstruction (ODE-FR) module with residual structure to improve the model representation ability. Extensive experimental results on Market-1501, DukeMTMC-reID, and MSMT17 demonstrate that our proposed method is superior to the state-of-the-art methods.
Xi Yang 0011, Wenjiao Dong, Gu Zheng, Nannan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 AdaNoise: Cycle-Consistent Image Translation With Domain-Adaptive Noise Perturbation
abstract
Image-to-image (I2I) translation aims to transform an input image into a target domain while preserving its structural details. Recent advances in diffusion-based generative models have significantly improved the perceptual quality of generated images; however, these approaches still face challenges in controllability and consistency, largely due to the inherent randomness introduced by stochastic noise during the generation process. Specifically, directly manipulating noise distributions without semantic alignment can lead to mode collapse, texture distortion, or loss of domain-specific features. To overcome these challenges, we propose the Adaptive Noise Framework (AdaNoise), a novel and cycle-consistent I2I translation approach guided by domain-adaptive noise modulation. AdaNoise introduces a Domain-Adaptive Noise Perturbation (DANP) module, which adaptively learns structured noise patterns aligned with the target domain distribution, enhancing both the expressiveness and reliability of the translation process. Through integration with a Cycle-Consistent Dual Diffusion (CDD) architecture, AdaNoise ensures faithful content reconstruction while allowing semantically meaningful domain shifts. The framework is designed to maintain a balance between generation quality and controllability, enabling more faithful and flexible image translation. Extensive experiments on two tasks, including SAR-to-optical image translation and low-light image enhancement, validate that AdaNoise not only surpasses existing state-of-the-art methods in terms of image fidelity and semantic preservation, but also achieves more controllable and diverse outputs across varying conditions, thus offering a robust and scalable solution for cross-domain visual generation.
Xi Yang 0011, Fei Gao 0006, Nannan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 CSHNet: A Novel Information Asymmetric Image Translation Method
abstract
Despite the considerable advancements in cross-domain image translation, a significant challenge remains in addressing information asymmetric translation tasks such as SAR-to-Optical and Sketch-to-Instance conversions. These tasks involve transforming data from a domain with limited information into one with more detailed and richer content. Traditional CNN-based methods, while effective at capturing intricate details, often struggle to grasp the overall structural composition of the image, leading to unintended blending or merging of distinct regions within the generated images. In light of these limitations, research has increasingly turned toward Transformers. Though Transformers excel at capturing global structures, they often lack the ability to preserve fine-grained details. Recognizing the importance of both detailed features and structural relationships in information asymmetric translation tasks, we introduce the CNN-Swin Hybrid Network (CSHNet). This network employs a novel bottleneck architecture featuring two key modules: Swin Embedded CNN (SEC) and CNN Embedded Swin (CES), which together form the SEC-CES-Bottleneck (SCB). Within this structure, SEC capitalizes on CNN’s capability for detailed feature extraction while incorporating the Swin Transformer’s inherent structural bias. In contrast, CES preserves the Swin Transformer’s strength in maintaining global structural integrity, while compensating for CNN’s tendency to emphasize detail. In addition to the SCB architecture, CSHNet integrates two essential components designed to improve cross-domain information retention and ensure structural consistency. The Interactive Guided Connection (IGC) fosters dynamic information exchange between SEC and CES, encouraging a deeper understanding of image details. At the same time, Adaptive Edge Perception Loss (AEPL) is implemented to preserve well-defined structural boundaries throughout the translation process. Experimental evaluations demonstrate that CSHNet surpasses current state-of-the-art methods, achieving superior results in both visualization and performance metrics across scene-level and instance-level datasets. Our code is available at: https://github.com/XduShi/CSHNet.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 Semantic-Interactive Clustering Optimization With SAM for Weakly Supervised Person Search
abstract
Weakly-supervised person search presents significant challenges when relying solely on bounding-box annotations, particularly due to inter-class confusion from clothing similarity and intra-class variations caused by illumination changes, which severely degrade cross-view matching accuracy. Existing clustering-based methods, constrained by their heavy dependence on color features, frequently produce unreliable pseudo-labels that ultimately limit model performance. To overcome these limitations, we present Segment Anything Model-based Semantic-Interactive Clustering Optimization (SAM-SICO), a novel framework that integrates the Segment Anything Model’s semantic segmentation capability with adaptive clustering optimization for weakly-supervised person search. Our framework harnesses the representational power of the Segment Anything Model (SAM) to enable detector-free semantic feature learning while significantly improving clustering precision. The proposed solution makes three key advances: the Semantic Contour Embedding (SCE) module leverages SAM’s zero-shot segmentation capability to produce highly accurate human body masks; the Relation-driven Semantic Feature Interaction (RSFI) mechanism effectively mitigates clothing-color bias through innovative dynamic affinity matrix construction across multiscale semantic masks and visual features; and the Adaptive Clustering Optimization (ACO) algorithm introduces parameter adaptation to optimize intra-class compactness and inter-class separation metrics. Experimental results show that our method outperforms existing state-of-the-art approaches on the PRW and CUHK-SYSU datasets. The source code is available at https://github.com//HawlsonZ/SAM-SICO.
Xi Yang 0011, Hexun Zhou, De Cheng, Menghui Tian, Nannan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 Sparse VMamba: Robust Spatio-Temporal Information Modeling for Event Camera Person Re-Identification
abstract
Event camera-based person re-identification (Re-ID) effectively addresses the challenges faced by traditional Re-ID systems, such as privacy leakage, low-light imaging degradation, and motion blur. However, traditional Convolutional Neural Networks (CNNs) struggle to model long-range spatio-temporal dependencies, while the Transformer architecture encounters fundamental conflicts with second-order computational complexity and the high temporal resolution of event streams. Additionally, sparse data leads to wasted computational resources and diluted effective data. In contrast, the Mamba architecture, with its long-term modeling capability and linear complexity, is better suited for event stream data. Therefore, we innovatively explore the potential of VMamba in event camera-based person Re-ID; however, directly using VMamba does not fully leverage the temporal asynchronicity and spatial sparsity inherent in event data. To address this, we design a novel Sparse VMamba framework to construct a more robust spatio-temporal information extraction mechanism. First, we develop a Spatio-Temporal Information Modeling (STIM) module that simultaneously employs CNNs and Gated Recurrent Units (GRUs) for modeling spatial and temporal information. Then, we enhance the robustness of sparse data feature extraction using two strategies: on one hand, we utilize Anti-Noise Contour Enhancement (ANCE) module to improve motion contour features and mitigate sensor pulse noise; on the other hand, we implement Direction-Aware Sparse Perception (DASP) module to encourage the model to extract robust person descriptors. Results on the Event-ReID-v1 and Event-ReID-v2 datasets validate the effectiveness of our approach.
Wenjiao Dong, Xi Yang 0011, Nannan Wang 0001
IEEE Trans. Inf. Forensics Secur.2
2026 ALIGNER: Learning Fine-Grained Cross-Modal Alignment for Text-Based Person Retrieval
Xi Yang 0011, Nannan Wang 0001
IEEE Trans. Inf. Forensics Secur.1
2026 Toward Semantically Enhanced Representation Learning for Text-Based Person Retrieval
abstract
Text-Based Person Retrieval (TBPR), which is a pivotal technology in the intelligent surveillance field, is aimed at retrieving target pedestrians based on free-form textual descriptions. While the existing methods attempt to align cross-modal features via multigranular interactions, their performance remains fundamentally limited by two core challenges: cross-modal semantic inconsistency and cross-modal semantic discriminability. To address these issues, we propose DSEE (Diversity Semantic Embedding Expansion), a novel framework for semantically enhanced representation learning. Unlike approaches that rely on constructing larger or more detailed datasets, DSEE establishes identity-centric cross-modal consistency through contrastive learning and generative synergy. The framework consists of two key modules: Bidirectional-guided Semantic Modeling (BSM) and Generative-driven Semantic Enhancement (GSE) modules. The BSM module constructs novel semantic embeddings by modeling similarity-based interactions between the image and text modalities. Specifically, it emphasizes identity-level similarity to guide the generation of enriched, discriminative semantic representations, thereby enhancing their semantic expressiveness and cross-modal alignment. The GSE module provides enriched semantic diversity through a generative text augmentation scheme based on visual inputs, while refining the semantic precision of the method via a dual-path attention mechanism that performs both intramodal refinement and cross-modal alignment. Extensive experiments demonstrate that DSEE achieves state-of-the-art performance on major benchmarks across diverse scenarios. Our work provides an effective paradigm for advancing TBPR applications in real-world settings.
Xi Yang 0011, Nannan Wang 0001
IEEE Trans. Image Process.2
2026 Toward Universal Semantic Communication via Matchable Semantic Subspace Transmission
abstract
Semantic communication targets reliable task execution at the receiver under stringent bandwidth and channel constraints. However, existing communication paradigms either focus on bit-level signal reconstruction, impeding the balance between task efficacy and bandwidth efficiency, or are limited by fixed vocabularies and lack generalization when facing unknown categories and open scenarios. To this end, we propose Universal Semantic Communication (UniSC), an open-vocabulary semantic communication framework that formulates transmission as a Matchable Semantic Subspace Transmission (MSST) problem. In this work, "universal" refers to the ability to handle arbitrary text-defined semantic categories beyond fixed vocabularies, rather than universality across all vision tasks. The transmitted representation is explicitly constrained to preserve cross-modal matchability after noisy transmission, rather than merely supporting latent recovery or closed-set inference. Concretely, UniSC comprises a Visual Semantic Engine (VSE), a Semantic Squeeze Network (SSN), a Noise-Adaptive Semantic Re-expansion (NASR) module, and a VLM-based Decoder. VSE and SSN project images into a compact semantic subspace for transmission. This subspace is optimized to preserve both robustness and cross-modal matchability under channel corruption. NASR denoises and lifts the received features back into a semantically complete visual space, from which the VLM-based Decoder performs open-category inference by matching arbitrary text queries rather than relying on a fixed classifier head. The VLM-based Decoder employs a Text Semantic Engine (TSE) to map natural language to text embeddings and, via a learnable Text-Visual Bridge (TVB), aligns them with the reconstructed visual structure for cross-modal matching. To improve cross-modal alignment and transmission robustness, a two-stage training strategy first establishes cross-modal anchors and then optimizes end-to-end robustness and compactness. Extensive experiments on semantic segmentation benchmarks demonstrate that UniSC achieves strong generalization and state-of-the-art performance under harsh channel conditions, outperforming existing methods in both low-SNR and extreme-compression regimes.
Xi Yang 0011, Songsong Duan, Nannan Wang 0001
IEEE Trans. Image Process.2
2026 Active Style-Content Dual-Branch Domain Adaptation for Semi-Supervised SAR Object Detection
abstract
Synthetic Aperture Radar (SAR) images offer unique advantages in all-weather, all-day remote sensing, but the high acquisition costs and time-consuming annotation processes limit their widespread implementation. Semi-supervised domain adaptation leverages abundant annotated optical images and a small number of labeled SAR images to achieve great performance on SAR images. However, existing semi-supervised domain adaptation object detection methods typically select SAR domain labeled samples randomly, making it difficult to fully exploit the valuable information and distinctive features inherent in the target domain data. Moreover, there is a significant style and content gap between optical and SAR images, and previous methods have not adapted to them in a task-specific manner. To this end, this paper proposes an active style-content dual-branch domain adaptation method specifically designed for semi-supervised object detection in SAR images. The proposed approach employs Task-aware Active Sampling (TAS) module to select the most valuable SAR samples, addressing inefficiencies in random sampling. Also, we employ a dual-branch framework to address the style and content gaps between optical and SAR images. Multi-layer Feature Alignment (MFA) module ensures style alignment by maintaining consistent feature representations across different visual styles, while Gaussian-SAM Image Fusion (G-SIF) module is employed to integrate content from the source domain into the target domain, effectively bridging the gap between optical and SAR images. Extensive experiments on multiple ship and aircraft datasets demonstrate the exceptional generalization capabilities of our proposed model.
Xi Yang 0011, Quantao Xie, Yirong Yang, Nannan Wang 0001
IEEE Trans. Image Process.1
2026 Overcoming Dual Incremental Challenges in Continual Person Search via Adapter and Prototype
abstract
The advancement of continual person search techniques has seen significant progress in recent years due to its practical applications in the real world. However, continual learning for person search presents significant challenges as it combines both person detection and re-identification (Re-ID) tasks, resulting in issues of domain and class incremental learning. To address these challenges, we propose a novel framework that uses an adapter-based Swin Transformer backbone, and incorporates two key components: Domain Aware Adapter (DAA) blocks and Virtual Prototype Replay-Online Instance Matching (VPR-OIM). Specifically, to solve the domain incremental problem in object detection, we introduce parallel DAA blocks to handle multiple domains, while a Domain Prototype Router (DPR) mechanism is used to dynamically route the feature to the domain-specific adapter. Additionally, for class incremental Re-ID, we extend the OIM loss with virtual prototype replay, which generates Gaussian distribution-based virtual features derived from historical prototypes, effectively enabling the model to preserve knowledge of previous identities while accommodating new identity categories. Overall, our proposed DAA and VPR-OIM simultaneously address the dual incremental challenges of continual person search. Experimental results demonstrate that our method significantly improves both person detection and Re-ID performance in continual learning settings, achieving state-of-the-art (SOTA) performance.
Xi Yang 0011, Hexun Zhou, De Cheng, Nannan Wang 0001
IEEE Trans. Image Process.1
2026 Distribution-Aware Prompt Learning for Vision-Language Models With Dynamic Boundary Prototype
abstract
Prompt learning has emerged as an effective strategy for adapting vision-language models (VLMs) which injects learnable semantic prompts into VLMs to guide the alignment between visual and textual representations. Although existing methods have shown strong performance across various tasks, they usually focus on the representative class-level samples and overlook the atypical and hard samples in visual feature space, which hinders generalization of VLMs. To address this issue, we propose the concept of dynamic boundary prototype, which highlights ambiguous samples that are far from the class centroid and is updated at each epoch. Accordingly, we propose a Distribution-Aware Prompt Learning (DAPL) framework to calibrate the distribution of visual feature space via the definition, optimization, and updating of dynamic boundary prototypes. Firstly, we introduce Boundary-Centroid Pulling to optimize the intra-class distribution by progressively reducing the distance between boundary and centroid prototypes, thereby enhancing structural consistency within each class. Secondly, to further enhance inter-class separability, a distance-weighted contrastive loss that places greater emphasis on distinguishing adjacent classes is designed, facilitating more effective fine-grained discrimination. Thirdly, we apply Low-Rank Adaptation Fine-Tuning to adapt the vision encoder through targeted modifications to its self-attention layers. Additionally, we adopt a progressive training strategy for stable optimization. DAPL is compatible with mainstream prompt learning methods such as CoOp, CoCoOp and PromptKD, and consistently improves their average performance across 11 benchmark datasets.
Xi Yang 0011, Xinyue Zhong, Nannan Wang 0001
IEEE Trans. Image Process.1
2026 Prompt-Driven Knowledge Distillation for Remote Sensing Object Detection
abstract
Remote sensing object detection requires precise identification of multi-scale and multi-directional targets in com-plex backgrounds, demanding the model that achieves both high accuracy and real-time performance. While knowledge distillation proves effective for compressing natural image models, it exhibits limitations in more realistic remote sensing scenarios, including inadequate adaptability, biases from long-tail data distributions, and the propagation of errors from the teacher model. To address these challenges, we propose a Prompt Driven Knowledge Distillation (PDKD) framework for remote sensing object detection. This framework leverages prompt-based mechanisms to guide the student model in effectively acquiring and assimilating the teacher's knowledge, which integrates three core components: (1) Scale-Decoupled Feature Prompting (SDFP) module dynamically adjusts feature representation capabilities through scale decoupling, enabling differentiated distillation for targets of varying scales; (2) Semantic Visual Co-Prompting (SVCP) module, based on CLIP's multimodal prior knowledge, constructs category-specific semantic prompt vectors to enhance the focus on features of long-tail categories; (3) Self-Correcting Prompting (SCP) module that suppresses error propagation through a cross self-distillation mechanism. The experiments on the DOTA dataset show that with a 1x training schedule, the model achieves a 49.0% $mAP$ . Source codes are available at https://github.com/Ningsui/PDKD.git.
Xi Yang 0011, Nannan Wang 0001
IEEE Trans. Image Process.1
2025 Dual Information Purification for Lightweight SAR Object Detection
abstract
Synthetic aperture radar (SAR) object detection requires accurate identification and localization of targets at various scales within SAR images. However, background clutter and speckle noise can obscure key features and mislead the knowledge distillation process. To address these challenges, we introduce the Dual Information Purification Knowledge Distillation (DIPKD) method, which improves the performance of the student model through three key strategies: denoising, enrichment, and decoupling. First, our Selective Noise Suppression (SNS) technique reduces speckle noise in global features by minimizing misleading information from the teacher model. Second, the Knowledge Level Decoupling (KLD) module separates features into target and non-target knowledge, balancing feature mapping and reducing background noise to enhance the extraction of critical information for the student model. Finally, the Reverse Information Transfer (RIT) module refines intermediate features in the student model, compensating for the loss of detailed local information. Experimental results demonstrate that DIPKD significantly outperforms existing distillation techniques in SAR object detection, achieving 60.2% and 51.4% mAP scores on the SSDD and HRSID datasets, respectively. Additionally, the student model shows performance improvements of 1.3% and 2.9% over the teacher model, highlighting the effectiveness of the information purification approach.
Xi Yang 0011, Songsong Duan, De Cheng
AAAI1
2025 Optimizing Label Assignment for Weakly Supervised Person Search
abstract
Weakly supervised person search aims to detect and match individuals using only bounding box annotations jointly. The existing methods mainly alternate between the clustering stage and the training stage, where the former is responsible for instance level label allocation tasks and the latter needs to undertake proposal level label allocation tasks. In the clustering phase, the conventional use of the DBSCAN algorithm for clustering pedestrian instance features often neglects key contextual information such as scene context and relative positioning of individuals. During the training phase, the Region Proposal Network assigns labels based on the MaxIoU, which tends to produce locally ambiguous labels. Finally, the proposals updated to the memory bank with extensive background information tend to interfere with the task of pseudo-label generation. To address these issues, this paper proposes an Optimizing Label Assignment (OLA) for weakly supervised person search. Firstly, in the clustering phase, Context Aware Clustering is introduced to integrate contextual information and constraints, enhancing the accuracy of clustering. Secondly, in the training phase, we adopt Prototype Matching based on Optimal Transport theory to optimize label distribution from a global perspective. Furthermore, we propose Dual Memory Bank Enhancement that effectively enhances the accuracy of label assignment. Extensive experiments conducted on the CUHK-SYSU and PRW datasets demonstrate that our method achieves state-of-the-art performance in weakly supervised person search.
Xi Yang 0011, Nannan Wang 0001
AAAI2
2025 Multi-Label Prototype Visual Spatial Search for Weakly Supervised Semantic Segmentation
abstract
Existing Weakly Supervised Semantic Segmentation (WSSS) relies on the CNN-based Class Activation Map (CAM) and Transformer-based self-attention map to generate class-specific masks for semantic segmentation. However, CAM and self-attention maps usually cause incomplete segmentation due to classification bias issue. To address this issue, we propose a Multi-Label Prototype Visual Spatial Search (MuP-VSS) method with a spatial query mechanism. Specifically, MuP-VSS consists of two key components: multi-label prototype representation and multi-label prototype optimization. The former designs a global embedding to learn the global tokens from the images, and then proposes a Prototype Embedding Module (PEM) to interact with patch tokens to understand the local semantic information. The latter utilizes the exclusivity and consistency principles of the multi-label prototypes to design three prototype losses to optimize them, which contain cross-class prototype (CCP) contrastive loss, cross-image prototype (CIP) contrastive loss, and patch-to-prototype (P2P) consistency loss. CCP loss models exclusivity of multi-label prototypes learned from a single image to enhance the discriminative properties of each class better. CCP loss learns the consistency of the same class-specific prototypes extracted from multiple images to enhance the semantic consistency. P2P loss is proposed to control the semantic response of the prototype to the image patches. Experimental results on Pascal VOC 2012 and MS COCO show that MuP-VSS significantly outperforms recent methods and achieves state-of-the-art performance.
Songsong Duan, Xi Yang 0011, Nannan Wang 0001
CVPR2
2025 풟ℐℋ-CLIP: Unleashing the Diversity of Multi-Head Self-Attention for Training-Free Open-Vocabulary Semantic Segmentation
Songsong Duan, Xi Yang 0011, Nannan Wang 0001
ICCV2
2025 Dual Domain Control via Active Learning for Remote Sensing Domain Incremental Object Detection
De Cheng, Xi Yang 0011, Nannan Wang 0001
ICCV3
2025 Surrogate Prompt Learning: Towards Efficient and Diverse Prompt Learning for Vision-Language Models
abstract
Prompt learning is a cutting-edge parameter-efficient fine-tuning technique for pre-trained vision-language models (VLMs). Instead of learning a single text prompt, recent works have revealed that learning diverse text prompts can effectively boost the performances on downstream tasks, as the diverse prompted text features can comprehensively depict the visual concepts from different perspectives. However, diverse prompt learning demands enormous computational resources. This efficiency issue still remains unexplored. To achieve efficient and diverse prompt learning, this paper proposes a novel Surrogate Prompt Learning (SurPL) framework. Instead of learning diverse text prompts, SurPL directly generates the desired prompted text features via a lightweight Surrogate Feature Generator (SFG), thereby avoiding the complex gradient computation procedure of conventional diverse prompt learning. Concretely, based on a basic prompted text feature, SFG can directly and efficiently generate diverse prompted features according to different pre-defined conditional signals. Extensive experiments indicate the effectiveness of the surrogate prompted text features, and show compelling performances and efficiency of SurPL on various benchmarks.
Liangchen Liu 0001, Nannan Wang 0001, Xi Yang 0011, Xinbo Gao 0001, Tongliang Liu
ICML3
2025 Fooling human detectors via robust and visually natural adversarial patches
Dawei Zhou 0004, Hongbin Qu, Nannan Wang 0001, Chunlei Peng, Zhuoqi Ma, Xi Yang 0011, Xinbo Gao 0001
Neurocomputing6
2025 Frequency-Based Comprehensive Prompt Learning for Vision-Language Models
abstract
This paper targets to learn multiple comprehensive text prompts that can describe the visual concepts from coarse to fine, thereby endowing pre-trained VLMs with better transfer ability to various downstream tasks. We focus on exploring this idea on transformer-based VLMs since this kind of architecture achieves more compelling performances than CNN-based ones. Unfortunately, unlike CNNs, the transformer-based visual encoder of pre-trained VLMs cannot naturally provide discriminative and representative local visual information. To solve this problem, we propose Frequency-based Comprehensive Prompt Learning (FCPrompt) to excavate representative local visual information from the redundant output features of the visual encoder. FCPrompt transforms these features into frequency domain via Discrete Cosine Transform (DCT). Taking the advantages of energy concentration and information orthogonality of DCT, we can obtain compact, informative and disentangled local visual information by leveraging specific frequency components of the transformed frequency features. To better fit with transformer architectures, FCPrompt further adopts and optimizes different text prompts to respectively align with the global and frequency-based local visual information via a dual-branch framework. Finally, the learned text prompts can thus describe the entire visual concepts from coarse to fine comprehensively. Extensive experiments indicate that FCPrompt achieves the state-of-the-art performances on various benchmarks.
Liangchen Liu 0001, Nannan Wang 0001, Chen Chen 0128, Decheng Liu, Xi Yang 0011, Xinbo Gao 0001, Tongliang Liu
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Bidirectional modality information interaction for Visible-Infrared Person Re-identification
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
Pattern Recognit.1
2025 Associative graph convolution network for point cloud analysis
Xi Yang 0011, Xingyilang Yin, Nannan Wang 0001, Xinbo Gao 0001
Pattern Recognit.1
2025 Perspectives of Calibrated Adaptation for Few-Shot Cross-Domain Classification
abstract
Current few-shot learning techniques predominantly leverage amortization techniques based on meta-learning frameworks, which effectively adapt to unknown tasks with limited examples. However, these approaches face significant challenges in cross-domain scenarios, where the data distributions between the source domain (training data) and the target domain (testing data) differ substantially. This domain shift can lead to models that overfit the global discriminative model while underfitting their local amortization on the adaptable few-shot structure. To mitigate this problem, our proposal makes an upgrade on Conditional Neural Adaptive Processes, reformulating its conditioning mechanism to better handle cross-domain adaptation. This results in calibrated amortization of task-specific feature extractors and the construction of a robust non-parametric classifier. In our implementation, we first employ generative modeling or deterministic self-attention to all labeled context features, establishing a strong task-level alignment that adapts the extractor across domains. Additionally, we introduce a novel channel-wise normalization to further enhance the adaptation process. Our experiments on the Meta-dataset benchmark demonstrate an average$6.9\sim 9$% improvement in out-of-distribution tasks, underscoring the effectiveness of exploiting calibrated adaptation in few-shot cross-domain classification.
Dechen Kong, Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Elaborate Information Refinement Network for Fine-Grained Object Detection in Remote Sensing Images
abstract
0pt Fine-grained object detection aims to localize and classify subcategories of objects by extracting more discriminative semantic features, which is particularly challenging for remote sensing images due to their complex backgrounds and arbitrarily oriented objects. Accurate localization with bounding box regression typically relies on detailed texture and edge information to delineate object boundaries, while fine-grained classification requires more elaborate semantic information. However, existing methods often share the same input features across the model, resulting in a mismatch between the requirements of the localization task and those of the fine-grained classification task. To address this problem, we propose a Elaborate Information Refinement Network (EIRNet), which not only effectively separates features for localization and fine-grained classification but also refines these features according to the specific requirements of each task. For fine-grained classification, we propose a Fine-grained Context Fusion Module (FCFM) to enhance the ability to extract discriminative features by expanding the receptive field. For localization, we introduce an Edge Information Sensing Module (EISM) to extract scale-invariant features by combining high-dimensional information with detailed edge information, thereby improving the network’s ability to accurately locate objects. Additionally, to extract richer fine-grained semantic information, we present a Feature Injection Module (FIM), which collects and fuses information across different levels using a unified structure. This module then distributes the refined features to appropriate levels, effectively reducing inherent information loss and enhancing the local information fusion capability of the network. Through extensive experiments, the proposed method demonstrates a mean average precision (mAP) of 80.23% on ShipRSImageNet and 50.52% on FAIR1M-v1.0, marking improvements of 4.10% and 2.54% over the current state-of-the-art approaches, respectively.
Xi Yang 0011, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.2
2025 Scale-Consistent Learnable PnP Network for Space Target Pose Estimation
abstract
Due to constraints from imaging devices, the most effective method for estimating the pose of space target RGB images is to establish a 2-D–3-D correspondence and then use the perspective-n-point (PnP) algorithm for pose recovery. However, traditional PnP algorithms are not differentiable, which hinders their integration with neural network training. Although recent work attempts to make PnP partially differentiable during the 2-D–3-D matching stage by mathematical methods, this leads to increased inevitably computational costs. To this end, we propose a scale-consistent learnable PnP (SCLP) network that facilitates end-to-end pose estimation for space targets. Our method incorporates a sparse keypoint learnable PnP (SKL-PnP) layer within a multiscale network, enabling PnP to function as a differentiable layer that integrates seamlessly with preceding neural components. Additionally, we also sample the 2-D–3-D correspondences to obtain sparse keypoint pairs, achieving a lightweight single-stage 6-D pose estimation algorithm. To manage the significant scale variations in space target images, we introduce Gaussian perception sampling (GPS) by assigning instances to different pyramid levels based on size. Furthermore, we propose a scale consistency regularization (SCR) module that aligns downsized feature maps with original ones to better address scale differences. Experimental results demonstrate that our approach achieves superior accuracy and efficiency on the SPEED and SwissCube datasets, showing significant improvements over state-of-the-art methods.
Xi Yang 0011, Jingyuan Wang 0002, Songsong Duan
IEEE Trans. Geosci. Remote. Sens.1
2025 SCIR: A Weakly Supervised Contextual Instance Refinement Method for Remote Sensing Object Detection
abstract
Most weakly supervised object detection (WSOD) methods currently prioritize the top-scoring object instance from proposals to train the corresponding object detector. The detector tends to focus on the entire object by analyzing the contextual information around the most discriminative activation regions. However, the traditional selective search algorithm yields top-scoring instance proposals that cover only a portion of the object, thereby diminishing the detector’s sensitivity to objects with a wide range of scale variations in remote sensing images (RSIs), particularly tiny object clusters. To address this issue, this paper proposes a novel WSOD method called SAM-guided proposal generation with Contextual Instance Refinement (SCIR) method for detecting objects with significant scale variations in RSIs. Specifically, a SAM-guided proposal generator (SPG) module is designed to generate high-quality proposals based on SAM masks rather than the traditional top-scoring one. Our SPG combines WSOD’s advantage of mining classification clues through inexact supervision with SAM’s capability of pre-learned world knowledge to provide automatic prompts for SAM. Meanwhile, we propose a contextual enhanced feature extractor (CEFE) module to capture the global context of the visual scene, which further activates the feature representation of the entire object. Finally, the feature map from CEFE and proposals generated by SPG are fed into a context-perceived instance refinement (CPIR) module. Our CPIR aims to shift the attention of the detection network from the local feature portion to the entire object by integrating local and global contextual information. Extensive experiments on the challenging DOTA and DIOR datasets demonstrate that our proposed SCIR achieves state-of-the-art performance and is quite effective on multi-scale object issues.
Xi Yang 0011, Zhongyuan Zhou, Songsong Duan, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.1
2025 SFDN: A Novel Semantic Feature Decouple Network for Fine-Grained Remote Sensing Object Detection
abstract
Fine-grained object detection (FGOD) aims to identify subcategories of objects by extracting more discriminative semantic information. Bounding box regression typically requires detailed texture and edge information to accurately delineate object boundaries, while classification requires richer semantic information. However, existing methods use the same input features in the model, resulting in an imbalance between the localization task and the fine-grained classification task. To address this issue, we propose a novel Semantic Feature Decouple Network (SFDN) that effectively separates semantic information for localization and fine-grained classification. For the localization task, we propose a Regression Feature Fusion Module (RFFM) to extract feature maps with more edge information. To enhance classification one, we propose a Fine-grained Feature Diversification Module (FFDM) to capture discriminative semantic information from feature maps by introducing fake attention maps. Aiming to extract richer fine-grained semantic information, we propose an Adaptive Local Perception Module (ALPM) to deeply extract multiscale semantic feature information by using dilated convolution at varying dilation rates. Extensive experiments demonstrate that the proposed network respectively achieves the mean average precision of 80.08% and 49.44% on the ShipRSImageNet and FAIR1M-v1.0 datasets, outperforming SOTA methods by 3.95% and 1.46%.
Xi Yang 0011, Zhongyuan Zhou, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.1
2025 Granularity-Aware Hyperbolic Representation for Text-Based Person Search
abstract
Text-based person search aims to identify specific target person from the database according to the given text description. Early work adopted separately pretrained encoders to extract visual and textual features, but benefit from the bloom of visual language pre-training, recent work uses unified pretrained visual language models such as CLIP as backbone. However, visual language models are generally pretrained from coarse-grained image-text pairs, while image-text pairs in text-based person search are more fine-grained to distinguish different persons. In addition, visual and linguistic concepts naturally organize themselves in a hierarchy, which is not explicitly captured by current large-scale vision and language models such as CLIP. To bridge this gap, we propose a novel Granularity-Aware Hyperbolic Representation learning method for mining granularity and capturing semantic hierarchy. Notably, we consider both token-level and instance-level granularity. For token-granularity alignment, we present a Bidirectional Attention Interaction module to explicitly learn the matching between fine-grained visual tokens and text tokens. For instance-granularity alignment, we equip the contrastive learning loss with Semantic Margin Softmax so that image-text pairs can perceive the similarity granularity of different samples during training. Besides, the global features of images and texts are mapped into hyperbolic space through Hyperbolic Representation Learning to embed tree-like data to capture semantic hierarchy. Extensive experiments verify the effectiveness of our proposed modules and show that our method achieves state-of-the-art results on the three widely acknowledged benchmarks, namely CUHK-PEDES, ICFG-PEDES, and RST-PReID. Our code is available at https://github.com/7chQ/GAHR.
Chenghuan Qi, Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.2
2025 Escaping Modal Interactions: An Efficient DESANet for Multi-Modal Object Re-Identification
abstract
Multi-modal object Re-ID aims to leverage the complementary information provided by multiple modalities to overcome challenging conditions and achieve high-quality object matching. However, existing multi-modal methods typically rely on various modality interaction modules for information fusion, which can reduce the efficiency of real-time monitoring systems. Additionally, practical challenges such as low-quality multi-modal data or missing modalities further complicate the application of object Re-ID. To address these issues, we propose the Complementary Data Enhancement and Modal-Aware Soft Alignment Network (DESANet), which is designed to be independent of interactive networks and adaptable to scenarios with missing modalities. This approach ensures a simple-yet-effective, and efficient multi-modal object Re-ID. DESANet consists of three key components: Firstly, the Dual-Color Space Data Enhancement (DCDE) module, which enhances multi-modal data by performing patch rotation in the RGB space and improving image quality in the HSV space. Secondly, the Salient Feature ReConstruction (SFRC) module, which addresses the issue of missing modalities by reconstructing features from one modality using the other two. Thirdly, the Modal-Aware Soft Alignment (MASA) module, which integrates multi-source data to avoid the blind fusion of features and prevents the propagation of noise from reconstructed modalities. Our approach achieves state-of-the-art performances on both person and vehicle datasets. Source code is available at https://github.com/DWJ11/DESANet.
Wenjiao Dong, Xi Yang 0011, De Cheng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.2
2025 Lightweight RGB-D Salient Object Detection From a Speed-Accuracy Tradeoff Perspective
abstract
Current RGB-D methods usually leverage large-scale backbones to improve accuracy but sacrifice efficiency. Meanwhile, several existing lightweight methods are difficult to achieve high-precision performance. To balance the efficiency and performance, we propose a Speed-Accuracy Tradeoff Network (SATNet) for Lightweight RGB-D SOD from three fundamental perspectives: depth quality, modality fusion, and feature representation. Concerning depth quality, we introduce the Depth Anything Model to generate high-quality depth maps,which effectively alleviates the multi-modal gaps in the current datasets. For modality fusion, we propose a Decoupled Attention Module (DAM) to explore the consistency within and between modalities. Here, the multi-modal features are decoupled into dual-view feature vectors to project discriminable information of feature maps. For feature representation, we develop a Dual Information Representation Module (DIRM) with a bi-directional inverted framework to enlarge the limited feature space generated by the lightweight backbones. DIRM models texture features and saliency features to enrich feature space, and employ two-way prediction heads to optimal its parameters through a bi-directional backpropagation. Finally, we design a Dual Feature Aggregation Module (DFAM) in the decoder to aggregate texture and saliency features. Extensive experiments on five public RGB-D SOD datasets indicate that the proposed SATNet excels state-of-the-art (SOTA) CNN-based heavyweight models and achieves a lightweight framework with 5.2 M parameters and 415 FPS. The code is available at https://github.com/duan-song/SATNet.
Songsong Duan, Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.2
2025 FA-Net: A Feature Alignment Network for Video-Based Visible-Infrared Person Re-Identification
abstract
Video-based visible-infrared person re-identification (VVI-ReID) aims to match target pedestrians between visible and infrared videos, which is significantly applied in 24-hour surveillance systems. The key of VVI-ReID is to learn modality invariant and spatio-temporal invariant sequence-level representation to solve the challenges such as modality differences, spatio-temporal misalignment, and domain shift noise. However, existing methods predominantly emphasize on reducing modality discrepancy while relatively neglect temporal misalignment and domain shift noise reduction. To this end, this paper proposes a VVI-ReID framework called Feature Alignment Network (FA-Net) from the perspective of feature alignment, aiming to mitigate temporal misalignment. FA-Net comprises two main alignment modules: Spatial-Temporal Alignment Module (STAM) and Modality Distribution Constraint (MDC). STAM integrates global and local features to ensure individuals' spatial representation alignment. Additionally, STAM also establishes temporal relationships by exploring inter-frame features to address cross-frame person feature matching. Furthermore, we introduce the Modality Distribution Constraint (MDC), which utilizes a symmetric distribution loss to align the distributions of features from different modalities. Besides, the SAM Guidance Augmentation (SAM-GA) strategy is designed to transform the image space of RGB and IR frames to provide more informative and less noisy frame information. Extensive experimental results demonstrate the effectiveness of the proposed method, surpassing existing state-of-the-art methods. Our code will be available at: https://github.com/code/FANet.
Xi Yang 0011, Wenjiao Dong, De Cheng, Nannan Wang 0001
IEEE Trans. Image Process.1
2025 IDENet: An Inter-Domain Equilibrium Network for Unsupervised Cross-Domain Person Re-Identification
abstract
Unsupervised person re-identification aims to retrieve a given pedestrian image from unlabeled data. For training on the unlabeled data, the method of clustering and assigning pseudo-labels has become mainstream, but the pseudo-labels themselves are noisy and will reduce the accuracy. To overcome this problem, several pseudo-label improvement methods have been proposed. But on the one hand, they only use target domain data for fine-tuning and do not make sufficient use of high-quality labeled data in the source domain. On the other hand, they ignore the critical fine-grained features of pedestrians and overfitting problems in the later training period. In this paper, we propose a novel unsupervised cross-domain person re-identification network (IDENet) based on an inter-domain equilibrium structure to improve the quality of pseudo-labels. Specifically, we make full use of both source domain and target domain information and construct a small learning network to equalize label allocation between the two domains. Based on it, we also develop a dynamic neural network with adaptive convolution kernels to generate adaptive residuals for adapting domain-agnostic deep fine-grained features. In addition, we design the network structure based on ordinary differential equations and embed modules to solve the problem of network overfitting. Extensive cross-domain experimental results on Market1501, PersonX, and MSMT17 prove that our proposed method outperforms the state-of-the-art methods.
Xi Yang 0011, Wenjiao Dong, Gu Zheng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2025 Hyperbolic Insights With Knowledge Distillation for Cross-Domain Few-Shot Learning
abstract
Cross-domain few-shot learning aims to achieve swift generalization between a source domain and a target domain using a limited number of images. Current research predominantly relies on generalized feature embeddings, employing metric classifiers in Euclidean space for classification. However, due to existing disparities among different data domains, attaining generalized features in the embedding becomes challenging. Additionally, the rise in data domains leads to high-dimensional Euclidean spaces. To address the above problems, we introduce a cross-domain few-shot learning method named Hyperbolic Insights with Knowledge Distillation (HIKD). By integrating knowledge distillation, it enhances the model's generalization performance, thereby significantly improving task performance. Hyperbolic space, in comparison to Euclidean space, offers a larger capacity and supports the learning of hierarchical structures among images, which can aid generalized learning across different data domains. So we map the Euclidean space features to the hyperbolic space via hyperbolic embedding and utilize hyperbolic fitting distillation method in the meta-training phase to obtain multi-domain unified generalization representation. In the meta-testing phase, accounting for biases between the source and target domains, we present a hyperbolic adaptive module to adjust embedded features and eliminate inter-domain gap. Experiments on the Meta-Dataset demonstrate that HIKD outperforms state-of-the-arts methods with the average accuracy of 80.6%.
Xi Yang 0011, Dechen Kong, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2025 Dense Information Learning Based Semi-Supervised Object Detection
abstract
Semi-Supervised Object Detection (SSOD) aims to improve the utilization of unlabeled data, and various methods, such as adaptive threshold techniques, have been extensively studied to increase exploitable information. However, these methods are passive, relying solely on the original image data. Additionally, existing approaches prioritize the predicted categories of the teacher model while overlooking the relationships between different categories in the prediction. In this paper, we introduce a novel approach called Dense Information Learning (DIL), which actively generates unlabeled data containing densely exploitable information and forces the network to have relation consistency under different perturbations. Specifically, Dense Information Augmentation (DIA) leverages the prior information of the network to create a foreground bank and actively incorporates exploitable information into the unlabeled data. DIA automatically performs information enhancement and filters noise. Furthermore, to encourage the network to maintain consistency at the manifold level under various perturbations, we introduce Relation Consistency Regularization (RCR). It considers both feature-level and image-level perturbations, guiding the network to focus on more discriminative features. Extensive experiments conducted on multiple datasets validate the effectiveness of our approach in leveraging information from unlabeled images. The proposed DIL improves the mAP by 12.6% and 10.0% relative to the supervised baseline method when utilizing 5% and 10% of labeled data on the MS-COCO dataset, respectively.
Xi Yang 0011, Qiubai Zhou, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2025 Uncertainty Quantification for Semi-Supervised Object Detection in Remote Sensing Images
abstract
Semi-supervised object detection (SSOD) aims to solve the data annotation challenge in object detection and can achieve remarkable progress in natural scenes; however, it remains unexplored in horizontal bounding box (HBB)-based remote sensing imagery where annotation tasks pose greater challenges. In remote sensing scenarios, objects exhibit arbitrary orientations, small scales, and dense distributions, leading to pseudoboxes with fuzzy boundaries and class imbalance issues. Therefore, we propose UNCertainty quantification (UNC) for SSOD in remote sensing images. UNC uses uncertainty to guide the network from both regression and classification perspectives: Semantic alignment SAM calibration (SASC) uses pseudoboxes as box prompts for the input of the segment anything model (SAM), achieving more precise boundaries. Subsequently, boundaries with lower regression uncertainty are selected as the final pseudoboxes, ensuring better alignment between the pseudoboxes and the ground truth. Dynamic uncertainty weighting (DUW) calculates class uncertainty and determines its correlation with the availability of instances per class. High uncertainty implies limited availability of instances, necessitating greater emphasis on instances of that class. Furthermore, we set a percentage uncertainty threshold to avoid overemphasis caused by individual classes. Extensive experiments conducted on the DIOR and DOTA HBB-based datasets demonstrate the effectiveness of our method in leveraging unlabeled image information. Specifically, compared with the supervised baseline method, the UNC method improves mAP by 12.4% and 8.6% when 5% and 10% of labeled data on DIOR, respectively.
Xi Yang 0011, Qiubai Zhou, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2025 S3OIL: Semi-Supervised SAR-to-Optical Image Translation via Multi-Scale and Cross-Set Matching
abstract
Image-to-image translation has achieved great success, but still faces the significant challenge of limited paired data, particularly in translatingSynthetic Aperture Radar(SAR) images to optical images. Furthermore, most existing semi-supervised methods place limited emphasis on leveraging the data distribution. To address those challenges, we propose aSemi-Supervised SAR-to-Optical Image Translation(S3OIL) method that achieves high-quality image generation using minimal paired data and extensive unpaired data while strategically exploiting the data distribution. To this end, we first introduce aCross-Set Alignment Matching(CAM) mechanism to create local correspondences between the generated results of paired and unpaired data, ensuring cross-set consistency. In addition, for unpaired data, we apply weak and strong perturbations and establish intra-setMulti-Scale Matching(MSM) constraints. For paired data, intra-modal semantic consistency (ISC) is presented to ensure alignment with the ground truth. Finally, we propose local and global cross-modal semantic consistency (CSC) to boost structural identity during translation. We conduct extensive experiments on SAR-to-optical datasets and another sketch-to-anime task, demonstrating that S3OIL delivers competitive performance compared to state-of-the-art unsupervised, supervised, and semi-supervised methods, both quantitatively and qualitatively. Ablation studies further reveal that S3OIL can ensure the preservation of both semantic content and structural integrity of the generated images. Our code is available at: https://github.com/XduShi/SOIL.
Xi Yang 0011, Ziyun Li 0002, Maoying Qiao, Fei Gao 0006, Nannan Wang 0001
IEEE Trans. Image Process.1
2025 Toward Generalizable Prompt Learning via Multi-Regularization Guided Knowledge Distillation
abstract
Prompt learning has made significant progress in vision-language models (VLMs), enabling pre-trained models like CLIP to perform cross-domain tasks with few-shot or even zero-shot learning. However, existing methods tend to overfit the training data after fine-tuning on the target domain, leading to a decline in generalization ability and limiting their performance on unseen categories.To address these challenges, we propose a multi-regularization guided knowledge distillation towards generalizable prompt learning. This approach enhances the model's adaptability and generalization through different stages of regularization while mitigating performance degradation caused by target domain training. Specifically, within the image encoder of CLIP, we introduce Residual Regularization, which binds additional residual connections to certain transformer blocks. This design provides greater flexibility, allowing the model to adjust to new data distributions when adapting to the target domain.Furthermore, during training, we impose Self-distillation Regularization to ensure that while adapting to the target domain, the model preserves its prior generalization knowledge. Specifically, we regularize the intermediate layer outputs of Transformer Blocks to prevent the model from excessively favoring target domain data. Additionally, we employ an unsupervised knowledge distillation strategy to enforce multi-level alignment between the teacher and student models by Direction Distillation Regularization. This ensures that both models maintain consistent visual feature orientations under the same textual features, thereby enhancing overall model stability and cross-domain adaptability.Experimental results demonstrate that our method achieves more stable classification performance in both cross-domain few-shot classification and domain adaptation settings.
Xi Yang 0011, Xinyue Zhong, Dechen Kong, Nannan Wang 0001
IEEE Trans. Image Process.1
2025 Generalizable Prompt Learning via Gradient Constrained Sharpness-Aware Minimization
abstract
This paper targets a novel trade-off problem in generalizable prompt learning for vision-language models (VLM), i.e., improving the performance on unseen classes while maintaining the performance on seen classes. Comparing with existing generalizable methods that neglect the seen classes degradation, the setting of this problem is stricter and fits more closely with practical applications. To solve this problem, we start from the optimization perspective, and leverage the relationship between loss landscape geometry and model generalization ability. By analyzing the loss landscapes of the state-of-the-art method and vanilla Sharpness-aware Minimization (SAM) based method, we conclude that the trade-off performance correlates to bothloss valueandloss sharpness, while each of them is indispensable. However, we find the optimizing gradient of existing methods cannot maintain high relevance to both loss value and loss sharpness during optimization, which severely affects their trade-off performance. To this end, we propose a novel SAM-based method for prompt learning, denoted as Gradient Constrained Sharpness-aware Context Optimization (GCSCoOp), to dynamically constrain the optimizing gradient, thus achieving above two-fold optimization objective simultaneously. Extensive experiments verify the effectiveness of GCSCoOp in the trade-off problem.
Liangchen Liu 0001, Nannan Wang 0001, Dawei Zhou 0004, Decheng Liu, Xi Yang 0011, Xinbo Gao 0001, Tongliang Liu
IEEE Trans. Multim.5
2025 TIENet: A Tri-Interaction Enhancement Network for Multimodal Person Reidentification
abstract
Multimodal person reidentification (ReID), which aims to learn modality-complementary information by utilizing multimodal images simultaneously for person retrieval, is crucial for achieving all-time and all-weather monitoring. Existing methods try to address this issue through modality fusion to absorb complementary information. However, most of these methods are limited to the spatial domain only and usually overlook the intra-/intermodal interactions during feature fusion, resulting in insufficient learning of modality-specific and complementary information. To address these issues, we propose a tri-interaction enhancement network (TIENet), which contains three modules: spatial-frequency interaction (SFI), intermodal mask interaction (IMMI), and intramodal feature fusion (IMFF). Specifically, the SFI boosts the modality-specific representation by integrating the amplitude-guided attention mechanism into the phase space, combined with spatial-domain convolution to achieve fine-grained information learning. Meanwhile, the IMMI enhances the richness of the feature descriptors by embedding the intermodal relationships to preserve complementary information. Finally, the IMFF module considers the structure of the human body and integrates intramodal contextual information. Extensive experimental results demonstrate the effectiveness of our method, achieving superior performances on RGBNT201 and MARKET1501_RGBNT datasets.
Xi Yang 0011, Wenjiao Dong, De Cheng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Point Deformable Network with Enhanced Normal Embedding for Point Cloud Analysis
abstract
Recently MLP-based methods have shown strong performance in point cloud analysis. Simple MLP architectures are able to learn geometric features in local point groups yet fail to model long-range dependencies directly. In this paper, we propose Point Deformable Network (PDNet), a concise MLP-based network that can capture long-range relations with strong representation ability. Specifically, we put forward Point Deformable Aggregation Module (PDAM) to improve representation capability in both long-range dependency and adaptive aggregation among points. For each query point, PDAM aggregates information from deformable reference points rather than points in limited local areas. The deformable reference points are generated data-dependent, and we initialize them according to the input point positions. Additional offsets and modulation scalars are learned on the whole point features, which shift the deformable reference points to the regions of interest. We also suggest estimating the normal vector for point clouds and applying Enhanced Normal Embedding (ENE) to the geometric extractors to improve the representation ability of single-point. Extensive experiments and ablation studies on various benchmarks demonstrate the effectiveness and superiority of our PDNet.
Xingyilang Yin, Xi Yang 0011, Liangchen Liu 0001, Nannan Wang 0001, Xinbo Gao 0001
AAAI2
2024 Pro2SAM: Mask Prompt to SAM with Grid Points for Weakly Supervised Object Localization
Xi Yang 0011, Songsong Duan, Nannan Wang 0001, Xinbo Gao 0001
ECCV (69)1
2024 Information Fusion with Knowledge Distillation for Fine-grained Remote Sensing Object Detection
abstract
Fine-grained remote sensing object detection aims to locate and identify specific targets with variable scale and orientation from complex background in the high-resolution and wide-swath images, which needs requirement of high precision and real-time processing simultaneously. Although traditional knowledge distillation technology show its effectiveness in model compression and accuracy preservation for natural images, the challenges of heavy background noise and intra-class similarity faced by remote sensing images limits the knowledge quality of teacher model and the learning ability of student model. To address these issues, we propose the Information Fusion with Knowledge Distillation (IFKD) method to enhance student model performance by integrating information from external images, frequency domain, and hyperbolic space. This includes three key modules: 1) External Disturbance Enhancement (EDE), which uses MobileSAM to enrich teachers' knowledge and reduce students' dependency on teachers; 2) Frequency Domain Reconstruction (FDR) to amplify key feature representations and reduce background noise interference by resampling low-frequency information; 3) Hyperbolic Similarity Mask (HSM) to increase intra-class differences, guiding students in analyzing and utilizing teachers' knowledge, and leveraging the exponential capabilities of hyperbolic space for performance improvement. Experimental results verify that the IFKD method significantly enhances performance in fine-grained recognition tasks compared to existing distillation techniques. Specially, 65.8% and 81.4% Ap_50 have achieved on optical ShipRSImageNet and SAR Aircraft-1.0 with our method, even which is 0.4% and 4.7% higher than the teacher.
Xi Yang 0011
ACM Multimedia2
2024 SA3DIP: Segment Any 3D Instance with Potential 3D Priors
abstract
The proliferation of 2D foundation models has sparked research into adapting them for open-world 3D instance segmentation. Recent methods introduce a paradigm that leverages superpoints as geometric primitives and incorporates 2D multi-view masks from Segment Anything model (SAM) as merging guidance, achieving outstanding zero-shot instance segmentation results. However, the limited use of 3D priors restricts the segmentation performance. Previous methods calculate the 3D superpoints solely based on estimated normal from spatial coordinates, resulting in under-segmentation for instances with similar geometry. Besides, the heavy reliance on SAM and hand-crafted algorithms in 2D space suffers from over-segmentation due to SAM's inherent part-level segmentation tendency. To address these issues, we propose SA3DIP, a novel method for Segmenting Any 3D Instances via exploiting potential 3D Priors. Specifically, on one hand, we generate complementary 3D primitives based on both geometric and textural priors, which reduces the initial errors that accumulate in subsequent procedures. On the other hand, we introduce supplemental constraints from the 3D space by using a 3D detector to guide a further merging process. Furthermore, we notice a considerable portion of low-quality ground truth annotations in ScanNetV2 benchmark, which affect the fair evaluations. Thus, we present ScanNetV2-INS with complete ground truth labels and supplement additional instances for 3D class-agnostic instance segmentation. Experimental evaluations on various 2D-3D datasets demonstrate the effectiveness and robustness of our approach. Our code and proposed ScanNetV2-INS dataset are available HERE.
Xi Yang 0011, Xu Gu 0002, Xingyilang Yin, Xinbo Gao 0001
NeurIPS1
2024 Feature-Level Adversarial Attacks and Ranking Disruption for Visible-Infrared Person Re-identification
abstract
Visible-infrared person re-identification (VIReID) is widely used in fields such as video surveillance and intelligent transportation, imposing higher demands on model security. In practice, the adversarial attacks based on VIReID aim to disrupt output ranking and quantify the security risks of models. Although numerous studies have been emerged on adversarial attacks and defenses in fields such as face recognition, person re-identification, and pedestrian detection, there is currently a lack of research on the security of VIReID systems. To this end, we propose to explore the vulnerabilities of VIReID systems and prevent potential serious losses due to insecurity. Compared to research on single-modality ReID, adversarial feature alignment and modality differences need to be particularly emphasized. Thus, we advocate for feature-level adversarial attacks to disrupt the output rankings of VIReID systems. To obtain adversarial features, we introduce \textit{Universal Adversarial Perturbations} (UAP) to simulate common disturbances in real-world environments. Additionally, we employ a \textit{Frequency-Spatial Attention Module} (FSAM), integrating frequency information extraction and spatial focusing mechanisms, and further emphasize important regional features from different domains on the shared features. This ensures that adversarial features maintain consistency within the feature space. Finally, we employ an \textit{Auxiliary Quadruple Adversarial Loss} to amplify the differences between modalities, thereby improving the distinction and recognition of features between visible and infrared images, which causes the system to output incorrect rankings. Extensive experiments on two VIReID benchmarks (i.e., SYSU-MM01, RegDB) and different systems validate the effectiveness of our method.
Xi Yang 0011, De Cheng, Nannan Wang 0001, Xinbo Gao 0001
NeurIPS1
2024 Unconstrained Facial Expression Recognition With No-Reference De-Elements Learning
abstract
Most unconstrained facial expression recognition (FER) methods take original facial images as inputs to learn discriminative features by well-designed loss functions, which cannot reflect important visual information in faces. Although existing methods have explored the visual information of constrained facial expressions, there is no explicit modeling of what visual information is important for unconstrained FER. To find out valuable information of unconstrained facial expressions, we pose a new problem of no-reference de-elements learning: we decompose any unconstrained facial image into the facial expression element and a neutral face without the reference of corresponding neutral faces. Importantly, the element provides visualization results to understand important facial expression information and improves the discriminative power of features. Moreover, we propose a simple yet effectiveDe-ElementsNetwork (DENet) to learn the element and introduce appropriate constraints to overcome no ground truth of corresponding neutral faces during the de-elements learning. We extensively evaluate the proposed method on in-the-wild FER datasets including RAF-DB, AffectNet, SFEW and FERPlus. The comparable results show that our method is promising to improve classification performance and achieves equivalent performance compared with state-of-the-art methods. Also, we demonstrate the strong generalization performance on realistic occlusion and pose variation datasets and the cross-dataset evaluation.
Hangyu Li 0001, Nannan Wang 0001, Xi Yang 0011, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Affect. Comput.3
2024 Unleashing the Feature Hierarchy Potential: An Efficient Tri-Hybrid Person Search Model
abstract
Person search aims to locate target pedestrians from scene images, involving detection and re-identification. The former seeks to separate the background and focus on the commonality between pedestrians, while the latter aims to identify the target and focus on the difference between pedestrians. To address the paradox of detection and re-identification in search tasks, we propose an efficient Tri-Hybrid person search model utilizing the feature hierarchy design. Our model introduces three feature hybrid models for various feature levels. Before the RoI-Align, we present “Spatial-Channel Hybrid” (SCH) and “Token-Channel Hybrid” (TCH). SCH perceives the boundary frame of pedestrians at multiple scales, thereby enhancing the information disparity between pedestrians and the background and refining the accuracy of the detection frame. TCH uses multi-layer perceptrons (MLP) and blends token and channel features, emphasizing detecting fine-grained semantic information for pedestrians. The interaction of multi-scale perception and fine-grained semantic information enhances the details of detected pedestrians, making them more suitable for similarity measurement in pedestrian matching. After the RoI-Align, we design the “CNN-Transformer Hybrid” to amalgamate global and local features to extract more comprehensive detailed features. Extensive experimental results on CUHK-SYSU and PRW demonstrate the effectiveness of the proposed method over the state-of-the-art performance. Specifically, our method achieves comparable performance on two benchmark datasets, CUHK-SYSU and PRW, with mAP scores of 94.62% and 57.84%, respectively.
Xi Yang 0011, Menghui Tian, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Domain-Aware Generalized Meta-Learning for Space Target Recognition
abstract
As the exploration and utilization of outer space persist, the proliferation of space targets has significantly increased, underscoring the growing importance of space situational awareness. However, space target images encounter numerous challenges, including overexposure, excessive shadowing, star noise, and motion blur, distinct from natural images. While existing models can address specific issues in space target recognition images, their ability for generalizing to unseen data remains relatively weak. Furthermore, the uniform background and minimal interclass differences in space target images impose significant constraints on recognition accuracy. To tackle these challenges, we propose a domain-aware generalized meta-learning for space target recognition. In the meta-training phase, we introduce a distillation module to generalize the prior knowledge of auxiliary domains. This module distills features and predictions from auxiliary domains, providing prior information to develop a model capable of generalization across diverse domains. In the meta-testing phase, the frozen generalized embedding function is connected with a feature bias module to mitigate domain bias issues. Building on the advanced awareness of the space target domain, which is marked by substantial intraclass variations and minimal interclass variations, we introduce a feature refinement module. This module resolves fine-grained issues by reconstructing features and augmenting the proto loss to narrow the intraclass data distance. In practice, our method is evaluated under out-of-distribution settings on the BUAA-SID-share1.0 dataset, achieving an impressive accuracy of 96.0%, surpassing existing space target recognition algorithms.
Xi Yang 0011, Dechen Kong, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.1
2024 An Effective and Lightweight Hybrid Network for Object Detection in Remote Sensing Images
abstract
In recent years, general object detection on nature images leverages convolutional neural networks (CNNs) and vision transformers (ViT) to achieve great progress and high accuracy. However, unlike nature scenarios, remote sensing systems usually deploy a large number of edge devices. Therefore, this condition encourages detectors to have lower parameters than high-complexity neural networks. To this end, we propose an effective and lightweight detection framework of hybrid network, which enhances representation learning to balance efficiency and accuracy of model. Specifically, to compensate low precision caused by lightweight neural networks, we design a boundary-aware context (BAC) module and a frequency self-attention refinement (FSAR) module to improve detector performance in a hybrid structure. The BAC module enhances the local features of the object by fusing the spatial context information with the original image, which not only improves the model accuracy but also effectively solves the problem of multiscale objects. To alleviate the interference of complex background, the FSAR module adopts an adaptive technology to filter out redundant information at different frequencies to improve overall detection performance. The comprehensive experiments on remote sensing datasets, i.e., NWPU VHR-10, LEVIR, and RSOD, indicate that the proposed method achieves state-of-the-art performance and balances between model size and accuracy.
Xi Yang 0011, Songsong Duan
IEEE Trans. Geosci. Remote. Sens.1
2024 Adaptive Mid-Level Feature Attention Learning for Fine-Grained Ship Classification in Optical Remote Sensing Images
abstract
Ship classification in optical remote sensing images is a critical task for various maritime applications, including anti-smuggling, maritime traffic control, and maritime rescue. However, fine-grained ship classification (FGSC) is challenging due to the complex background, intraclass similarity, and interclass difference. In this article, we propose a novel mid-level feature attention learning method for FGSC. Our method incorporates mid-level feature casual attention (MFCA) and mid-level channel attention (MCA) to identify discriminative regions and local features corresponding to subtle visual features. The MFCA constrains the learning process of mid-level features through comparison with attention maps and counterfactual attention maps, while the MCA uses a discriminative component to extract discriminative features from channel information and a diversity component to focus feature channels on more obvious feature regions. Besides, an adaptive weight is added to dynamically adjust the influence of MFCA and MCA in the model. Our method can be trained end-to-end and requires no annotations other than category information. Extensive experiments on two large-scale FGSC datasets, FGSC-23 and FGSCR-42, demonstrate that the proposed method achieves state-of-the-art performance, outperforming existing methods by a significant margin.
Xi Yang 0011, Zilong Zeng, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.1
2024 Two-Way Assistant: A Knowledge Distillation Object Detection Method for Remote Sensing Images
abstract
Due to resource constraints on edge devices, lightweight detection models are gaining popularity in remote sensing. However, achieving efficient performance with these models is challenging compared to traditional detection models. Knowledge distillation (KD) is a promising solution to this issue. However, previous KD methods in remote sensing often suffer from background noise and fail to address feature disparities among detectors. To address the above issues, we introduce the Two-Way Assistant (TWA) distillation method for remote sensing object detection. TWA comprises two crucial modules: the Compression Assistant Module (CPAM) and the Multiscale Adaptive Assistant Module (MAAM). CPAM reduces background information and category interference by compressing and redistributing teacher model features to the student model. MAAM enhances feature knowledge through multi-scale fusion, addressing feature disparities. Through extensive experiments conducted on two distinct types of remote sensing datasets, optical LEVIR and SAR SSDD, our TWA demonstrates favourable performance across both single-stage and two-stage detectors. Especially, it achieves a performance of 82.5% (optical LEVIR dataset) and 95.4% (SAR SSDD dataset) in theAP50metric, superior to existing state-of-the-art methods.
Xi Yang 0011
IEEE Trans. Geosci. Remote. Sens.1
2024 Dual-Adversarial Representation Disentanglement for Visible Infrared Person Re-Identification
abstract
Heterogeneous pedestrian images are captured by visible and infrared cameras with different spectrums, which play an important role in night-time video surveillance. However, visible infrared person re-identification (VI-REID) is still a challenging problem due to the considerable cross-modality discrepancies. To extract modality-invariant features which are discriminative for the person identity, recent studies are inclined to regard modality-specific features as noise and discard them. Actually, the modality-specific characteristics containing background and color information are indispensable for learning modality-shared features. In this paper, we propose a novel Dual-Adversarial Representation Disentanglement (DARD) model to separate modality-specific features from tangled pedestrian representations and effectively learn the robust modality-invariant representations. Specifically, our method employs dual-adversarial learning, incorporating image-level channel exchange and feature-level magnitude change to introduce variations in modality-specific representations. This deliberate perturbation raises the learning difficulty for the model to learn modality-shared features. Simultaneously, to control the changing scope of modality-specific features, bi-constrained noise alleviation is introduced during adversarial learning, keeping the balance of feature generation and adversary. The proposed dual-adversarial learning methodology enhances the robustness against cross-modality visual discrepancy and strengthens the discriminative power of the learned modality-shared representations without introducing additional network parameters. This improvement further elevates the retrieval performance of VI-REID. Extensive experiments with insightful analysis on two cross-modality re-identification datasets verify the effectiveness and superiority of the proposed DARD method.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.2
2024 Semi-Supervised Learning With Heterogeneous Distribution Consistency for Visible Infrared Person Re-Identification
abstract
Visible infrared person re-identification (VI-ReID) exposes considerable challenges because of the modality gaps between the person images captured by daytime visible cameras and nighttime infrared cameras. Several fully-supervised VI-ReID methods have improved the performance with extensive labeled heterogeneous images. However, the identity of the person is difficult to obtain in real-world situations, especially at night. Limited known identities and large modality discrepancies impede the effectiveness of the model to a great extent. In this paper, we propose a novel Semi-Supervised Learning framework with Heterogeneous Distribution Consistency (HDC-SSL) for VI-ReID. Specifically, through investigating the confidence distribution of heterogeneous images, we introduce a Gaussian Mixture Model-based Pseudo Labeling (GMM-PL) method, which adaptively adjusts different thresholds for each modality to label the identity. Moreover, to facilitate the representation learning of unutilized data whose prediction is lower than the threshold, Modality Consistency Regularization (MCR) is proposed to ensure the prediction consistency of the cross-modality pedestrian images and handle the modality variance. Extensive experiments with different label settings on two VI-ReID datasets demonstrate the effectiveness of our method. Particularly, HDC-SSL achieves competitive performance with state-of-the-art fully-supervised VI-ReID methods on RegDB dataset with only 1 visible label and 1 infrared label per class.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.2
2024 Adapting Few-Shot Classification via In-Process Defense
abstract
Most few-shot learning methods employ either adaptive approaches or parameter amortization techniques. However, their reliance on pre-trained models presents a significant vulnerability. When an attacker's trigger activates a hidden backdoor, it may result in the misclassification of images, profoundly affecting the model's performance. In our research, we explore adaptive defenses against backdoor attacks for few-shot learning. We introduce a specialized stochastic process tailored to task characteristics that safeguards the classification model against attack-induced incorrect feature extraction. This process functions during forward propagation and is thus termed an "in-process defense." Our method employs an adaptive strategy, effectively generating task-level representations, enabling rapid adaptation to pre-trained models, and proving effective in few-shot classification scenarios for countering backdoor attacks. We apply latent stochastic processes to approximate task distributions and derive task-level representations from the support set. This task-level representation guides feature extraction, leading to backdoor trigger mismatching and forming the foundation of our parameter defense strategy. Benchmark tests on Meta-Dataset reveal that our approach not only withstands backdoor attacks but also shows an improved adaptation in addressing few-shot classification tasks.
Xi Yang 0011, Dechen Kong, Ren Lin, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2024 Image-Level Adaptive Adversarial Ranking for Person Re-Identification
abstract
The potential vulnerability of deep neural networks and the complexity of pedestrian images, greatly limits the application of person re-identification techniques in the field of smart security. Current attack methods often focus on generating carefully crafted adversarial samples or only disrupting the metric distances between targets and similar pedestrians. However, both aspects are crucial for evaluating the security of methods adapted for person re-identification tasks. For this reason, we propose an image-level adaptive adversarial ranking method that comprehensively considers two aspects to adapt to changes in pedestrians in the real world and effectively evaluate the robustness of models in adversarial environments. To generate more refined adversarial samples, our image representation enhancement module leverages channel-wise information entropy, assigning varying weights to different channels to produce images with richer information content, along with a generative adversarial network to create adversarial samples. Subsequently, for adaptive perturbation of ranking, the adaptive weight confusion ranking loss is presented to calculate the weights of distances between positive or negative samples and query samples. It endeavors to push positive samples away from query samples and bring negative samples closer, thereby interfering with the ranking of system. Notably, this method requires no additional hyperparameter tuning or extra data training, making it an adaptive attack strategy. Experimental results on large-scale datasets such as Market1501, CUHK03, and DukeMTMC demonstrate the effectiveness of our method in attacking ReID systems.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2024 Towards Specific Domain Prompt Learning via Improved Text Label Optimization
abstract
Prompt learning has emerged as a thriving parameter-efficient fine-tuning technique for adapting pre-trained vision-language models (VLMs) to various downstream tasks. However, existing prompt learning approaches still exhibit limited capability for adapting foundational VLMs to specific domains that require specialized and expert-level knowledge. Since this kind of specific knowledge is primarily embedded in the pre-defined text labels, we infer that foundational VLMs cannot directly interpret semantic meaningful information from these specific text labels, which causes the above limitation. From this perspective, this paper additionally models text labels with learnable tokens and casts this operation into traditional prompt learning framework. By optimizing label tokens, semantic meaningful text labels are automatically learned for each class. Nevertheless, directly optimizing text label still remains two critical problems, i.e., insufficient optimization and biased optimization. We further address these problems by proposing Modality Interaction Text Label Optimization (MITLOp) and Color-based Consistency Augmentation (CCAug) respectively, thereby effectively improving the quality of the optimized text labels. Extensive experiments indicate that our proposed method achieves significant improvements in VLM adaptation on specific domains.
Liangchen Liu 0001, Nannan Wang 0001, Decheng Liu, Xi Yang 0011, Xinbo Gao 0001, Tongliang Liu
IEEE Trans. Multim.4
2024 Cooperative Separation of Modality Shared-Specific Features for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) is a challenging task because the different imaging principles of visible and infrared images bring about huge modality discrepancy. Existing methods primarily address this issue by generating intermediate images to align modality features and establish connections between the visible and infrared modalities. However, the quality of these generated images is often unstable, limiting the effectiveness of such approaches. To overcome this limitation, we propose a novel method called modality shared-specific features cooperative separation. It consists of two key modules: the saliency response module and the cooperative separation module, aimed at alleviating the modality gap. The saliency response module incorporates a location attention mechanism and local features to construct contextual connections and extract local saliency information. Then, the cooperative separation module employs a more concise dual-MLPs as generator to effectively separate shared-specific features. Additionally, we introduce a shared feature refinement mechanism in both the generator and discriminator. By coordinating the shared-specific features, our method achieves secondary separation and extracts purer modality-shared features without specific information. Extensive experiments conducted on the SYSU-MM01 and RegDB public datasets demonstrate that our proposed method performs excellently in VI-ReID.
Xi Yang 0011, Wenjiao Dong, Meijie Li, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.1
2024 SSRR: Structural Semantic Representation Reconstruction for Visible-Infrared Person Re-Identification
abstract
Visible-infrared Person Re-identification (VI-ReID) aims to retrieve the images of pedestrian with the same identity from different modalities and cameras given a pedestrian image. To reduce modality discrepancy, existing methods often perform hard partitioning to mine more detail. However, these methods employ only uniform partitioning, without considering pedestrian structure, and lose a lot of pedestrian semantic information. To this end, this paper proposes a structural semantic representation reconstruction (SSRR) method to capture pedestrian semantic information by focusing on pedestrian structure. Specifically, based on the fine-grained features obtained by hard partitioning, we carry out structural reconstruction to obtain the reconstructed features containing semantic information. By adopting the direct link reconstruction structure, the reciprocal learning of fine-grained features and semantic features is ensured. Semantic features are reconstructed based on fine-grained features, and semantic information is beneficial to fine-grained features to better capture pedestrian-related details. In addition, local consistency loss is introduced to ensure the consistency of fine-grained features in the same component location, further enhancing the discriminant of the learned reconstructed representation. Extensive experiments confirm the superiority of our method on two public datasets SYSU-MM01 and RegDB.
Xi Yang 0011, Menghui Tian, Meijie Li, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.1
2024 STFE: A Comprehensive Video-Based Person Re-Identification Network Based on Spatio-Temporal Feature Enhancement
abstract
Video-based person re-identification (Re-ID) is designed to retrieve target pedestrians in video sequences under non-overlapping cameras. At present, mainstream approaches post-process the feature map extracted by the convolutional neural network backbone to obtain a global representation or a fine-grained local representation for higher accuracy. However, they still suffer from challenges, such as information loss for global-based methods and spatio-temporal feature fragmentation for local-based methods. To alleviate these problems, this article proposes a Spatio-Temporal Feature Enhancement (STFE) network from a spatio-temporal comprehensive perspective, combining the advantages of the above methods to obtain more comprehensive information from video tracklets. STFE consists of two main modules: Feature Space Projection Module (FSPM) and Global Low-frequency Enhancement Module (GLEM). FSPM mathematically converts continuous video information into a discrete feature space and selectively retains more useful information, thus avoiding spatio-temporal information loss. Meanwhile, FSPM applies global features instead of dividing feature maps spatially, thereby avoiding spatio-temporal feature fragmentation. In addition, GLEM which is based on transformer, acts as a broadband low-pass filter to mine richer global comprehensive information. Finally, by combining FSPM with GLEM, STFE can obtain spatio-temporal comprehensive video representation. Extensive experiments were conducted on two widely-used video Re-ID datasets. The experimental results verify our idea and demonstrate the effectiveness of the proposed STFE with 95.5% Rank-1 accuracy on MARS benchmarks, which surpasses previousstate-of-the-artsby a large margin of +4%.
Xi Yang 0011, Liangchen Liu 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.1
2024 SCSP: An Unsupervised Image-to-Image Translation Network Based on Semantic Cooperative Shape Perception
abstract
This paper introduces a novel approach to unsupervised image-to-image translation, aiming to overcome the limitations of existing methods in accurately capturing the shape of the source domain and the style of the target domain. The proposed method, called Semantic Cooperative Shape Perception (SCSP), focuses on enhancing the quality of generated images by addressing two key aspects. Firstly, the SCSP model employs a fusion generator that divides the mapping process into a unique texture part and a shared semantic part. By using different network structures and constraints, each part learns specific information. The unique texture generator emphasizes the style and texture details of the target domain, while the shared semantic generator focuses on the semantic information present in the source domain. This separation enables the sub-generators to extract and restore different aspects of the target domain more effectively. Secondly, a shape perception loss is introduced to improve the similarity of semantic images. It enhances the shared semantic generator's ability to perceive semantic information related to the same object by imposing constraints on the semantic graph of both the generated and input images. Therefore, the proposed method ensures semantic consistency during the translation process, leading to improved authenticity and image quality. Experimental results on four datasets, including horse2zebra, tiger2leopard, summer2winter, and photo2vangogh, demonstrate that the SCSP model achieves state-of-the-art visualization results and favorable evaluation metrics.
Xi Yang 0011, Dong Yang 0012
IEEE Trans. Multim.1
2024 Improving Cross-Modal Constraints: Text Attribute Person Search With Graph Attention Networks
abstract
Nowadays, video surveillance systems are widely deployed in public areas. However, in the unreachable corner of surveillance cameras, it still seems impossible to find the suspects only depending on eyewitness memory. Therefore, the technology that can detect particular pedestrians only by text-based attributes, or text-attribute person search, attracts lots of attention from academia. Most existing text-attribute person search methods focus on learning better feature representations by designing better network structures or using local information but lack direct constraints between modalities. This paper proposes a feature embedding motivated and graph attention network-based model, optimizing the feature extraction process by its attention mechanism. Meanwhile, this paper studies the effectiveness of the attention mechanism in feature alignment, and thus redesigns the cross-attention module, simplifying the complexity of the model and constraining the inter-modality gap in maximum by the self-attention mechanism of the graph attention network. In this way, the method simultaneously offsets the influence of modal-specific features and optimizes the number of parameters. Thus, the method improves performance and reduces time costs. Meanwhile, according to the inherent feature of attributes, this article introduces a novel embedding space, which effectively enhances the discrimination ability of the model. Extensive experiments illustrate the superiority of our model in two widely used text-attribute person search benchmarks among the state-of-the-art methods.
Xi Yang 0011, Dong Yang 0012
IEEE Trans. Multim.1
2024 Elaborate Teacher: Improved Semi-Supervised Object Detection With Rich Image Exploiting
abstract
Semi-Supervised Object Detection (SSOD) has shown remarkable results by leveraging image pairs with a teacher-student framework. An excellent strong augmentation method can generate richer images and alleviate the influence of noise in pseudo-labels. However, existing data augmentation methods for SSOD do not consider instance-level information, thus, they cannot make full use of unlabeled data. Besides, the current teacher-student framework in SSOD solely relies on pseudo-labeling techniques, which may disregard some uncertain information. In this article, we introduce a new method called Elaborate Teacher which generates and exploits image pairs in a more refined manner. To enrich strongly augmented images, a novel data augmentation method called Information-Aware Mixup Representation (IAMR) is proposed. IAMR utilizes the teacher model's predictions as prior information and considers instance-level information, which can be seamlessly integrated with existing SSOD data augmentation methods. Furthermore, to fully exploit the information in unlabeled data, we propose the Enhanced Scale Consistency Regularization (ESCR), which considers the consistency from both semantic space and feature space. Elaborate Teacher introduces a fresh data augmentation method, complemented by consistency regularization, which boosts the performance of semi-supervised object detectors. Extensive experiments on thePASCAL VOCandMS-COCOdatasets demonstrate the effectiveness of our method in leveraging unlabeled image information. Our method consistently outperforms the baseline method and improves mAP by 11.6% and 9.0% relative to the supervised baseline method when using 5% and 10% of labeled data onMS-COCO, respectively.
Xi Yang 0011, Qiubai Zhou, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.1
2024 Toward Pixel-Level Precision for Binary Super-Resolution With Mixed Binary Representation
abstract
Binary neural network (BNN) is an effective method for reducing model computational and memory cost, which has achieved much progress in the super-resolution (SR) field. However, there is still a noticeable performance gap between a binary SR network and its full-precision counterpart. Considering that the information density in quantization features is far lower than full-precision features, we aim to improve the precision of quantization features to produce rich-enough output activations for SR task. First, we make several observations that a multibit value could be approximated by multiple 1-bit values, and the computation power of binary convolution could be improved by approximating the multibit convolution process. Then, we propose a mixed binary representation set to approximate multibit activations, which is effective in compensating the quantization precision loss. Finally, we present a new precision-driven binary convolution (PDBC) module, which increases the convolution precision and protects image detail information without extra computation. Compared with normal binary convolution, our method could largely reduce the information loss caused by binarization. In experiments, our methods consistently show superior performance over the baseline models and can surpass state-of-the-art methods in terms of peak signal to noise ratio (PSNR) and visual quality.
Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Address the Unseen Relationships: Attribute Correlations in Text Attribute Person Search
abstract
Text attribute person search aims to identify the particular pedestrian by textual attribute information. Compared to person re- identification tasks which requires imagery samples as its query, text attribute person search is more useful under the circumstance where only witness is available. Most existing text attribute person search methods focus on improving the matching correlation and alignments by learning better representations of person-attribute instance pairs, with few consideration of the latent correlations between attributes. In this work, we propose a graph convolutional network (GCN) and pseudo-label-based text attribute person search method. Concretely, the model directly constructs the attribute correlations by label co- occurrence probability, in which the nodes are represented by attribute embedding and edges are by the filtered correlation matrix of attribute labels. In order to obtain better representations, we combine the cross-attention module (CAM) and the GCN. Furthermore, to address the unseen attribute relationships, we update the edge information through the instances through testing set with high predicted probability thus to better adapt the attribute distribution. Extensive experiments illustrate that our model outperforms the existing state-of-the-art methods on publicly available person search benchmarks: Market-1501 and PETA.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Task-Specific Heterogeneous Network for Object Detection in Aerial Images
abstract
Object detection in aerial images has attracted increasing attention in recent years. Due to the complex background and arbitrary-oriented objects, it is challenging to accurately locate the objects of interest in the images. Many methods have been developed for improving localization accuracy of oriented objects. However, classification and localization tasks require different features due to the unique characteristics of aerial images, which are still not fully considered in previous methods. Therefore, we propose a Task-Specific Heterogeneous (TSH) network for aerial object detection. Specifically, we design an Interference-Suppression Module (ISM) to reduce both the background and inter-class interference, which can provide discriminative features for classification. To produce more reliable localization confidence, we propose a Joint-learning Quality Estimation (JQE) module to adaptively combine the classification and regression features, thereby achieving accurate classification and localization quality estimation simultaneously. Moreover, we propose a Point-Based Localization (PBL) branch. In the PBL, the learnable points can effectively adapt to objects with diverse shapes and orientations, and the dynamic information aggregation module can enhance the relationships between the dispersible points to promote localization accuracy. The proposed TSH is evaluated extensively on four widely used aerial datasets, demonstrating its state-of-the-art performance. Ablation study and visualizations further verify the effectiveness of our method.
Xi Yang 0011, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 A Refined Hybrid Network for Object Detection in Aerial Images
abstract
Aerial object detection is a challenging task that needs to detect objects with large variations in scale and orientation. Previous dense object detectors rely on heuristic Non-Maximum Suppression (NMS) to filter out redundant detections. This may reduce the recall rate for objects with arbitraty orientations and large aspect ratios. Recently proposed sparse object detectors treat object detection as a set prediction task, effectively eliminating the need for hand-crafted components. However, applying this paradigm directly to aerial images achieves inferior performance. In this paper, we develop an effective refined hybrid network for object detection in aerial images. Our method combines the advantages of both dense and sparse detectors, achieving outstanding performance for aerial objects with large variations. Specifically, considering the highly diverse orientations of objects, we first apply a dynamic query generation module to produce high-quality oriented queries. These queries can effectively locate the foreground objects in an image, ensuring a high recall rate. Then, the object queries are sent to a query decoder for further refinement. This refinement stage adopts one-to-one matching to eliminate the negative impact caused by NMS. Moreover, an adaptive feature fusion module is designed to learn stronger modeling capabilities for rotated objects at different scales. In addition, we propose a practical mixed query sampling strategy that utilizes many-to-one assignment as an auxiliary scheme to help the detector training. Extensive experiments conducted on several aerial datasets demonstrate the superior performance of the proposed method in comparison with other state-of-the-art approaches.
Xi Yang 0011, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 FABNet: Frequency-Aware Binarized Network for Single Image Super-Resolution
abstract
Remarkable achievements have been obtained with binary neural networks (BNN) in real-time and energy-efficient single-image super-resolution (SISR) methods. However, existing approaches often adopt the Sign function to quantize image features while ignoring the influence of image spatial frequency. We argue that we can minimize the quantization error by considering different spatial frequency components. To achieve this, we propose a frequency-aware binarized network (FABNet) for single image super-resolution. First, we leverage the wavelet transformation to decompose the features into low-frequency and high-frequency components and then employ a "divide-and-conquer" strategy to separately process them with well-designed binary network structures. Additionally, we introduce a dynamic binarization process that incorporates learned-threshold binarization during forward propagation and dynamic approximation during backward propagation, effectively addressing the diverse spatial frequency information. Compared to existing methods, our approach is effective in reducing quantization error and recovering image textures. Extensive experiments conducted on four benchmark datasets demonstrate that the proposed methods could surpass state-of-the-art approaches in terms of PSNR and visual quality with significantly reduced computational costs. Our codes are available at https://github.com/xrjiang527/FABNet-PyTorch.
Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Image Process.5
2023 Frequency Information Disentanglement Network for Video-Based Person Re-Identification
abstract
Recently, most video-based person re-identification (Re-ID) methods adopt complex model or multi-scaled information to explore more discriminative spatio-temporal clues, thus achieving better retrieval accuracy. However, we witness that these approaches involve significant higher computation costs but only improve limited performances. Therefore, the overarching goal at this stage is to solve video Re-ID on the trade-off between accuracy and efficiency, thereby boosting the application in real scenarios. Frequency transform provides advantages of simplified representation, identification of hidden information and noise filtering in signal processing. Motivated by this, we treat the complex spatio-temporal feature as signal and convert it to frequency domain. By directly analyzing frequency clues, complex feature extraction procedures can be avoided. Specifically, this paper proposes a novel paradigm by categorizing video features into low/high and spatial/temporal frequency information. Then, with the help of 3D DCT, we theoretically establish the transform equivalence relationship between spatio-temporal domain and frequency domain. Finally, this paper proposes a simple and intuitive Frequency Information Disentanglement Network (FIDN) for video Re-ID. By extracting and applying both low and high frequency spatio-temporal features from a disentangling way, FIDN achieves comprehensive and discriminative video representation. Extensive experiments indicate that FIDN reaches the state-of-the-arts with only one convolution layer addition against baseline.
Liangchen Liu 0001, Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.2
2022 Towards Semi-Supervised Deep Facial Expression Recognition with An Adaptive Confidence Margin
abstract
Only parts of unlabeled data are selected to train models for most semi-supervised learning methods, whose confidence scores are usually higher than the pre-defined threshold (i.e., the confidence margin). We argue that the recognition performance should be further improved by making full use of all unlabeled data. In this paper, we learn an Adaptive Confidence Margin (Ada-CM) to fully leverage all unlabeled data for semi-supervised deep facial expression recognition. All unlabeled samples are partitioned into two subsets by comparing their confidence scores with the adaptively learned confidence margin at each training epoch: (1) subset I including samples whose confidence scores are no lower than the margin; (2) subset II including samples whose confidence scores are lower than the margin. For samples in subset I, we constrain their predictions to match pseudo labels. Meanwhile, samples in subset II participate in the feature-level contrastive objective to learn effective facial expression features. We extensively evaluate Ada-CM on four challenging datasets, showing that our method achieves state-of-the-art performance, especially surpassing fully-supervised baselines in a semi-supervised manner. Ablation study further proves the effectiveness of our method. The source code is available at https://github.com/hangyu94/Ada-CM.
Hangyu Li 0001, Nannan Wang 0001, Xi Yang 0011, Xiaoyu Wang 0002, Xinbo Gao 0001
CVPR3
2022 SAR-to-optical image translation based on improved CGAN
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
Pattern Recognit.1
2022 Dually Distribution Pulling Network for Cross-Resolution Person Reidentification
abstract
Person reidentification (Re-ID) aims at recognizing the same identity across different camera views. However, the cross resolution of images [high resolution (HR) and low resolution (LR)] is unavoidable in a realistic scenario due to the various distances among cameras and pedestrians of interest, thus leading to cross-resolution person Re-ID problems. Recently, most cross-resolution person Re-ID methods focus on solving the resolution mismatch problem, while the distribution mismatch between HR and LR images is another factor that significantly impacts the person Re-ID performance. In this article, we propose a dually distribution pulling network (DDPN) to tackle the distribution mismatch problem. DDPN is composed of two modules, that is: 1) super-resolution module and 2) person Re-ID module. They attempt to pull the distribution of LR images closer to the distribution of HR images from image and feature aspects, respectively, through optimizing the maximum mean discrepancy losses. Extensive experiments have been conducted on three benchmark datasets and the results demonstrate the effectiveness of DDPN. Remarkably, DDPN shows a great advantage when compared to the state-of-the-art methods, for instance, we achieve rank-1 accuracy of 76.9% on VR-Market1501, which outperforms the best existing cross-resolution person Re-ID method by 10%.
Yingzhi Tang, Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Cybern.2
2022 RBDF: Reciprocal Bidirectional Framework for Visible Infrared Person Reidentification
abstract
Visible infrared person reidentification (VI-REID) plays a critical role in night-time surveillance applications. Most methods attempt to reduce the cross-modality gap by extracting the modality-shared features. However, they neglect the distinct image-level discrepancies among heterogeneous pedestrian images. In this article, we propose a reciprocal bidirectional framework (RBDF) to achieve modality unification before discriminative feature learning. The bidirectional image translation subnetworks can learn two opposite mappings between visible and infrared modality. Particularly, we investigate the characteristics of the latent space and design a novel associated loss to pull close the distribution between the intermediate representations of two mappings. Mutual interaction between two opposite mappings helps the network generate heterogeneous images that have high similarity with the real images. Hence, the concatenation of original and generated images can eliminate the modality gap. During the feature learning procedure, the attention mechanism-based feature embedding network can learn more discriminative representations with the identity classification and feature metric learning. Experimental results indicate that our method achieves state-of-the-art performance. For instance, we achieve 54.41% mAP and 57.66% rank-1 accuracy on SYSU-MM01 dataset, outperforming the existing works by a large margin.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Cybern.2
2022 An Efficient and Lightweight CNN Model With Soft Quantification for Ship Detection in SAR Images
abstract
Convolutional neural networks have been widely used for synthetic aperture radar (SAR) target detection. Typical methods based on convolutional neural network have obtained favorable detection accuracy at the cost of high model complexity, and thus are difficult to be directly applied to real-time satellites on board as well as maritime rescue. To deal with this problem, this paper proposes an efficient and lightweight target detection network incorporating soft quantization. Firstly, to compensate for the lack of accuracy caused by lightweight networks, a feature fusion module called split bidirectional feature pyramid network is proposed to alleviate the interference of complex background on SAR images. Meanwhile, to adapt the lightweight network and the feature fusion module, a linear transformation module is presented to enhance the linear representation of the model via learnable parameters. Eventually, to make the model size smaller, a soft quantization algorithm is proposed to reduce the accuracy degradation caused by quantization errors. We validate the robustness of the model in several publicly available datasets. Experimental results show that our model achieves 97.0% detection accuracy on SAR ship detection dataset, with a 0.9% accuracy improvement compared to mainstream methods using less than 15x the number of parameters and less than 6x the number of flops.
Xi Yang 0011, Chengzeng Chen, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.1
2022 Learning Deep Resonant Prior for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image super-resolution (HSISR) task has been widely studied, and significant progress has been made by leveraging the deep convolution neural network (CNN) techniques. Nevertheless, the scarcity of training images hinders the research progress of HSISR task. Moreover, the differences in imaging conditions and the number of spectral bands among different datasets, make it very difficult to construct a unified deep neural network. In this paper, we first present a non-training based HSISR method based on deep prior knowledge, which captures the image prior to restore the high resolution image by using the intrinsic characteristics of CNN. Then, we append a special network input processing module onto the HSI super-resolution network to automatically adjust the structure of the input so that the choice of network structure is no longer limited, while the network design focuses on exploiting the spatial information of hyperspectral images and the correlation between spectral bands, making the method more suitable for HSISR tasks and greatly extending its applications. Extensive experiment results on the hyperspectral image datasets illustrate the effectiveness of the proposed method, and we have got comparable results with the state-of-the-art methods while requiring no training samples.
Zhaori Gong, Nannan Wang 0001, De Cheng, Jingwei Xin, Xi Yang 0011, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 An Enhanced SiamMask Network for Coastal Ship Tracking
abstract
Coastal ship tracking is significant for cargo transportation and route determination. However, there are few effective tracking methods for specialized ship tracking. Although Siamese networks have been commonly used for object tracking in the field of deep learning, the results for the ship are not accurate due to the lack of contour and edge information. In addition, scale variation and seawater cause unstable ship movements, which aggravates the reduction in tracking accuracy. Therefore, we propose an enhanced SiamMask network for coastal ship tracking. Compared to the previous Siamese network, our algorithm has the following three advantages. First, we apply the unity of visual object tracking and semisupervised object segmentation to the ship tracking task, which completes target tracking while outputting edge shape information. Second, we propose a refined feature pyramid network that utilizes a feature fusion module and enhanced residual module (ERM) to solve the problems of scale variation in datasets. Third, we propose an attentionwise cross correlation with a multidimension attention module (MDAM) to focus more on ship targets and suppress nontargets through autonomous learning at width, height, and channel levels, which creates a tradeoff between the accuracy and the robustness of tracking algorithms. The experimental results show that our method achieves leading performance in LMD-TShips, outperforming most of the state-of-the-art trackers. Code is available athttps://gitee.com/EnhancedSiamShipTracking/code.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 SRDN: A Unified Super-Resolution and Motion Deblurring Network for Space Image Restoration
abstract
Space target super-resolution (SR) is a domain-specific single image SR problem aiming to help distinguish the satellite and spacecrafts from numerous space debris. Compared to the other object SR problem, images for space target are always in low quality with varies of degradation condition, as a result of long distance and motion blur, which significantly reduces the manual classification reliability, especially for these small targets, e.g., satellite payloads. To address this challenge, we present an end-to-end SR and deblurring network (SRDN). Concretely, focusing on the low-resolution (LR) space target images with blind motion blur, we integrate the SR and deblur function together, improving the image quality by a unified generative adversarial network (GAN)-based framework. We implement a deblur module by using contrastive learning to extract degradation feature and add symmetrical downsampling and upsampling modules to the SR network in order to restore texture information, while shortcut connections are redesigned to maintain the global similarity. Extensive experiments on the public satellite dataset, BUAA-SID-share1.5, demonstrate that our network outperforms the state-of-the-art SR and deblur methods.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 FG-GAN: A Fine-Grained Generative Adversarial Network for Unsupervised SAR-to-Optical Image Translation
abstract
Synthetic aperture radar (SAR) and optical sensing are two important means of Earth observation. SAR can be used for all-day and all-weather Earth observation, but it has the disadvantages of speckle noise and geometric distortion, which are not conducive to human eye recognition. Optical image conforms to the characteristics of human visual observation, but it is easily affected by climate and time. Therefore, to integrate the advantages of the two, researchers have carried out extensive work on SAR-to-optical (S2O) image translation. Most of the existing methods for S2O image translation are supervised and need paired training samples, limiting its large-scale application in remote sensing field. Thus, we give priority to an unsupervised S2O image translation method. Meanwhile, we find that the images generated by unsupervised methods suffer from significant detail deficiencies. To solve this problem, we propose a fine-grained generative adversarial network (FG-GAN) introducing three strategies to enhance the detailed information in generated optical images. First, we design an unbalanced generator (UBG) with complex encoder networks and relatively simple decoder networks. The complex encoder extracts abundant feature information, while the decoder obtains key details by filtering these features. Second, to match the learning ability of the generator, we present a multiscale discriminator (MSD) to enhance the discriminant ability of the network. Third, we propose a comprehensive normalization group (CNG) to promote the physical representation consistency of SAR and optical images. Extensive experiments have been conducted, and the results show that our method is superior to the state-of-the-art (SOTA) methods on both subjective and objective evaluation indicators. Moreover, our FG-GAN has a significant improvement on classification accuracy, indicating its potential in facilitating the performance of practical remote sensing tasks.
Xi Yang 0011, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.1
2022 A Robust One-Stage Detector for Multiscale Ship Detection With Complex Background in Massive SAR Images
abstract
With the development of synthetic aperture radar (SAR) imaging and deep learning, SAR ship detection based on convolutional neural networks (CNNs) has been extensively applied in the last few years. Nevertheless, there are two main obstacles in SAR ship detection: 1) the SAR images have too much noise, such as the interference from land area, making it difficult to distinguish ship objects from the surrounding background, and 2) due to the multiscale characteristics of ship objects, there are numerous false negatives in the detection results, especially for small objects. To alleviate the above problems, we propose a one-stage ship detector with strong robustness against scale changes and various interferences. First, to mitigate the disturbance from complex background, a coordinate attention module (CoAM) is introduced for obtaining more representative semantic features to accurately locate and distinguish ship objects. Second, a receptive field increased module (RFIM) is devised to capture multiscale contextual information to improve the detection performance for ships with diverse scales. Finally, we verify the robustness of our method on several public SAR datasets, i.e., SAR-Ship-Dataset, high-resolution SAR images dataset (HRSID), and SAR ship detection dataset (SSDD). The experimental results demonstrate that the proposed method has a competitive performance, exceeding other state-of-the-art methods by at least 2.6% AP50on HRSID.
Xi Yang 0011, Xin Zhang 0129, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 A Cascade Rotated Anchor-Aided Detector for Ship Detection in Remote Sensing Images
abstract
Automatic ship detection in high-resolution remote sensing images has attracted increasing research interest due to its numerous practical applications. However, there still exist challenges when directly applying state-of-the-art object detection methods to real ship detection, which greatly limits the detection accuracy. In this article, we propose a novel cascade rotated anchor-aided detection network to achieve high-precision performance for detecting arbitrary-oriented ships. First, we develop a data preprocessing embedded cascade structure to reduce large amounts of false positives generated on blank areas in large-size remote sensing images. Second, to achieve accurate arbitrary-oriented ship detection, we design a rotated anchor-aided detection module. This detection module adopts a coarse-to-fine architecture with a cascade refinement module (CRM) to refine the rotated boxes progressively. Meanwhile, it utilizes an anchor-aided strategy similar to anchor-free, thus breaking through the bottlenecks of anchor-based methods and leading to a more flexible detection manner. Besides, a rotated align convolution layer is introduced in CRM to extract features from rotated regions accurately. Experimental results on the challenging DOTA and HRSC2016 data sets show that the proposed method achieves 84.12% and 90.79% AP, respectively, outperforming other state-of-the-art methods.
Xi Yang 0011, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Object Detection for Aerial Images With Feature Enhancement and Soft Label Assignment
abstract
Object detection in aerial images, different from general object detection, faces with several challenges such as arbitrary-oriented objects and extremely imbalanced foreground-background distribution. Although some recent proposed aerial object detection methods achieve promising results, they are mainly anchor-based detectors which rely heavily on pre-defined anchor boxes and the final detection performance is sensitive to anchor-related hyper-parameters. In contrast, in this paper, we present an anchor-free detector with feature enhancement and soft label assignment (FSDet) which adopts a simpler design and achieves competitive performance. Specifically, to address the feature misalignment for detecting oriented objects, we propose an oriented feature refinement module to align the features with oriented objects. To alleviate the background issue, we design a class-aware context aggregation module to integrate the intra-class context information and suppress background context. Moreover, we propose a soft label assignment mechanism to measure the weight of training samples within the arbitrary-oriented objects, which can concentrate more on representative items with regard to their potential to detect oriented objects, achieving a more stable optimization during training. Extensive experiments on several datasets suggest that the proposed method is superior to the state-of-the-art methods and achieves a better trade-off between speed and accuracy.
Xi Yang 0011, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 A Universal Ship Detection Method With Domain-Invariant Representations
abstract
Although ship detection methods based on deep learning have achieved remarkable progress, the design of the universal ship detection (USD) system is rarely studied. The main challenge of USD lies in the notorious domain bias and shift problem across multiple domains. This article implements USD based on domain-invariant representations to alleviate this issue. Specifically, the proposed method integrates a multilevel domain classification network (MDCN) and a domain-centric cut-paste module (DCM). First, the backbone network is facilitated to learn domain-independent image features from multiple domains through MDCN, thereby reducing the disturbance of domain-specific features to universal detector. Furthermore, the proposed method combines the domain-related synthetic samples generated by DCM to provide MDCN with strong supervision information, which further motivates the network to be more attentive to the domain-invariant representations at the instance level. Finally, we conduct experiments on multiple ship datasets in the synthetic aperture radar (SAR) and optical domains to verify the effectiveness of the method. The results show that the proposed method outperforms baseline by around 2.95% average precision (AP50), which achieves an effective USD system by complementing the information between domain-invariant representations.
Xin Zhang 0129, Xi Yang 0011, Dong Yang 0012, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 CRS-CONT: A Well-Trained General Encoder for Facial Expression Analysis
abstract
Existing facial expression recognition (FER) methods train encoders with different large-scale training data for specific FER applications. In this paper, we propose a new task in this field. This task aims to pre-train a general encoder to extract any facial expression representations without fine-tuning. To tackle this task, we extend the self-supervised contrastive learning to pre-train a general encoder for facial expression analysis. To be specific, given a batch of facial expressions, some positive and negative pairs are firstly constructed based on coarse-grained labels and a FER-specified data augmentation strategy. Secondly, we propose the coarse-contrastive (CRS-CONT) learning, where the features of positive pairs are pulled together, while pushed away from the features of negative pairs. Moreover, one key event is that the excessive constraint on the coarse-grained feature distribution will affect fine-grained FER applications. To address this, a weight vector is designed to control the optimization of the CRS-CONT learning. As a result, a well-trained general encoder with frozen weights could preferably adapt to different facial expressions and realize the linear evaluation on any target datasets. Extensive experiments on both in- the-wild and in- the-lab FER datasets show that our method provides superior or comparable performance against state-of-the-art FER methods, especially on unseen facial expressions and cross-dataset evaluation. We hope that this work will help to reduce the training burden and develop a new solution against the fully-supervised feature learning with fine-grained labels. Code and the general encoder will be publicly available at https://github.com/hangyu94/CRS-CONT.
Hangyu Li 0001, Nannan Wang 0001, Xi Yang 0011, Xinbo Gao 0001
IEEE Trans. Image Process.3
2022 Towards Multi-Domain Face Synthesis Via Domain-Invariant Representations and Multi-Level Feature Parts
abstract
Cross-domain face synthesis plays a positive role in the real world. It is challenging to synthesize high-quality faces across multiple domains based on limited paired data because the multiple mappings between different domains may interfere with each other. Cognitive science investigates that the brain can recognize the same person with multiple different expressions by extracting invariant information on the face and we humans perceive instances by decomposing them into parts. Motivated by these cognition, we propose a unified semi-supervised framework for multi-domain face synthesis by extracting a domain-invariant representation and exploiting parts of multi-level features. Specifically, realized by adversarial training with additional ability to utilize domain-specific information, a encoder is trained to remove domain-specific information and extract the domain-invariant representation from multiple inputs. Then, we utilize the multi-level feature parts extracted from inputs and reconstructed faces via a pre-trained recognition model to ensure that the domain-invariant representation contains enough useful semantic information. we also utilize the feature parts extracted from inputs and limited paired data to compose pseudo features in target domain for supervising the synthesis, which makes our framework suitable for large amounts of unpaired training data. By exploiting this framework, we can achieve face synthesis between multiple domains using some paired data together with a large training database without ground truth target faces. Experimental results demonstrate our framework achieves great performances on qualitative and quantitative evaluations under both artificial and uncontrolled environments, and our framework has competitive performances in single translation compared with specialized methods for translation between two specific domains.
Dawei Zhou 0004, Nannan Wang 0001, Chunlei Peng, Yi Yu 0001, Xi Yang 0011, Xinbo Gao 0001
IEEE Trans. Multim.5
2022 Flexible Body Partition-Based Adversarial Learning for Visible Infrared Person Re-Identification
abstract
Person re-identification (Re-ID) aims to retrieve images of the same person across disjoint camera views. Most Re-ID studies focus on pedestrian images captured by visible cameras, without considering the infrared images obtained in the dark scenarios. Person retrieval between visible and infrared modalities is of great significance to public security. Current methods usually train a model to extract global feature descriptors and obtain discriminative representations for visible infrared person Re-ID (VI-REID). Nevertheless, they ignore the detailed information of heterogeneous pedestrian images, which affects the performance of Re-ID. In this article, we propose a flexible body partition (FBP) model-based adversarial learning method (FBP-AL) for VI-REID. To learn more fine-grained information, FBP model is exploited to automatically distinguish part representations according to the feature maps of pedestrian images. Specially, we design a modality classifier and introduce adversarial learning which attempts to discriminate features between visible and infrared modality. Adaptive weighting-based representation learning and threefold triplet loss-based metric learning compete with modality classification to obtain more effective modality-sharable features, thus shrinking the cross-modality gap and enhancing the feature discriminability. Extensive experimental results on two cross-modality person Re-ID data sets, i.e., SYSU-MM01 and RegDB, exhibit the superiority of the proposed method compared with the state-of-the-art solutions.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2021 Training Binary Neural Network without Batch Normalization for Image Super-Resolution
abstract
Recently, binary neural network (BNN) based super-resolution (SR) methods have enjoyed initial success in the SR field. However, there is a noticeable performance gap between the binarized model and the full-precision one. Furthermore, the batch normalization (BN) in binary SR networks introduces floating-point calculations, which is unfriendly to low-precision hardwares. Therefore, there is still room for improvement in terms of model performance and efficiency. Focusing on this issue, in this paper, we first explore a novel binary training mechanism based on the feature distribution, allowing us to replace all BN layers with a simple training method. Then, we construct a strong baseline by combining the highlights of recent binarization methods, which already surpasses the state-of-the-arts. Next, to train highly accurate binarized SR model, we also develop a lightweight network architecture and a multi-stage knowledge distillation strategy to enhance the model representation ability. Extensive experiments demonstrate that the proposed method not only presents advantages of lower computation as compared to conventional floating-point networks but outperforms the state-of-the-art binary methods on the standard SR networks.
Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Xinbo Gao 0001
AAAI5
2021 Syncretic Modality Collaborative Learning for Visible Infrared Person Re-Identification
abstract
Visible infrared person re-identification (VI-REID) aims to match pedestrian images between the daytime visible and nighttime infrared camera views. The large cross-modality discrepancies have become the bottleneck which limits the performance of VI-REID. Existing methods mainly focus on capturing cross-modality sharable representations by learning an identity classifier. However, the heterogeneous pedestrian images taken by different spectrum cameras differ significantly in image styles, resulting in inferior discriminability of feature representations. To alleviate the above problem, this paper explores the correlation between two modalities and proposes a novel syncretic modality collaborative learning (SMCL) model to bridge the cross-modality gap. A new modality that incorporates features of heterogeneous images is constructed automatically to steer the generation of modality-invariant representations. Challenge enhanced homogeneity learning (CEHL) and auxiliary distributional similarity learning (ADSL) are integrated to project heterogeneous features on a unified space and enlarge the inter-class disparity, thus strengthening the discriminative power. Extensive experiments on two cross-modality benchmarks demonstrate the effectiveness and superiority of the proposed method. Especially, on SYSU-MM01 dataset, our SMCL model achieves 67.39% rank-1 accuracy and 61.78% mAP, surpassing the cutting-edge works by a large margin.
Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
ICCV2
2021 Viewing from Frequency Domain: A DCT-based Information Enhancement Network for Video Person Re-Identification
abstract
Video-based person re-identification (Re-ID) aims to match the target pedestrians under non-overlapping camera system by video tracklets. The key issue of video Re-ID focuses on exploring effective spatio-temporal features. Generally, the spatio-temporal information of a video sequence can be divided into two aspects: the discriminative information in each frame and the shared information over the whole sequence. To make full use of the rich information in video sequences, this paper proposes a Discrete Cosine Transform based Information Enhancement Network (DCT-IEN) to achieve more comprehensive spatio-temporal representation from frequency domain. Inspired by the principle that average pooling is one of the special frequency components in DCT (the lowest frequency component), DCT-IEN first adopts discrete cosine transform to convert the extracted feature maps into frequency domain, thereby retaining more information that embedded in different frequency components. With the help of DCT frequency spectrum, two branches are adopted to learn the final video representation: Frequency Selection Module (FSM) and Lowest Frequency Enhancement Module (LFEM). FSM explores the most discriminative features in each frame by aggregating different frequency components with attention mechanism. LFEM enhances the shared feature over the whole video sequence by frame feature regularization. By fusing these two kinds of features together, DCT-IEN finally achieves comprehensive video representation. We conduct extensive experiments on two widely used datasets. The experimental results verify our idea and demonstrate the effectiveness of DCT-IEN for video-based Re-ID.
Liangchen Liu 0001, Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
ACM Multimedia2
2021 LBAN-IL: A novel method of high discriminative representation for facial expression recognition
Hangyu Li 0001, Nannan Wang 0001, Yi Yu 0001, Xi Yang 0011, Xinbo Gao 0001
Neurocomputing4
2021 Learning lightweight super-resolution networks with weight pruning
Nannan Wang 0001, Jingwei Xin, Xiaobo Xia, Xi Yang 0011, Xinbo Gao 0001
Neural Networks5
2021 A CenterNet++ model for ship detection in SAR images
Haoyuan Guo, Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
Pattern Recognit.2
2021 Hierarchical Deep Embedding for Aurora Image Retrieval
abstract
Retrieving informative images from the large-scale aurora data is of great significance in the field of space physics. In this article, we propose a hierarchical deep embedding (HDE) model to assist scientists for their aurora image retrieval. Other than conventional bag-of-words (BoW) models employing local cues individually, HDE performs visual matching in a hierarchical way, that is, only keypoints which are similar on local, regional, and global simultaneously can be treated as a true match. The added contextual evidences can effectively alleviate the occurrence of false matches and improve the precision of visual matching. Specifically, to complement the local SIFT feature, the convolutional neural network (CNN) is refined with a polar region pooling (PRP) layer to extract features from regional patches and global image, forming a group of hierarchical deep features with strong discriminative power. Also, an improved polar meshing (IPM) scheme is presented to determine the positions of keypoints, which is more suitable for images captured by circular fisheye lens and capable of reflecting the physical information in aurora images. Extensive experiments are conducted on the big aurora data, which indicate that the proposed HDE model greatly promotes the retrieval accuracy with acceptable memory cost and efficiency. In addition, the effectiveness of the IPM scheme and the superiority of the hierarchical deep feature integration are separately demonstrated.
Xi Yang 0011, Xinbo Gao 0001, Bin Song 0001, Bing Han 0003
IEEE Trans. Cybern.1
2021 An Efficient Graph-Based Algorithm for Time-Varying Narrowband Interference Suppression on SAR System
abstract
Synthetic aperture radar (SAR) as a wideband radar system is subject to complicated interferences, such as radio frequency interference or other narrowband interferences (NBIs). In order to suppress the NBI, voluminous literature focused on its signal models and characteristics, such as the sinusoidal model and relatively constant frequencies. However, in practice, the interference environment is commonly complicated. It is hard to model the interferences accurately and mitigate them clearly in an easy way, especially for the time-varying interferences. In this article, a novel graph-based algorithm is proposed to mitigate the time-varying NBIs by using graph theory, which constructs the connections between different azimuth samples of NBIs. As a result, the locally time-varying interferences can be clustered in a nonlinear low-dimensional manifold and effectively removed by the proposed algorithm. In addition, the case of the globally time-varying interference is also analyzed in detail with strict derivations to demonstrate its low-rank property. Furthermore, the matrix factorization scheme is introduced to improve the efficiency of the proposed algorithm, and the closed-form solutions are derived for each iteration. The real SAR data with measured NBIs are provided to demonstrate the effectiveness and efficiency of the proposed algorithm.
Yan Huang 0018, Lei Zhang 0019, Xi Yang 0011, Zhanye Chen, Jie Li 0027, Wei Hong 0002
IEEE Trans. Geosci. Remote. Sens.3
2021 Adaptively Learning Facial Expression Representation via C-F Labels and Distillation
abstract
Facial expression recognition is of significant importance in criminal investigation and digital entertainment. Under unconstrained conditions, existing expression datasets are highly class-imbalanced, and the similarity between expressions is high. Previous methods tend to improve the performance of facial expression recognition through deeper or wider network structures, resulting in increased storage and computing costs. In this paper, we propose a new adaptive supervised objective named AdaReg loss, re-weighting category importance coefficients to address this class imbalance and increasing the discrimination power of expression representations. Inspired by human beings' cognitive mode, an innovative coarse-fine (C-F) labels strategy is designed to guide the model from easy to difficult to classify highly similar representations. On this basis, we propose a novel training framework named the emotional education mechanism (EEM) to transfer knowledge, composed of a knowledgeable teacher network (KTN) and a self-taught student network (STSN). Specifically, KTN integrates the outputs of coarse and fine streams, learning expression representations from easy to difficult. Under the supervision of the pre-trained KTN and existing learning experience, STSN can maximize the potential performance and compress the original KTN. Extensive experiments on public benchmarks demonstrate that the proposed method achieves superior performance compared to current state-of-the-art frameworks with 88.07% on RAF-DB, 63.97% on AffectNet and 90.49% on FERPlus.
Hangyu Li 0001, Nannan Wang 0001, Xinpeng Ding, Xi Yang 0011, Xinbo Gao 0001
IEEE Trans. Image Process.4
2021 A Two-Stream Dynamic Pyramid Representation Model for Video-Based Person Re-Identification
abstract
Video-based person re-identification (Re-ID) leverages rich spatio-temporal information embedded in sequence data to further improve the retrieval accuracy comparing with single image Re-ID. However, it also brings new difficulties. 1) Both spatial and temporal information should be considered simultaneously. 2) Pedestrian video data often contains redundant information and 3) suffers from data quality problems such as occlusion, background clutter. To solve the above problems, we propose a novel two-stream Dynamic Pyramid Representation Model (DPRM). DPRM mainly consists of three sub-models, i.e., Pyramidal Distribution Sampling Method (PDSM), Dynamic Pyramid Dilated Convolution (DPDC) and Pyramid Attention Pooling (PAP). PDSM is applied for more effective data pre-processing according to sequence semantic distribution. DPDC and PAP can be considered as two streams to describe the motion context and static appearance of a video sequence, respectively. By fusing the two-stream features together, we finally achieve comprehensive spatio-temporal representation. Notably, dynamic pyramid strategy is applied throughout the whole model. This strategy exploits multi-scale features under attention mechanism to maximally capture the most discriminative features and mitigate the impact of video data quality problems such as partial occlusion. Extensive experiments demonstrate the outperformance of DPRM. For instance, it achieves 83.0% mAP and 89.0% Rank-1 accuracy on MARS dataset and reaches state of the art.
Xi Yang 0011, Liangchen Liu 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2020 ABP: Adaptive Body Partition Model For Visible Infrared Person Re-Identification
abstract
Person re-identification (Re-ID) aims to match pedestrian images cross multiple cameras. Most Re-ID studies focus on visible pedestrian images, without considering the images obtained by infrared cameras in the dark. To solve the cross-modality person Re-ID problem, current methods usually exploit global feature descriptors to obtain discriminative representations. However, they ignore the fine-grained information of heterogeneous images. In this paper, we propose an adaptive body partition (ABP) model to automatically detect and distinguish effective part representations. Instead of utilizing a two-stream convolutional neural network (CNN) to extract modality-specific information, we directly design an end-to-end one-stream CNN to simultaneously learn multi-modality sharable features and map them on a common space. Global loss, part losses and threefold triplet loss are integrated to enhance the feature discriminability and minimize the cross-modality gap. Extensive experimental results on two cross-modality Re-ID datasets exhibit the superiority of the proposed method compared with the state-of-the-art solutions.
Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
ICME2
2020 Image super-resolution via multi-view information fusion networks
Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Yi Yu 0001, Xinbo Gao 0001
Neurocomputing4
2020 Person Re-Identification with Feature Pyramid Optimization and Gradual Background Suppression
Yingzhi Tang, Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
Neural Networks2
2020 HCNN-PSI: A hybrid CNN with partial semantic information for space target recognition
Xi Yang 0011, Tan Wu, Nannan Wang 0001, Yan Huang 0018, Bin Song 0001, Xinbo Gao 0001
Pattern Recognit.1
2020 A novel deformable body partition model for MMW suspicious object detection and dynamic tracking
Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
Signal Process.1
2020 A Rotational Libra R-CNN Method for Ship Detection
abstract
Recently, ship detection methods based on deep learning have attracted significant attention due to their superior accuracy over traditional methods. However, there still exist two problems affecting its robustness in practical application. 1) The size of ships in one image varies greatly, i.e., different sizes; 2) Numerous ships gather in limited field-of-view, i.e., dense distribution. To address these problems, we propose a rotational Libra R-convolutional neural network (CNN) method. Our idea is to balance the three levels of neural networks for predicting the location of ships with rotational angle information, which refers to the feature level, sample level, and objective level. First, to extract a discriminative feature and improve its robustness against the impact of different sizes of ships, the concept of balanced feature pyramid is introduced. Second, to generate reliable proposals for feature pyramid and efficiently mine hard negative samples, we employ intersection over union (IoU)-balanced sampling. Finally, to eliminate the redundant background and detect densely distributed ships, we bring in a rotational region detection branch with balanced L1 loss. In general, we develop the balanced learning with rotational region detection to achieve consistent improvement on accuracy and visualization. Experimental results on DOTA data set show that the proposed method achieves the state-of-the-art accuracy.
Haoyuan Guo, Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.2
2020 Reweighted Tensor Factorization Method for SAR Narrowband and Wideband Interference Mitigation Using Smoothing Multiview Tensor Model
abstract
For the interference suppression problem on synthetic aperture radar (SAR) systems, traditional methods have focused on how to remove one kind of interference through nonparametric methods and parametric methods. However, complicated interferences, including both narrowband interferences (NBIs) and wideband interferences (WBIs), severely affect SAR imaging in practical scenarios. Also, the spectra of the complicated interferences can be continuously distributed, which are even harder to mitigate from the received signal. Hence, in this article, we propose a smoothing multiview (SMV) tensor model in range-azimuth-space domain to represent the intrinsically unified characteristics of the NBIs and the WBIs for SAR systems, reserving more azimuth degrees-of-freedom (DOFs) than the previous MV tensor model. The proposed SMV tensor model can enhance the potential low-rank property of the complicated interferences, even though the interferences may be continuously distributed in low-dimensional domains. Moreover, due to the larger scale of the SMV model than those of the traditional models, a complex reweighted tensor factorization (CRTF) algorithm is proposed to factorize the large-scale tensor into the product of two small-scale tensors, achieving both better computational efficiency and better low-rank approximation of complicated interferences. Finally, the measured SAR data with different kinds of simulated complicated interferences are employed to demonstrate the effectiveness and efficiency of the newly designed SMV model and the proposed method compared with the MV model and the complex tensor robust principal component analysis (CT-RPCA) method.
Yan Huang 0018, Lei Zhang 0019, Jie Li 0027, Zhanye Chen, Xi Yang 0011
IEEE Trans. Geosci. Remote. Sens.5
2020 D2N4: A Discriminative Deep Nearest Neighbor Neural Network for Few-Shot Space Target Recognition
abstract
With the rapid development of space exploration worldwide, there is a sudden increase in the type and number of spacecraft, thus leading to a more complex space environment. To enhance the ability of space situational awareness, the most important step is to effectively recognize space targets of interests from various spacecraft and debris. Traditional space target recognition approaches adopt manual feature extraction with limited data, resulting in a semantic gap between low-level visual features and high-level semantic representation. Although deep learning models alleviate this problem with a unified framework for combined learning feature extraction and classification simultaneously, it is easy to overfit and leads to poor generalization results when faced with a situation of small examples. To address these issues, we present an end-to-end few-shot deep learning framework for space target recognition, i.e., discriminative deep nearest neighbor neural network (D2N4). Our D2N4 aims to improve the discriminability of the deeply learned features with mainly two strategies. On the one hand, we add an intraclass compactness principle by introducing center loss to efficiently pull deep features of the same classes to their centers and, thus overcoming significant intraclass variation of space target. On the other hand, we introduce the global pooling information for each deep local descriptor to reduce interference from local background noise, thus enhancing the model robustness. In practice, under the joint supervision of soft-max loss and center loss, the deep embedding module and image-to-class metric module are trained in an end-to-end way. Extensive experiments on the space target data set BUAA-SID-share1.0 demonstrate that our simple and effective approach outperforms previous space target recognition methods and is more efficient than recent few-shot approaches. In addition, the proposed framework is equally applicable to natural images and achieves state-of-the-art performance on data sets CUB-200-2010, Stanford Dogs, and Stanford Cars.
Xi Yang 0011, Xiaoting Nan, Bin Song 0001
IEEE Trans. Geosci. Remote. Sens.1
2020 Aurora Image Search With a Saliency-Weighted Region Network
abstract
On account of the remarkable performance of convolutional neural network (CNN) features for natural image searches, utilizing it for other images collected with the anamorphic lens has become a research hotspot. This article selects the aurora images generated from a circular fisheye lens as a typical example. By considering the imaging principle and geomagnetic information, a saliency-weighted region network (SWRN) is presented and introduced into the Mask R-CNN pipeline. Our SWRN selects salient regions with important semantic information and weights them both hierarchically and spatially. Hence, regions encompassing the search target are strengthened while uninformative regions are discarded, which benefits the suppression of background interference and reduction of computational complexity. In practice, by aggregating the outputs of SWRN with post-processing, a compact CNN feature is generated to represent the aurora image. Large-scale aurora image search experiments are conducted, and the results prove that our method performs better than the state-of-the-art methods on both accuracy and efficiency.
Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.1
2020 CGAN-TM: A Novel Domain-to-Domain Transferring Method for Person Re-Identification
abstract
Person re-identification (re-ID) is a technique aiming to recognize person cross different cameras. Although some supervised methods have achieved favorable performance, they are far from practical application owing to the lack of labeled data. Thus, unsupervised person re-ID methods are in urgent need. Generally, the commonly used approach in existing unsupervised methods is to first utilize the source image dataset for generating a model in supervised manner, and then transfer the source image domain to the target image domain. However, images may lose their identity information after translation, and the distributions between different domains are far away. To solve these problems, we propose an image domain-to-domain translation method by keeping pedestrian's identity information and pulling closer the domains' distributions for unsupervised person re-ID tasks. Our work exploits the CycleGAN to transfer the existing labeled image domain to the unlabeled image domain. Specially, a Self-labeled Triplet Net is proposed to maintain the pedestrian identity information, and maximum mean discrepancy is introduced to pull the domain distribution closer. Extensive experiments have been conducted and the results demonstrate that the proposed method performs superiorly than the state-ofthe- art unsupervised methods on DukeMTMC-reID and Market- 1501.
Yingzhi Tang, Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
IEEE Trans. Image Process.2
2020 A Novel Symmetry Driven Siamese Network for THz Concealed Object Verification
abstract
Security inspection aims to improve the high detection rate as well as reduce the false alarm rate. However, it still suffers from two challenges affecting its robustness. 1) Existing security inspection methods are mostly designed for natural images, which cannot reflect the uniqueness and imaging principle of THz images. 2) Existing methods is sensitive to noise interference and pose variations. This work revisits these challenges and presents a novel symmetry driven Siamese network (SDSN) for THz concealed object verification. Our idea is to employ a specially designed network architecture for THz concealed object verification. First, to reflect the uniqueness and the special property of THz images, Siamese network with Contrastive loss is used for feature extraction along with symmetrical prior information consideration, which can learn symmetrical metrics from the same person. Second, to alleviate the impact of noise interference and pose variations, the adaptive identity normalization (A-IDN) is proposed to normalize the symmetrical metrics each person. Finally, to enhance the generalization of network, an adaptive selective threshold based on Gaussian mixture model (AST-GMM) is designed, which serves as a classifier for the final classification results. Extensive experiments show that SDSN significantly improves the accuracy. Specially, SDSN outperforms the state-of-the-art methods without symmetrical prior information on THz security dataset.
Xi Yang 0011, Haoyuan Guo, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2019 Person re-Identification with Gradual Background Suppression
abstract
Person re-identification plays an important role in public security. However, owing to the interference of background clutters, its performance still needs to be improved. Several mask-based methods aim to solve this problem by totally removing the background clutters, but the promotion is limited because of the mask sharpening effect. In this paper, we propose a novel person re-identification method with Gradual Background Suppression (GBS). The GBS adopts several CNN branches to extract deep features of images with different weight distributions between background and human body. Thus, it can not only reduce the background clutters but also keep the smoothness of target pedestrians. Afterwards, deep features from different CNN branches are integrated with a fusion scheme, and the fused feature is capable of balancing the influence of background clutter and mask sharpening. Extensive experiments have been conducted and the results prove the superiority of the proposed GBS over the background removal approach. Comparing with the state-of-the-art methods, our method achieves remarkable performance with 6.6%, 7.58% and 8.26% improvement of mAP on dataset Market-1501, CUHK03-labeled and CUHK03-detected, respectively.
Yingzhi Tang, Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
ICME2
2019 T-SCNN: A Two-Stage Convolutional Neural Network for Space Target Recognition
abstract
Space target recognition plays an important role in the field of space security and exploration. With the rapid development of artificial intelligence technique and explosive increase of image dataset, object recognition based on deep learning has achieved favorable performance. However, the recognition of deep space targets in visible spectrum images still remains in the traditional manual interpretation approach, thus leading to low efficiency and inevitable subjective errors. In this paper, we propose an artificial intelligence method for space target recognition, called Two-Stage Convolutional Neural Network (T-SCNN). Our T-SCNN is composed of two stages, i.e., target locating and target recognition. In the stage of target locating, we first detect all suspected targets from the total image dataset by presenting a minimum bounding rectangle with threshold (MBRT) approach, then cut out all regions encompassing targets to generate target images for training. In the stage of target recognition, we send target images to the well-trained recognition network for identification. Additionally, data augmentation is conducted in the CNN training to satisfy its data quantity requirement. Extensive experiments are performed on our synthetic space target image dataset, and the result demonstrate that the proposed method achieves high accuracy within a short time.
Tan Wu, Xi Yang 0011, Bin Song 0001, Nannan Wang 0001, Xinbo Gao 0001, Liyang Kuang, Xiaoting Nan, Dong Yang 0012
IGARSS2
2019 BoSR: A CNN-based aurora image retrieval method
Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
Neural Networks1
2019 CNN with spatio-temporal information for fast suspicious object detection and recognition in THz security images
Xi Yang 0011, Tan Wu, Lei Zhang 0019, Dong Yang 0012, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
Signal Process.1
2018 Saliency Deep Embedding for Aurora Image Search
abstract
Deep neural networks have achieved remarkable success in the field of image search. However, the state-of-the-art algorithms are trained and tested for natural images captured with ordinary cameras. In this paper, we aim to explore a new search method for images captured with circular fisheye lens, especially the aurora images. To reduce the interference from uninformative regions and focus on the most interested regions, we propose a saliency proposal network (SPN) to replace the region proposal network (RPN) in the recent Mask R-CNN. In our SPN, the centers of the anchors are not distributed in a rectangular meshing manner, but exhibit spherical distortion. Additionally, the directions of the anchors are along the deformation lines perpendicular to the magnetic meridian, which perfectly accords with the imaging principle of circular fisheye lens. Extensive experiments are performed on the big aurora data, demonstrating the superiority of our method in both search accuracy and efficiency.
Xi Yang 0011, Xinbo Gao 0001, Bin Song 0001, Nannan Wang 0001, Dong Yang 0012
ICME1
2018 Aurora image search with contextual CNN feature
Xi Yang 0011, Xinbo Gao 0001, Bin Song 0001, Dong Yang 0012
Neurocomputing1
2016 Ground moving target detection in MIMO-SAR system
abstract
Multiple-input multiple-output synthetic aperture radar (MIMO-SAR) has drawn widely attention for its increased degrees of freedom. This enhanced architecture offers not only the opportunity to map wider images swaths with improved spatial resolution, but also enables novel SAR modes to resolve some of the contradicting user requirements. In this paper, the performance of ground moving target indication (GMTI) in MIMO-SAR system is analyzed. It becomes evident that the virtual channels can be utilized to fulfill both the image swath and the required signal-to-noise-radio. An analysis on orthogonal frequency division multiplexing (OFDM) chirp signal designing is proposed and a novel GMTI method based on low-rank property is also proposed. Simulation and experimental results show its good performance, which implies that MIMO-SAR would be a good choice for future multichannel systems.
Dong Yang 0012, Xi Yang 0011, Xiaomin Tan, Hongxing Dang
IGARSS2
2016 Shape-Constrained Sparse and Low-Rank Decomposition for Auroral Substorm Detection
abstract
An auroral substorm is an important geophysical phenomenon that reflects the interaction between the solar wind and the Earth's magnetosphere. Detecting substorms is of practical significance in order to prevent disruption to communication and global positioning systems. However, existing detection methods can be inaccurate or require time-consuming manual analysis and are therefore impractical for large-scale data sets. In this paper, we propose an automatic auroral substorm detection method based on a shape-constrained sparse and low-rank decomposition (SCSLD) framework. Our method automatically detects real substorm onsets in large-scale aurora sequences, which overcomes the limitations of manual detection. To reduce noise interference inherent in current SLD methods, we introduce a shape constraint to force the noise to be assigned to the low-rank part (stationary background), thus ensuring the accuracy of the sparse part (moving object) and improving the performance. Experiments conducted on aurora sequences in solar cycle 23 (1996-2008) show that the proposed SCSLD method achieves good performance for motion analysis of aurora sequences. Moreover, the obtained results are highly consistent with manual analysis, suggesting that the proposed automatic method is useful and effective in practice.
Xi Yang 0011, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001, Bing Han 0003, Jie Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2015 Strong Clutter Suppression via RPCA in Multichannel SAR/GMTI System
abstract
Clutter suppression and ground moving target indication are challenging tasks in multichannel synthetic aperture radar (SAR) systems. In recent years, robust principal component analysis (RPCA) has attracted much attention for its good performance in distinguishing the different parts from a set of correlative database. Therefore, we propose a fast RPCA-based detection method for multichannel SAR under a strong clutter background in this letter even with channel unbalance or platform motion error. Subsequently, as the existing space-time adaptive processing (STAP) method would fail when the training samples are contaminated by the moving target, we apply the RPCA-based method in the range-Doppler domain to improve the performance of STAP. Since the regions of targets can be detected via RPCA, the remaining samples, which can be regarded as only clutter, are used to estimate the covariance matrix for further processing. The final experiments based on real measured data set show its good performance under the strong clutter background. Although the RPCA-based result differs from that of the STAP method, they can work cooperatively to get a more robust detection performance.
Dong Yang 0012, Xi Yang 0011, Guisheng Liao, Shengqi Zhu 0001
IEEE Geosci. Remote. Sens. Lett.2
2015 Polar Embedding for Aurora Image Retrieval
abstract
Exploring the multimedia techniques to assist scientists for their research is an interesting and meaningful topic. In this paper, we focus on the large-scale aurora image retrieval by leveraging the bag-of-visual words (BoVW) framework. To refine the unsuitable representation and improve the retrieval performance, the BoVW model is modified by embedding the polar information. The superiority of the proposed polar embedding method lies in two aspects. On the one hand, the polar meshing scheme is conducted to determine the interest points, which is more suitable for images captured by circular fisheye lens. Especially for the aurora image, the extracted polar scale-invariant feature transform (polar-SIFT) feature can also reflect the geomagnetic longitude and latitude, and thus facilitates the further data analysis. On the other hand, a binary polar deep local binary pattern (polar-DLBP) descriptor is proposed to enhance the discriminative power of visual words. Together with the 64-bit polar-SIFT code obtained via Hamming embedding, the multifeature index is performed to reduce the impact of false positive matches. Extensive experiments are conducted on the large-scale aurora image data set. The experimental result indicates that the proposed method improves the retrieval accuracy significantly with acceptable efficiency and memory cost. In addition, the effectiveness of the polar-SIFT scheme and polar-DLBP integration are separately demonstrated.
Xi Yang 0011, Xinbo Gao 0001, Qi Tian 0001
IEEE Trans. Image Process.1
2015 An Efficient MRF Embedded Level Set Method for Image Segmentation
abstract
This paper presents a fast and robust level set method for image segmentation. To enhance the robustness against noise, we embed a Markov random field (MRF) energy function to the conventional level set energy function. This MRF energy function builds the correlation of a pixel with its neighbors and encourages them to fall into the same region. To obtain a fast implementation of the MRF embedded level set model, we explore algebraic multigrid (AMG) and sparse field method (SFM) to increase the time step and decrease the computation domain, respectively. Both AMG and SFM can be conducted in a parallel fashion, which facilitates the processing of our method for big image databases. By comparing the proposed fast and robust level set method with the standard level set method and its popular variants on noisy synthetic images, synthetic aperture radar (SAR) images, medical images, and natural images, we comprehensively demonstrate the new method is robust against various kinds of noises. In particular, the new level set method can segment an image of size 500 × 500 within 3 s on MATLAB R2010b installed in a computer with 3.30-GHz CPU and 4-GB memory.
Xi Yang 0011, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001, Jie Li 0001
IEEE Trans. Image Process.1
2014 A shape-initialized and intensity-adaptive level set method for auroral oval segmentation
Xi Yang 0011, Xinbo Gao 0001, Jie Li 0001, Bing Han 0003
Inf. Sci.1
2014 SAR Imaging With Undersampled Data via Matrix Completion
abstract
High-resolution synthetic aperture radar (SAR) imagery of a wide area of surveillance is a difficult large-data problem. In the past few years, researchers have applied compressive sensing (CS) to SAR, as it exploits redundancy in signals. To further extend the sparse problem from the vector to the matrix, a new theory called matrix completion (MC) has attracted much attention, which can complete a matrix from a small set of corrupted entries based on the assumption that the matrix is essentially of low rank. Inspired by this technique, a novel SAR imaging algorithm is proposed in this letter to deal with the undersampled data. After representing the data of a range cell as a matrix, the phase is compensated to keep the matrix holding the property of low rank. Subsequently, MC can be utilized to recover the full-aperture data in the new constructed matrix. Since the data are completely unsampled in the corresponding azimuth cells, the proposed method has effectively conquered the restriction of previous applications that each received channel must have a small number of samples. The final results in both simulation and real-data experiments show that the targets can be well focused even in the scenario of discarding a large percentage of the received pulses. Moreover, when compared with CS, the method is not required to design the complicated measurement matrix.
Dong Yang 0012, Guisheng Liao, Shengqi Zhu 0001, Xi Yang 0011, Xuepan Zhang
IEEE Geosci. Remote. Sens. Lett.4
2014 Improving Level Set Method for Fast Auroral Oval Segmentation
abstract
Auroral oval segmentation from ultraviolet imager images is of significance in the field of spatial physics. Compared with various existing image segmentation methods, level set is a promising auroral oval segmentation method with satisfactory precision. However, the traditional level set methods are time consuming, which is not suitable for the processing of large aurora image database. For this purpose, an improving level set method is proposed for fast auroral oval segmentation. The proposed algorithm combines four strategies to solve the four problems leading to the high-time complexity. The first two strategies, including our shape knowledge-based initial evolving curve and neighbor embedded level set formulation, can not only accelerate the segmentation process but also improve the segmentation accuracy. And then, the latter two strategies, including the universal lattice Boltzmann method and sparse field method, can further reduce the time cost with an unlimited time step and narrow band computation. Experimental results illustrate that the proposed algorithm achieves satisfactory performance for auroral oval segmentation within a very short processing time.
Xi Yang 0011, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001
IEEE Trans. Image Process.1