Yinghui Xing

dblp:218/2673 · DBLP profile ↗
← Back
41ranked-venue papers
16as first author
40since 2021 · last 2026
0000-0001-6021-8261ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 22 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 9 first-author · 12 since 2021
YearPublicationVenuePosition
2026 Attention Retention for Continual Learning with Vision Transformers
abstract
Continual learning (CL) empowers AI systems to progressively acquire knowledge from non-stationary data streams. However, catastrophic forgetting remains a critical challenge. In this work, we identify attention drift in Vision Transformers as a primary source of catastrophic forgetting, where the attention to previously learned visual concepts shifts significantly after learning new tasks. Inspired by neuroscientific insights into the selective attention in the human visual system, we propose a novel attention-retaining framework to mitigate forgetting in CL. Our method constrains attention drift by explicitly modifying gradients during backpropagation through a two-step process: 1) extracting attention maps of the previous task using a layer-wise rollout mechanism and generating instance-adaptive binary masks, and 2) when learning a new task, applying these masks to zero out gradients associated with previous attention regions, thereby preventing disruption of learned visual concepts. For compatibility with modern optimizers, the gradient masking process is further enhanced by scaling parameter updates proportionally to maintain their relative magnitudes. Experiments and visualizations demonstrate the effectiveness of our method in mitigating catastrophic forgetting and preserving visual concepts. It achieves state-of-the-art performance and exhibits robust generalizability across diverse CL scenarios.
Yue Lu 0008, Xiangyu Zhou 0002, Shizhou Zhang, Yinghui Xing, Guoqiang Liang 0001, Wencong Zhang
AAAI4
2026 Better Matching, Less Forgetting: A Quality-Guided Matcher for Transformer-based Incremental Object Detection
abstract
Incremental Object Detection (IOD) aims to continuously learn new object classes without forgetting previously learned ones. A persistent challenge is catastrophic forgetting, primarily attributed to background shift in conventional detectors. While pseudo-labeling mitigates this in dense detectors, we identify a novel, distinct source of forgetting specific to DETR-like architectures: background foregrounding. This arises from the exhaustiveness constraint of the Hungarian matcher, which forcibly assigns every ground truth target to one prediction, even when predictions primarily cover background regions (i.e., low IoU). This erroneous supervision compels the model to misclassify background features as specific foreground classes, disrupting learned representations and accelerating forgetting. To address this, we propose a Quality-guided Min-Cost Max-Flow (Q-MCMF) matcher. To avoid forced assignments, Q-MCMF builds a flow graph and prunes implausible matches based on geometric quality. It then optimizes for the final matching that minimizes cost and maximizes valid assignments. This strategy eliminates harmful supervision from background foregrounding while maximizing foreground learning signals. Extensive experiments on the COCO dataset under various incremental settings demonstrate that our method consistently outperforms existing state-of-the-art approaches.
Qirui Wu, Shizhou Zhang, De Cheng, Yinghui Xing, Lingyan Ran, Dahu Shi, Peng Wang 0015
AAAI4
2026 DuGI-MAE: Improving Infrared Mask Autoencoders via Dual-Domain Guidance
abstract
Infrared imaging plays a critical role in low-light and adverse weather conditions. However, due to the distinct characteristics of infrared images, existing foundation models such as Masked Autoencoder (MAE) trained on visible data perform suboptimal in infrared image interpretation tasks. To bridge this gap, an infrared foundation model known as InfMAE was developed and pre-trained on large-scale infrared datasets. Despite its effectiveness, InfMAE still faces several limitations, including the omission of informative tokens, insufficient modeling of global associations, and neglect of non-uniform noise. In this paper, we propose a Dual-domain Guided Infrared foundation model based on MAE (DuGI-MAE). First, we design a deterministic masking strategy based on token entropy, preserving only high-entropy tokens for reconstruction to enhance informativeness. Next, we introduce a Dual-Domain Guidance (DDG) module, which simultaneously captures global token relationships and adaptively filters non-uniform background noise commonly present in infrared imagery. To facilitate large-scale pretraining, we construct Inf-590K, a comprehensive infrared image dataset encompassing diverse scenes, various target types, and multiple spatial resolutions. Pretrained on Inf-590K, DuGI-MAE demonstrates strong generalization capabilities across various downstream tasks, including infrared object detection, semantic segmentation, and small target detection. Experimental results validate the superiority of the proposed method over both supervised and self-supervised comparison methods.
Yinghui Xing, Xiaoting Su, Shizhou Zhang, Donghao Chu, Di Xu 0010
AAAI1
2026 YOLO-IOD: Towards Real Time Incremental Object Detection
abstract
Current methodologies for incremental object detection (IOD) primarily rely on Faster R-CNN or DETR series detectors; however, these approaches do not accommodate the real-time YOLO detection frameworks. In this paper, we first identify three primary types of knowledge conflicts that contribute to catastrophic forgetting in YOLO-based incremental detectors: foreground-background confusion, parameter interference, and misaligned knowledge distillation. Subsequently, we introduce YOLO-IOD, a real-time Incremental Object Detection (IOD) framework that is constructed upon the pretrained YOLO-World model, facilitating incremental learning via a stage-wise parameter-efficient finetuning process. Specifically, YOLO-IOD encompasses three principal components: 1) Conflict-Aware Pseudo-Label Refinement (CPR), which mitigates the foreground-background confusion by leveraging the confidence levels of pseudo labels and identifying potential objects relevant to future tasks. 2) Importance-based Kernel Selection (IKS), which identifies and updates the pivotal convolution kernels pertinent to the current task during the current learning stage. 3)Cross-Stage Asymmetric Knowledge Distillation (CAKD), which addresses the misaligned knowledge distillation conflict by transmitting the features of the student target detector through the detection heads of both the previous and current teacher detectors, thereby facilitating asymmetric distillation between existing and newly introduced categories. We further introduce LoCo COCO, a more realistic benchmark that eliminates data leakage across stages. Experiments on both conventional and LoCo COCO benchmarks show that YOLO-IOD achieves superior performance with minimal forgetting.
Shizhou Zhang, Xueqiang Lv, Yinghui Xing, Qirui Wu, Di Xu 0010, Chen Zhao 0009, Yanning Zhang 0001
AAAI3
2026 VPT-NSP2++: Importance-Aware Visual Prompt Tuning in Null Space for Continual Learning
abstract
Continual learning (CL) enables AI models to adapt to evolving environments while mitigating catastrophic forgetting, which is a critical capability for dynamic real-world applications. With the growing popularity of pre-trained Vision Transformer (ViT) models and visual prompt tuning (VPT) technique in CL, this work explores a CL method on top of the ViT-based foundation model, through VPT mechanism with theoretical guarantees. Inspired by the orthogonal projection method, we aim to leverage this approach for VPT to enhance CL performance, particularly in long-term scenarios. However, since the orthogonal projection is originally designed for linear operations in CNNs, applying it to ViTs poses challenges induced by the non-linear self-attention mechanism and the distribution drift within LayerNorm. To address these issues, we deduced two orthogonality conditions to achieve the prompt gradient orthogonal projection, which provide a theoretical guarantee of maintaining stability. Considering the strict orthogonal constraints can diminish model capacity and reduce plasticity, we further propose an importance-aware orthogonal regularization framework. By applying varying degrees of orthogonal constraints to different parameters based on their importance to old and new tasks, the framework adaptively enhances model capacity and thereby promotes long-sequence CL while improving the stability-plasticity trade-off. To implement the proposed approach, a null-space-based approximation solution is employed to efficiently achieve the prompt gradient orthogonal projection. Extensive experiments on various class-incremental learning benchmarks demonstrate that our method achieves state-of-the-art performance across diverse CL scenarios.
Shizhou Zhang, Yue Lu 0008, De Cheng, Yinghui Xing, Nannan Wang 0001, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Nearest-neighbor class prototype prompt and simulated logits for continual learning
Yue Lu 0008, Shizhou Zhang, Yinghui Xing, Guoqiang Liang 0001, Yanning Zhang 0001
Pattern Recognit.4
2026 Less Is More: Infrared and Visible Images Fusion via Semantic-Guided Mixture of Multi-Feature Experts
abstract
Infrared (IR) and visible image fusion (IVIF) has become prevalent in recent years. By leveraging the complementary characteristics of infrared and visible images, we can obtain visually-appealing fused images, which further facilitate subsequent scene understanding and object detection from day to night. Integrating complementary information while simultaneously eliminating redundancy is a crucial challenge in fusion. Most of available deep learning based methods, after being trained, execute static inference on all pairs of infrared and visible images. They struggle to effectively handle redundancy of modality across diverse scenarios, resulting in superfluous information such as thermal noise in infrared images and artifacts in visible images. In this paper, we propose an IVIF method based on a semantic-guided mixture of multi-feature experts, where multiple types of features are extracted, each assigned to a dedicated expert network specialized in processing a specific type of features. Through an expert routing mechanism, these experts are chosen dynamically, ensuring that the most significant features of each image modality are routed to a specific group of experts. In order to align fusion task with subsequent semantic segmentation task, we introduce a segmentation head to semantically guide the selection of the complementary features. Extensive experiments on five infrared and visible image fusion and segmentation benchmarks demonstrate the effectiveness of our method, both for image fusion and subsequent semantic segmentation tasks. The code will be available at https://github.com/ZhilongNiu/SD-MoMFE.
Yinghui Xing, Zhilong Niu, Shizhou Zhang, Yanning Zhang 0001
IEEE Trans. Image Process.1
2025 Training Consistent Mixture-of-Experts-Based Prompt Generator for Continual Learning
abstract
Visual prompt tuning-based continual learning (CL) methods have shown promising performance in exemplar-free scenarios, where their key component can be viewed as a prompt generator. Existing approaches generally rely on freezing old prompts, slow updating and task discrimination for prompt generators to preserve stability and minimize forgetting. In contrast, we introduce a novel approach that trains a consistent prompt generator to ensure stability during CL. Consistency means that for any instance from an old task, its corresponding instance-ware prompt generated by the prompt generator remains consistent even as the generator continually updates in a new task. This ensures that the representation of a specific instance remains stable across tasks and thereby prevents forgetting. We employ a mixture of experts (MoE) as the prompt generator, which contains a router and multiple experts. By deriving conditions sufficient to achieve the consistency for the MoE prompt generator, we demonstrate that: during training in a new task, if the router and experts update in the directions orthogonal to the subspaces spanned by old input features and gating vectors, respectively, the consistency can be theoretically guaranteed. To implement this orthogonality, we project parameter gradients to those orthogonal directions using the orthogonal projection matrices computed via the null space method. Extensive experiments on four class-incremental learning benchmarks validate the effectiveness and superiority of our approach.
Yue Lu 0008, Shizhou Zhang, De Cheng, Guoqiang Liang 0001, Yinghui Xing, Nannan Wang 0001, Yanning Zhang 0001
AAAI5
2025 Dual-Granularity Semantic Guided Sparse Routing Diffusion Model for General Pansharpening
abstract
Pansharpening aims at integrating complementary information from panchromatic and multispectral images. Available deep-learning based pansharpening methods typically perform exceptionally with particular satellite datasets. At the same time, it has been observed that these models also exhibit scene dependence, for example, if the majority of the training samples come from the urban scenes, the model’s performance may decline in the river scene. To address the domain gap produced by varying satellite sensors and distinct scenes, we propose a dual-granularity semantic guided sparse routing diffusion model for general pansharpening. By utilizing the large Vision-Language Models (VLMs) in the field of geoscience, e.g, GeoChat, we introduce the dual granularity semantics to generate dynamic sparse routing scores for adaptation of different satellite sensors and scenes. This scene-level and region-level dual-granularity semantic information serves as guidance for dynamically activating specialized experts within the diffusion model. Extensive experiments on WorldView-3, QuickBird, and GaoFen-2 datasets show the effectiveness of our proposed method. Notably, the proposed method outperforms the comparison approaches in adapting to new satellite sensors and scenes. The codes are available at https://github.com/codgodtao/SGDiff.
Yinghui Xing, Litao Qu, Shizhou Zhang, Di Xu 0010, Yingkun Yang, Yanning Zhang 0001
CVPR1
2025 Revisiting Generative Replay for Class Incremental Object Detection
abstract
Generative replay has gained significant attention in class-incremental learning; however, its application to Class Incremental Object Detection (CIOD) remains limited due to the challenges in generating complex images with precise spatial arrangements. In this study, motivated by the observation that the forgetting of prior knowledge is predominantly present in the classification sub-task as opposed to the localization sub-task, we revisit the generative replay method for class incremental object detection. Our method utilize a standard Stable Diffusion model to generate image-level replay data for all old and new tasks. Accordingly, the old detector and a stage-wise detector are conducted on the synthetic images respectively to determine the bounding box positions through pseudo-labeling. Furthermore, we propose to use a Similarity-based Cross Sampling mechanism to select valuable confusing data between old and new tasks to more effectively mitigate catastrophic forgetting and reduce the false alarm rate for the new task. Finally, all synthetic and real data are integrated for current-stage detector training, where the images generated for previous tasks are highly beneficial in minimizing the forgetting of existing knowledge, while those synthesized for the new task can help bridge the domain gap between real and synthetic images. We conducted extensive experiments on PASCAL VOC 2007 and MS COCO benchmark datasets in multiple settings to showcase the efficacy of our proposed approach, which achieves state-of-the-art results. The code is available at https://github.com/qiangzailv/RGR-IOD.
Shizhou Zhang, Xueqiang Lv, Yinghui Xing, Qirui Wu, Di Xu 0010, Yanning Zhang 0001
CVPR3
2025 Gradient Decomposition and Alignment for Incremental Object Detection
Wenlong Luo, Shizhou Zhang, De Cheng, Yinghui Xing, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001
ICCV4
2025 Demystifying Catastrophic Forgetting in Two-Stage Incremental Object Detector
abstract
Catastrophic forgetting is a critical chanllenge for incremental object detection (IOD). Most existing methods treat the detector monolithically, relying on instance replay or knowledge distillation without analyzing component-specific forgetting. Through dissection of Faster R-CNN, we reveal a key insight: Catastrophic forgetting is predominantly localized to the RoI Head classifier, while regressors retain robustness across incremental stages. This finding challenges conventional assumptions, motivating us to develop a framework termed NSGP-RePRE. Regional Prototype Replay (RePRE) mitigates classifier forgetting via replay of two types of prototypes: coarse prototypes represent class-wise semantic centers of RoI features, while fine-grained prototypes model intra-class variations. Null Space Gradient Projection (NSGP) is further introduced to eliminate prototype-feature misalignment by updating the feature extractor in directions orthogonal to subspace of old inputs via gradient projection, aligning RePRE with incremental learning dynamics. Our simple yet effective design allows NSGP-RePRE to achieve state-of-the-art performance on the Pascal VOC and MS COCO datasets under various settings. Our work not only advances IOD methodology but also provide pivotal insights for catastrophic forgetting mitigation in IOD. Code will be available soon.
Qirui Wu, Shizhou Zhang, De Cheng, Yinghui Xing, Di Xu 0010, Peng Wang 0015, Yanning Zhang 0001
ICML4
2025 Amplitude-aware Domain Style Replay for Lifelong Person Re-identification
abstract
Lifelong Person Re-identification (LReID) focuses on continuously adapting to new domains over time while preserving knowledge from previously seen domains, particularly under the domain incremental learning setting. The major challenge of LReID is catastrophic forgetting, typically caused by large domain shifts during training. To address this, we propose a novel Amplitude-aware Domain Style Replay (ADSR) framework, which introduces a Fourier-based Style Transfer (FST) mechanism to generate synthetic data that reflects the style of previously encountered domains. These proxy images help retain prior knowledge without the need to store actual past data. Our method transfers stylistic information-mainly encoded in the amplitude spectrum-from old domains to new ones, creating old-stylized images that preserve the content of new domain data while adopting the visual style of earlier domains. To further boost generalization, we design a Self-Stylization Normalization (SSN) module that adapts the current domain's style distribution, making the model more robust to stylistic variations. Additionally, we introduce a Multi-Granularity Transfer (MGT) module that uses K-Means clustering to extract multiple representative style features from each domain, enabling compact yet comprehensive storage and replay of domain-specific information. Extensive experiments on multiple LReID benchmarks show that ADSR achieves superior performance over existing approaches, effectively reducing forgetting and improving cross-domain generalization. Our code is available at https://github.com/cclong8/MM2025-ADSR.
De Cheng, Shizhou Zhang, Yinghui Xing, Di Xu 0010, Yanning Zhang 0001
ACM Multimedia4
2025 Prompt-Based Modality Alignment for Effective Multi-Modal Object Re-Identification
abstract
A critical challenge for multi-modal Object Re-Identification (ReID) is the effective aggregation of complementary information to mitigate illumination issues. State-of-the-art methods typically employ complex and highly-coupled architectures, which unavoidably result in heavy computational costs. Moreover, the significant distribution gap among different image spectra hinders the joint representation of multi-modal features. In this paper, we propose a framework named as PromptMA to establish effective communication channels between different modality paths, thereby aggregating modal complementary information and bridging the distribution gap. Specifically, we inject a series of learnable multi-modal prompts into the Image Encoder and introduce a prompt exchange mechanism to enable the prompts to alternately interact with different modal token embeddings, thus capturing and distributing multi-modal features effectively. Building on top of the multi-modal prompts, we further propose Prompt-based Token Selection (PBTS) and Prompt-based Modality Fusion (PBMF) modules to achieve effective multi-modal feature fusion while minimizing background interference. Additionally, due to the flexibility of our prompt exchange mechanism, our method is well-suited to handle scenarios with missing modalities. Extensive evaluations are conducted on four widely used benchmark datasets and the experimental results demonstrate that our method achieves state-of-the-art performances, surpassing the current benchmarks by over 15% on the challenging MSVR310 dataset and by 6% on the RGBNT201. The code is available at https://github.com/FHR-L/PromptMA.
Shizhou Zhang, Wenlong Luo, De Cheng, Yinghui Xing, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Image Process.4
2025 Frequency-Guided Spatial Adaptation for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) aims to segment camouflaged objects which exhibit very similar patterns with the surrounding environment. Recent research works have shown that enhancing the feature representation via the frequency information can greatly alleviate the ambiguity problem between the foreground objects and the background. With the emergence of vision foundation models, like InternImage, Segment Anything Model etc, adapting the pretrained model on COD tasks with a lightweight adapter module shows a novel and promising research direction. Existing adapter modules mainly care about the feature adaptation in the spatial domain. In this paper, we propose a novel frequency-guided spatial adaptation method for COD task. Specifically, we transform the input features of the adapter into frequency domain. By grouping and interacting with frequency components located within non overlapping circles in the spectrogram, different frequency components are dynamically enhanced or weakened, making the intensity of image details and contour features adaptively adjusted. At the same time, the features that are conducive to distinguishing object and background are highlighted, indirectly implying the position and shape of camouflaged object. We conduct extensive experiments on four widely adopted benchmark datasets and the proposed method outperforms 26 state-of-the-art methods with large margins. Code will be released.
Shizhou Zhang, Dexuan Kong, Yinghui Xing, Yue Lu 0008, Lingyan Ran, Guoqiang Liang 0001, Hexu Wang, Yanning Zhang 0001
IEEE Trans. Multim.3
2024 Cross-Platform Video Person ReID: A New Benchmark Dataset and Adaptation Approach
Shizhou Zhang, Wenlong Luo, De Cheng, Qingchun Yang, Lingyan Ran, Yinghui Xing, Yanning Zhang 0001
ECCV (27)6
2024 Complementary Fusion Network Based on Frequency Hybrid Attention for Pansharpening
abstract
Pansharpening is a feasible way to obtain the high-resolution (HR) multispectral (MS) images by using panchromatic (PAN) images to sharpen low-resolution MS images. Despite its great advances, most existing pansharpening methods neglect the importance of integrating local and non-local characteristics of images, resulting in the imbalance of spatial and spectral distribution. In this paper, we propose a complementary fusion network (CFNet) based on frequency hybrid attention mechanism for pansharpening. By introducing the frequency transformation and the deformable cross-attention, our model takes image-wide receptive field into consideration to explore global feature learning. Combined with the convolutional layers with local receptive field, CFNet can well capture local and non-local features. Experimental results demonstrate that the proposed method outperforms the comparison methods in terms of visual and quantitative qualities.
Yinghui Xing, Litao Qu, Kai Zhang 0010, Yan Zhang 0127, Xiuwei Zhang 0001, Yanning Zhang 0001
ICASSP1
2024 Dual-Branch Task Residual Enhancement with Parameter-Free Attention for Zero-Shot Multi-label Image Recognition
Shizhou Zhang, Kairui Dang, De Cheng, Yinghui Xing, Qirui Wu, Dexuan Kong, Yanning Zhang 0001
ICPR (22)4
2024 Visual Prompt Tuning in Null Space for Continual Learning
abstract
Existing prompt-tuning methods have demonstrated impressive performances in continual learning (CL), by selecting and updating relevant prompts in the vision-transformer models. On the contrary, this paper aims to learn each task by tuning the prompts in the direction orthogonal to the subspace spanned by previous tasks' features, so as to ensure no interference on tasks that have been learned to overcome catastrophic forgetting in CL. However, different from the orthogonal projection in the traditional CNN architecture, the prompt gradient orthogonal projection in the ViT architecture shows completely different and greater challenges, i.e., 1) the high-order and non-linear self-attention operation; 2) the drift of prompt distribution brought by the LayerNorm in the transformer block. Theoretically, we have finally deduced two consistency conditions to achieve the prompt gradient orthogonal projection, which provide a theoretical guarantee of eliminating interference on previously learned knowledge via the self-attention mechanism in visual prompt tuning. In practice, an effective null-space-based approximation solution has been proposed to implement the prompt gradient orthogonal projection. Extensive experimental results demonstrate the effectiveness of anti-forgetting on four class-incremental benchmarks with diverse pre-trained baseline models, and our approach achieves superior performances to state-of-the-art methods. Our code is available at https://github.com/zugexiaodui/VPTinNSforCL
Yue Lu 0008, Shizhou Zhang, De Cheng, Yinghui Xing, Nannan Wang 0001, Peng Wang 0015, Yanning Zhang 0001
NeurIPS4
2024 DDF: A Novel Dual-Domain Image Fusion Strategy for Remote Sensing Image Semantic Segmentation With Unsupervised Domain Adaptation
abstract
The semantic segmentation of remote sensing (RS) images is a challenging and hot issue due to the large amount of unlabeled data and domain variation. Unsupervised domain adaptation (UDA) has proven to be advantageous in leveraging unlabeled information from the target domain. However, traditional approaches of independently fine-tuning UDA models in the source and target domains have a limited effect on the result. In this article, we propose a hybrid training strategy that boosts self-training methods with domain fusion images. First, we introduce a novel dual-domain image fusion (DDF) strategy to effectively utilize the original image, the style-transferred image, and the intermediate-domain information. Second, to further refine the precision of pseudolabels, we present a region-specific reweighting strategy that assigns different weights to pseudolabel regions based on their spatial context. Finally, we conduct a series of extensive benchmark experiments and ablation studies on the ISPRS Vaihingen and Potsdam datasets. These results show the efficiency of our approach and establish a practical basis for implementing semantic segmentation in remote sensors.
Lingyan Ran, Lushuang Wang, Tao Zhuo, Yinghui Xing, Yanning Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Empower Generalizability for Pansharpening Through Text-Modulated Diffusion Model
abstract
Pansharpening is crucial to remote sensing applications by fusing high-resolution (HR) panchromatic (PAN) images with low-resolution multispectral (LRMS) images to generate HR multispectral (HRMS) images. Recently, diffusion probabilistic models (DPMs) have provided high-quality results than regression-based methods when trained on specific pairwise data for their specific purpose. However, their performance degrades when applied to a new satellite dataset, which represents different imaging properties and spectral ranges, limiting the generalization ability of them. For better generalizability of pansharpening, in this article, we propose a text-modulated diffusion model (TMDiff) for unified pansharpening of different satellites. TMDiff takes a text-modulated 3-D UNet (TM3DU) as denoising network to gradually recover HRMS through iterative refinement over multiple time steps. By introducing satellite’s physical properties as text prompts, TM3DU is able to learn meta-knowledge across different satellites and thus can sharpen LRMS images with diverse spatial and spectral attributes. Extensive experiments on various satellite datasets demonstrate the state-of-the-art performance of our model in both qualitative and quantitative metrics. Furthermore, our model exhibits superior generalization ability to unseen datasets, highlighting its practical significance. Code is available athttps://github.com/codgodtao/TMDiff.
Yinghui Xing, Litao Qu, Shizhou Zhang, Jiapeng Feng, Xiuwei Zhang 0001, Yanning Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Improving Reliability of Heterogeneous Change Detection by Sample Synthesis and Knowledge Transfer
abstract
Detecting changes in heterogeneous images without the supervision of changed label is a challenging yet critical task for quick responding natural disaster relief. Nevertheless, most of available unsupervised heterogeneous change detection methods strong rely on the quality of pseudo labels, and they suffer from performance degradation, even irreversible model collapse, when encounter the low-quality pseudo labels, leading to unreliable detection results. In order to improve the reliability of unsupervised heterogeneous change detection, in this paper, we propose a novel change detection paradigm based on sample synthesis and knowledge transfer. We address the issue of label reliability by artificially creating a changed region and assigning labels rather than constructing pseudo labels. These constructed labels guide the network in automatically learning the correspondence between heterogeneous images, confirming the reliability of changed regions. Moreover, an augmentation with synthetic samples on real samples makes it possible to generate more transferable samples while reducing the domain gap coarsely. A dual-branch joint training with feature contrastive learning is further developed to transfer the knowledge of changes from the synthetic sample domain to real sample domain. Experimental results on five public datasets demonstrate that our proposed method has superior performance when compared with available state-of-the-art methods. Our code is available at https://github.com/zhangqiiii/SS-KT.
Yinghui Xing, Lingyan Ran, Xiuwei Zhang 0001, Hanlin Yin, Yanning Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 SCAFNet: Semantic-Guided Cascade Adaptive Fusion Network for Infrared Small Target Detection
abstract
Infrared small target detection is a crucial component of infrared target tracking and search. It is challenging due to the complex backgrounds, low contrast between targets and backgrounds, and the small, dim nature of the targets. Therefore, effectively representing the targets and enhancing the distinction between targets and backgrounds is essential. Existing deep-learning (DL)-based methods struggle to capture the subtle details of weak targets, neglecting the complementary characteristics of multilevel features, which leads to inaccurate localization of targets. In this article, we propose a semantic-guided cascade adaptive fusion network (SCAFNet) to address these challenges. To improve the representation of small targets in the deeper layers, we introduce a multiresolution auxiliary enhancement (MAE) encoder to progressively enhance detailed information within the deep features. After extracting multiscale features, an adaptive fusion (AdaFus) decoder is proposed to fuse them. It has a semantic-guided cascade fusion (SGCF) module to integrate feature maps at three different resolutions. Specifically, SGCF first employs rich semantic features from the high-level feature map to guide the spatial distribution of the low-level feature maps, thereby improving the distinction between the target and the background. Then, AdaFus weights are generated to guide the fusion process, ensuring that the final feature map combines rich semantic information with precise spatial details. Furthermore, we perform long-distance modeling on the feature map to achieve detailed reconstruction, which aids in restoring the shape information of the target. The effectiveness of our method is validated through experiments on various public infrared small target detection datasets.
Shizhou Zhang, Yinghui Xing, Liangkui Lin, Xiaoting Su, Yanning Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 CrossDiff: Exploring Self-SupervisedRepresentation of Pansharpening via Cross-Predictive Diffusion Model
abstract
Fusion of a panchromatic (PAN) image and corresponding multispectral (MS) image is also known as pansharpening, which aims to combine abundant spatial details of PAN and spectral information of MS images. Due to the absence of high-resolution MS images, available deep-learning-based methods usually follow the paradigm of training at reduced resolution and testing at both reduced and full resolution. When taking original MS and PAN images as inputs, they always obtain sub-optimal results due to the scale variation. In this paper, we propose to explore the self-supervised representation for pansharpening by designing a cross-predictive diffusion model, named CrossDiff. It has two-stage training. In the first stage, we introduce a cross-predictive pretext task to pre-train the UNet structure based on conditional Denoising Diffusion Probabilistic Model (DDPM). While in the second stage, the encoders of the UNets are frozen to directly extract spatial and spectral features from PAN and MS images, and only the fusion head is trained to adapt for pansharpening task. Extensive experiments show the effectiveness and superiority of the proposed model compared with state-of-the-art supervised and unsupervised methods. Besides, the cross-sensor experiments also verify the generalization ability of proposed self-supervised representation learners for other satellite datasets. Code is available at https://github.com/codgodtao/CrossDiff.
Yinghui Xing, Litao Qu, Shizhou Zhang, Kai Zhang 0010, Yanning Zhang 0001, Lorenzo Bruzzone
IEEE Trans. Image Process.1
2024 MS-DETR: Multispectral Pedestrian Detection Transformer With Loosely Coupled Fusion and Modality-Balanced Optimization
abstract
Multispectral pedestrian detection is an important task for many around-the-clock applications, since the visible and thermal modalities can provide complementary information especially under low light conditions. Due to the presence of two modalities, misalignment and modality imbalance are the most significant issues in multispectral pedestrian detection. In this paper, we propose MultiSpectral pedestrian DEtection TRansformer (MS-DETR) to fix above issues. MS-DETR consists of two modality-specific backbones and Transformer encoders, followed by a multi-modal Transformer decoder, and the visible and thermal features are fused in the multi-modal Transformer decoder. To well resist the misalignment between multi-modal images, we design a loosely coupled fusion strategy by sparsely sampling some keypoints from multi-modal features independently and fusing them with adaptively learned attention weights. Moreover, based on the insight that not only different modalities, but also different pedestrian instances tend to have different confidence scores to final detection, we further propose an instance-aware modality-balanced optimization strategy, which preserves visible and thermal decoder branches and aligns their predicted slots through an instance-wise dynamic loss. Our end-to-end MS-DETR shows superior performance on the challenging KAIST, CVC-14 and LLVIP benchmark datasets. The source code is available athttps://github.com/YinghuiXing/MS-DETR.
Yinghui Xing, Song Wang 0002, Shizhou Zhang, Guoqiang Liang 0001, Xiuwei Zhang 0001, Yanning Zhang 0001
IEEE Trans. Intell. Transp. Syst.1
2024 Dual Modality Prompt Tuning for Vision-Language Pre-Trained Model
abstract
With the emergence of large pretrained vison-language models such as CLIP, transferable representations can be adapted to a wide range of downstream tasks via prompt tuning. Prompt tuning probes for beneficial information for downstream tasks from the general knowledge stored in the pretrained model. A recently proposed method named Context Optimization (CoOp) introduces a set of learnable vectors as text prompts from the language side. However, tuning the text prompt alone can only adjust the synthesized “classifier”, while the computed visual features of the image encoder cannot be affected, thus leading to suboptimal solutions. In this article, we propose a novel dual-modality prompt tuning (DPT) paradigm through learning text and visual prompts simultaneously. To make the final image feature concentrate more on the target visual concept, a class-aware visual prompt tuning (CAVPT) scheme is further proposed in our DPT. In this scheme, the class-aware visual prompt is generated dynamically by performing the cross attention between text prompt features and image patch token embeddings to encode both the downstream task-related information and visual instance information. Extensive experimental results on 11 datasets demonstrate the effectiveness and generalization ability of the proposed method.
Yinghui Xing, Qirui Wu, De Cheng, Shizhou Zhang, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Multim.1
2023 SSML-QNet: Scale-Separative Metric Learning Quadruplet Network for Multi-modal Image Patch Matching
abstract
Multi-modal image matching is very challenging due to the significant diversities in visual appearance of different modal images. Typically, the existing well-performed methods mainly focus on learning invariant and discriminative features for measuring the relation between multi-modal image pairs. However, these methods often take the features as a whole and largely overlook the fact that different scale features for a same image pair may have different similarity, which may lead to sub-optimal results only. In this work, we propose a Scale-Separative Metric Learning Quadruplet network (SSML-QNet) for multi-modal image patch matching. Specifically, SSML-QNet can extract both relevant and irrelevant features of imaging modality with the proposed quadruplet network architecture. Then, the proposed Scale-Separative Metric Learning module separately encodes the similarity of different scale features with the pyramid structure. And for each scale, cross-modal consistent features are extracted and measured by coordinate and channel-wise attention sequentially. This makes our network robust to appearance divergence caused by different imaging mechanism. Experiments on the benchmark dataset (VIS-NIR, VIS-LWIR, Optical-SAR, and Brown) have verified that the proposed SSML-QNet is able to outperform other state-of-the-art methods. Furthermore, the cross-dataset transferring experiments on these four datasets also have shown that the proposed method has powerful ability of cross-dataset transferring.
Xiuwei Zhang 0001, Hanlin Yin, Yinghui Xing, Yanning Zhang 0001
IJCAI6
2023 Automatic Network Architecture Search for RGB-D Semantic Segmentation
abstract
Recent RGB-D semantic segmentation networks are usually manually designed. However, due to limited human efforts and time costs, their performance might be inferior for complex scenarios. To address this issue, we propose the first Neural Architecture Search (NAS) method that designs the network automatically. Specifically, the target network consists of an encoder and a decoder. The encoder is designed with two independent branches, where each branch specializes in extracting features from RGB and depth images, respectively. The decoder fuses the features and generates the final segmentation result. Besides, for automatic network design, we design a grid-like network-level search space combined with a hierarchical cell-level search space. By further developing an effective gradient-based search strategy, the network structure with hierarchical cell architectures is discovered. Extensive results on two datasets show that the proposed method outperforms the state-of-the-art approaches, which achieves a mIoU score of 55.1% on the NYU-Depth v2 dataset and 50.3% on the SUN-RGBD dataset.
Wenna Wang, Tao Zhuo, Xiuwei Zhang 0001, Mingjun Sun, Hanlin Yin, Yinghui Xing, Yanning Zhang 0001
ACM Multimedia6
2023 Ground-to-Aerial Person Search: Benchmark Dataset and Approach
abstract
In this work, we construct a large-scale dataset for Ground-to-Aerial Person Search, named G2APS, which contains 31,770 images of 260,559 annotated bounding boxes for 2,644 identities appearing in both of the UAVs and ground surveillance cameras. To our knowledge, this is the first dataset for cross-platform intelligent surveillance applications, where the UAVs could work as a powerful complement for the ground surveillance cameras. To more realistically simulate the actual cross-platform Ground-to-Aerial surveillance scenarios, the surveillance cameras are fixed about 2 meters above the ground, while the UAVs capture videos of persons at different location, with a variety of view-angles, flight attitudes and flight modes. Therefore, the dataset has the following unique characteristics: 1) drastic view-angle changes between query and gallery person images from cross-platform cameras; 2) diverse resolutions, poses and views of the person images under 9 rich real-world scenarios. On basis of the G2APS benchmark dataset, we demonstrate detailed analysis about current two-step and end-to-end person search methods, and further propose a simple yet effective knowledge distillation scheme on the head of the ReID network, which achieves state-of-the-art performances on both of the G2APS and the previous two public person search datasets, i.e., PRW and CUHK-SYSU. The dataset and source code available on https://github.com/yqc123456/HKD_for_person_search.
Shizhou Zhang, Qingchun Yang, De Cheng, Yinghui Xing, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001
ACM Multimedia4
2023 AugFCOS: Augmented fully convolutional one-stage object detection network
Xiuwei Zhang 0001, Yinghui Xing, Wenna Wang, Hanlin Yin, Yanning Zhang 0001
Pattern Recognit.3
2023 FastICENet: A real-time and accurate semantic segmentation model for aerial remote sensing river ice image
Xiuwei Zhang 0001, Lingyan Ran, Yinghui Xing, Wenna Wang, Zeze Lan, Hanlin Yin, Houjun He, Qixing Liu, Baosen Zhang, Yanning Zhang 0001
Signal Process.4
2023 Pansharpening via Frequency-Aware Fusion Network With Explicit Similarity Constraints
abstract
The process of fusing a high spatial resolution (HR) panchromatic (PAN) image and a low spatial resolution (LR) multispectral (MS) image to obtain an HRMS image is known as pansharpening. With the development of convolutional neural networks, the performance of pansharpening methods has been improved, however, the blurry effects and the spectral distortion still exist in their fusion results due to the insufficiency in details learning and the frequency mismatch between MS and PAN. Therefore, the improvement of spatial details at the premise of reducing spectral distortion is still a challenge. In this paper, we propose a frequency-aware fusion network (FAFNet) together with a novel high-frequency feature similarity loss to address above mentioned problems. FAFNet is mainly composed of two kinds of blocks, where the frequency aware blocks aim to extract features in the frequency domain with the help of discrete wavelet transform (DWT) layers, and the frequency fusion blocks reconstruct and transform the features from frequency domain to spatial domain with the assistance of inverse DWT (IDWT) layers. Finally, the fusion results are obtained through a convolutional block. In order to learn the correspondence, we also propose a high-frequency feature similarity loss to constrain the HF features derived from PAN and MS branches, so that HF features of PAN can reasonably be used to supplement that of MS. Experimental results on three datasets at both reduced- and full-resolution demonstrate the superiority of the proposed method compared with several state-of-the-art pansharpening models. The codes are available at https://github.com/YinghuiXing/FAFNet.
Yinghui Xing, Yan Zhang 0127, Houjun He, Xiuwei Zhang 0001, Yanning Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Progressive Modality-Alignment for Unsupervised Heterogeneous Change Detection
abstract
Change detection based on heterogeneous images is of great importance in some applications, such as disaster monitoring and damage assessment. However, due to the huge modality discrepancy in heterogeneous images, it is difficult to accurately detect the changed regions. In this paper, we analyze the interference of modality-alignment and changed areas to each other, and propose a progressive modality-alignment based unsupervised change detection model for heterogeneous images. Specifically, the modality alignment is achieved in an iterative manner, which can improve the detection accuracy progressively. To reduce the influence of modality discrepancy and the changed regions to each other, a pseudo-label self-learning strategy is designed, where the pseudo-labels learned by the model itself are used to act as a guidance of change detection, and they are in turn refined by the proposed progressive model. Experimental results on different real heterogeneous images verify the effectiveness and robustness of proposed method.
Yinghui Xing, Lingyan Ran, Xiuwei Zhang 0001, Hanlin Yin, Yanning Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 ZeRGAN: Zero-Reference GAN for Fusion of Multispectral and Panchromatic Images
abstract
In this article, we present a new pansharpening method, a zero-reference generative adversarial network (ZeRGAN), which fuses low spatial resolution multispectral (LR MS) and high spatial resolution panchromatic (PAN) images. In the proposed method, zero-reference indicates that it does not require paired reduced-scale images or unpaired full-scale images for training. To obtain accurate fusion results, we establish an adversarial game between a set of multiscale generators and their corresponding discriminators. Through multiscale generators, the fused high spatial resolution MS (HR MS) images are progressively produced from LR MS and PAN images, while the discriminators aim to distinguish the differences of spatial information between the HR MS images and the PAN images. In other words, the HR MS images are generated from LR MS and PAN images after the optimization of ZeRGAN. Furthermore, we construct a nonreference loss function, including an adversarial loss, spatial and spectral reconstruction losses, a spatial enhancement loss, and an average constancy loss. Through the minimization of the total loss, the spatial details in the HR MS images can be enhanced efficiently. Extensive experiments are implemented on datasets acquired by different satellites. The results demonstrate that the effectiveness of the proposed method compared with the state-of-the-art methods. The source code is publicly available at https://github.com/RSMagneto/ZeRGAN.
Wenxiu Diao, Feng Zhang 0028, Jiande Sun 0001, Yinghui Xing, Kai Zhang 0010, Lorenzo Bruzzone
IEEE Trans. Neural Networks Learn. Syst.4
2022 Wavefusion: Wavelet Assistant Fusion Model for Pan-Sharpening
abstract
Pan-sharpening refers to obtain a high-resolution multispectral (HRMS) image by fusing a panchromatic (PAN) image and a low-resolution multispectral (LRMS) image. Recently, convolutional neural networks (CNNs) have achieved great success in pan-sharpening. However, the down-sampling operations in commonly used CNN-based models lead to information loss, and the corresponding up-sampling operations usually introduce some undesirable artifacts, resulting in suboptimal fusion results. In this paper, we propose a simple but effective wavelet assistant fusion model (WaveFusion) to address aforementioned issue. The proposed model consists of three parts, namely a wavelet feature extraction (WFE) part, a wavelet feature fusion (WFF) part and a reconstruction part. With the assistance of the wavelet transform and also a simple alignment operation, WaveFusion obtains the best fusion result compared with some state-of-the-art methods, especially for the fusion at the full resolution.
Yinghui Xing, Yan Zhang 0127, Yanning Zhang 0001
IGARSS1
2022 Hyperspectral and Multispectral Image Fusion via Variational Tensor Subspace Decomposition
abstract
The fusion of hyperspectral image (HSI) and multispectral image (MSI) refers to enhance the spatial resolution of HSI with the help of a corresponding MSI that has a high spatial resolution to finally obtain an HSI with high resolution in both spatial and spectral domains. In this letter, we propose a variational tensor subspace decomposition-based fusion method to fully explore the differences and correlations among three modes of the HSI tensor. Experimental results on two HSI datasets show that the proposed method can achieve superior performance compared with existing state-of-the-art fusion methods with high computational efficiency.
Yinghui Xing, Yan Zhang 0127, Shuyuan Yang 0001, Yanning Zhang 0001
IEEE Geosci. Remote. Sens. Lett.1
2022 SSA-Net: Spatial Scale Attention Network for Image-Based Geo-Localization
abstract
Image-based geo-localization is estimating the location of a query image by matching it to a large amount of images in geo-tagged database. This matching task is very challenging due to the vast differences in visual appearance or modality of image pairs on different platforms, for example, one image from the RGB camera, the other from the light detection and ranging (LiDAR) sensor. The spatial layout of the scene can provide important clues and significantly reduce matching ambiguity. Therefore, we propose a novel deep network that embeds spatial configuration of the scenes into feature representation. Specifically, we design a spatial-scale attention (SSA) module to highlight the salience correspondence layout features at different scales. The encoded features not only represent the emergence of certain objects, but also reflect the relative locations of the objects. By this way, we learn more discriminative deep feature representations, leading to a higher recall. The experimental results on two standard cross-view benchmark datasets (CVUSA and CVACT) and a cross-modal dataset (GRAL) demonstrate that our method performs better than the state-of-the-art methods. Remarkably, the recall rate@top-1 improves from 27.6% to 40.5% on the GRAL dataset.
Xiuwei Zhang 0001, Xiangchuang Meng, Hanlin Yin, Yuanzeng Yue, Yinghui Xing, Yanning Zhang 0001
IEEE Geosci. Remote. Sens. Lett.6
2022 Dual-Collaborative Fusion Model for Multispectral and Panchromatic Image Fusion
abstract
The aim of multispectral (MS) and panchromatic (PAN) image fusion is to obtain an MS image that has high resolution in both spectral and spatial domains. During the fusion process, there are two important issues, i.e., spectral information preservation and spatial information enhancement. In this article, we propose a dual-collaborative fusion model that considers not only the spectral correlation collaboration but also the spatial-spectral collaboration. First, the features of PAN and MS images are extracted by a shared feature embedding network. Then, in order to enhance the spatial details, the PAN features are decomposed into four subbands, and the collaborative relationships among subbands are fully explored to refine the features. After the refinement of the subbands, the high-frequency components are directly taken as the inputs of the reconstruction network, while the low-frequency components are transformed by the guidance generation network to accomplish the spatial-spectral collaboration and also make preparations for the spectral adjustment. To explore the spectral correlation collaboration, a novel graph convolutional network is designed for the modulation of intraspectral relationships. Finally, the adjusted MS features are combined with the high-frequency components of PAN features to reconstruct the high-resolution MS image. Experimental results show that the proposed method outperforms traditional state-of-the-art pan-sharpening methods as well as the available deep learning-based ones.
Yinghui Xing, Shuyuan Yang 0001, Zhixi Feng, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2022 ADHR-CDNet: Attentive Differential High-Resolution Change Detection Network for Remote Sensing Images
abstract
With the development of deep learning, change detection technology has gained great progress. However, how to effectively extract multi-scale substantive changed features and accurately detect small changed objects as well as the accurate details is still a challenge. To solve the problem, we propose Attentived Differential High-Resolution Change Detection Network (ADHR-CDNet) for remote sensing images. In ADHR-CDNet, a novel high-resolution backbone with a Differential Pyramid Module (DPM) is proposed to extract multi-level and multi-scale substantive changed features. The backbone structure with four interconnected sub-network branches of different resolution is helpful to extract multi-level and multi-scale features. DPM is capable of distinguishing between substantive changes and pseudo changes induced by illumination, shadow, seasonal variation, and so on. Then, a novel Multi-Scale Spatial feature Attention Module (MSSAM) is presented to effectively fuse the spatial detail information of different scale features produced by our backbone to generate finer prediction. We conduct quantitative and qualitative experiments on three public change detection datasets: the Lebedev, the LEVIR-CD, and the WHU Building dataset. The proposed ADHR-CDNet reaches F1-score of 97.2% (improved 3.1%) on the Lebedev dataset, 91.4% (improved 1.6%) on the LEVIR-CD dataset, and 90.9% (improved 1.2%) on the WHU Building dataset. The experimental results demonstrate that our method performs much better than the state-of-the-art methods. The visualization comparison results show that our method can effectively detect small changed objects and significantly improve the details of detected changed objects. Our code is available at https://github.com/w-here/ASGO-113lab/tree/main/ADHR-CDNet.
Xiuwei Zhang 0001, Mu Tian, Yinghui Xing, Yuanzeng Yue, Hanlin Yin, Runliang Xia, Yanning Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Learning Spectral Cues for Multispectral and Panchromatic Image Fusion
abstract
Recently, deep learning based multispectral (MS) and panchromatic (PAN) image fusion methods have been proposed, which extracted features automatically and hierarchically by a series of non-linear transformations to model the complicated imaging discrepancy. But they always pay more attention to the extraction and compensation of spatial details and use the mean squared error or mean absolute error as a loss function, regardless of the preservation of spectral information contained in multispectral images. For the sake of the improvements in both spatial and spectral resolution, this paper presents a novel fusion model that takes the spectral preservation into consideration, and learns the spectral cues from the process of generating a spectrally refined multispectral image, which is constrained by a spectral loss between the generated image and the reference image. Then these spectral cues are used to modulate the PAN features to obtain final fusion result. Experimental results on reduced-resolution and full-resolution datasets demonstrate that the proposed method can obtain a better fusion result in terms of visual inspection and evaluation indices when compared with current state-of-the-art methods.
Yinghui Xing, Shuyuan Yang 0001, Yan Zhang 0127, Yanning Zhang 0001
IEEE Trans. Image Process.1
2018 Pansharpening With Multiscale Geometric Support Tensor Machine
abstract
In this paper, a new pansharpening method is proposed by constructing a set of multiscale geometric support tensor filters (MGSTFs). First, a least-square ridgelet support tensor machine is developed to derive a series of MGSTFs. Then the source images are formulated as tensors and filtered by MGSTFs to capture geometric and salient features of images. These features are then fused at each scale and direction to obtain the fused products. The distortions can be reduced by exploring the tensor formulation of multispectral data and endowing the filters’ directionality to capture the geometric details of images. Some experiments are carried out on several groups of QuickBird and GeoEye-1 images, and the results show that our proposed method can simultaneously reduce spectral distortions and preserve spatial details in the fused image.
Yinghui Xing, Min Wang 0007, Shuyuan Yang 0001, Kai Zhang 0010
IEEE Trans. Geosci. Remote. Sens.1