VLDB 2026 Research / reviewers in the wild / expert
Weiying Xie
dblp:150/3937
· DBLP profile ↗
132ranked-venue papers
29as first author
104since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 55 · 13 first-author · 39 since 2021Artificial intelligence and machine learning · 47 · 14 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 4 first-author · 38 since 2021Systems, architecture and hardware · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TOP-RL: Task-Optimized Progressive Token Pruning with Reinforcement Learning for Vision Language ModelsabstractIn recent years, Large Vision-Language Models (LVLMs) have significantly advanced multimodal tasks. However, their inference requires intensive processing of numerous visual tokens and incurs substantial computational overhead. Existing methods typically compress visual tokens either at the input stage or in early model layers, ignoring variations across tasks and depths. To address these limitations, we introduce TOP-RL, a Task-Optimized Progressive token pruning framework based on Reinforcement Learning. TOP-RL formulates visual token pruning as a multi-stage Markov Decision Process (MDP). It employs an agent trained with dense and fine-grained reward signals to progressively generate differentiable binary masks. This enables TOP-RL to adaptively select crucial visual tokens tailored to each task, effectively balancing accuracy and computational efficiency. Extensive experiments on leading multimodal datasets and advanced LVLMs validate that TOP-RL effectively learns task-optimized pruning policies, significantly boosting inference efficiency while preserving robust performance. For instance, LLaVA-NeXT equipped with TOP-RL achieves a 1.9x speedup in inference time and a 9.3x reduction in FLOPs, with 96% performance preserved. Hengyi Wang, Weiying Xie, Yaotao Wei, Kai Jiang 0001, Mingxiang Cao, Chenhe Hao, Leyuan Fang |
AAAI | 2 |
| 2026 | MMFormer: Multi-Modality semi-Supervised vision transformer in remote sensing imagery classification
Daixun Li, Weiying Xie, Leyuan Fang, Yunke Wang, Mingxiang Cao, Jitao Ma, Yunsong Li 0001, Chang Xu 0002 |
Neural Networks | 2 |
| 2026 | FA-Mamba: frequency attention driven Mamba for multimodal remote sensing classification
Danian Yang, Daixun Li, Jitao Ma, Yibing Lu, Yunsong Li 0001, Leyuan Fang, Weiying Xie |
Neural Networks | 7 |
| 2026 | Generating Any Changes in the Noise DomainabstractChange detection is essential in Earth observation, yet current models heavily rely on large-scale annotated datasets. Generative models offer a promising alternative by synthesizing training data, but generating temporally coherent image pairs with realistic, semantically meaningful changes remains a significant challenge. Existing approaches typically simulate changes by generating pre- and post-change label maps using either heuristic rules (e.g., copy-pasting) or text prompts. However, the former offers limited change diversity, while the latter often fails to maintain spatial consistency between image pairs. We observe that the noise space of diffusion models encodes strong generative capacity and spatial controllability: localized perturbations in the noise can yield meaningful, interpretable changes in corresponding image regions. Motivated by this, we propose Noise2Change, a framework for simulating change directly in the noise domain. The key idea is to manipulate the semantic composition of the initial noise sampled from the noise domain, such that the diffusion process generates structurally consistent pre- and post-change images reflecting realistic transformations. Since the unperturbed noise is shared between both images, the resulting pairs exhibit strong temporal alignment and semantic coherence, effectively addressing the trade-off between realism and consistency. Concretely, we employ a discrete diffusion model to extract high-level semantics from the initial noise. Guided by these semantics, we introduce a change simulation strategy that optimizes the noise to encode intended changes. The modified noise is then used to drive the diffusion process, yielding pre- and post-change label maps with natural structural transitions. These maps are passed through a unified framework for image generation and label refinement, producing highly aligned image-label pairs. Our framework supports diverse change types across a wide range of scenarios. Extensive experiments on multiple change detection tasks demonstrate that our method achieves superior performance compared to existing generative approaches. Jun Yue 0004, Pedram Ghamisi, Weiying Xie, Leyuan Fang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Domain adapter for visual object tracking based on hyperspectral video
Langkun Chen, Gang He 0002, Weiying Xie, Yunsong Li 0001 |
Pattern Recognit. | 6 |
| 2026 | Glob-Diffusion: A Global Consistent Diffusion Model for Large-Scale Image GenerationabstractLarge-scale images play a crucial role in geospatial surveying, as they cover an extensively broad view and diverse objects. Due to computational limitations, existing methods rely on generating large-scale images in patches. However, the lack of global guidance in these methods often leads to significant logical errors among different patches. To address this issue, we propose a Global Consistency Diffusion model (Glob-Diffusion) for large-scale image generation. The core idea is to utilize the global consistency of small-scale images to guide the generation of large-scale images. Specifically, we introduce a Hierarchical Distributed Guidance (HDG) module that extracts patch prompts with different semantic hierarchies from small-scale images, distributedly embedding them into the generation of large-scale images to maintain global consistency across various regions. In addition, we further design a Region Guided Adapter (RGA) that dynamically optimizes the guidance strength of patch prompts by comparing differences across generated regions, effectively improving the realism of large-scale images. Our method demonstrates remarkable visual synthesis results across various natural scenes, effectively preserving global consistency in large-scale images, and also significantly enhancing the generation quality of large-scale remote sensing images. Code will be available at https://github.com/kyh433/Glob-Diffusion. Yuhan Kang, Hengcan Shi, Hao Liu 0123, Weiying Xie, Leyuan Fang, Lorenzo Bruzzone |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Cross-Modal Visual Perception Consistency: A Language-Enhanced Approach for Heterogeneous Change DetectionabstractHeterogeneous remote sensing image change detection (HRSICD) seeks to identify surface changes by comparing images captured at different times. However, CD faces significant challenges due to heterogeneity arising from varying sensor types and imaging conditions. Recently, powerful vision-language models like CLIP have emerged, with strong semantic decoding abilities. Opening new possibilities for using linguistic information as an auxiliary in visual tasks, potentially driving breakthroughs in HCD. Capitalizing on this prospect, we investigate graph learning with vision-language features and introduce LEVPC, the first language-enhanced visual perception consistency framework for HCD. First, we create a mutual information-guided graph aggregation module. Specifically, it builds modality-invariant structured relationships among visual nodes by using language features as connecting bridges, providing a consistent foundation for comparing changes. To reduce modeling bias from heterogeneity, language is used as an anchor to aggregate features, ensuring a unified expression of visual representations. In summary, language guides the generation and aggregation of multiple subgraphs from visual inputs, ultimately building robust representations of structural relationships within a shared semantic space. Moreover, a change semantic compensation module is introduced, which analyses the change intensity between bi-temporal data from a vision-language perspective. And then adds change-related semantic descriptions for salient change regions, enhancing the expressiveness of visual change features. Experiments on multiple datasets validate the superior performance of LEVPC in HCD, achieving an average increase of 2.6% in Kappa. The code will be publicly available at https://github.com/sylXIDIAN/LEVPC. Siyao Li, Weiying Xie, Jitao Ma, Leyuan Fang, Yunsong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | MultiGS: Multi-Dimensional Information-Aware Gradient SparsificationabstractGradient sparsification (GS) is an effective method for reducing communication overhead in distributed training. For the first time, we introduce the concept of Multi-dimensional information into GS and propose a new gradient sparsification method named Multi-dimensional information-aware Gradient sparsification (MultiGS), which achieves high compression ratio with negligible accuracy loss and is applicable to mainstream network architectures. MultiGS reconstructs the layer-wise gradient by combining the high-frequency components of the local gradient and the low-frequency components of the sparsified global gradient that effectively addresses the issue of stale gradients and alleviates model bifurcation. Through the convergence proof of MultiGS for smooth non-convex problems and comparison with momentum SGD in convergence speed, we show that such new perspective approach is theoretically reasonable and practically effective. As validated with several mainstream model families (i.e., ResNets, VGGNet, LSTM, Vision Transformer, and Large Language Models), our MultiGS shows better accuracy over previous GS methods. Moreover, empirical results show that when a sufficient number of training nodes are available, MultiGS accelerates the distributed training by more than 3×, which is better than existing sparsification method. Jitao Ma, Donglai Liu, Weiying Xie, Yunsong Li 0001, Leyuan Fang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | BSDM: Background Suppression Diffusion Model for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) is widely used in Earth observation and deep space exploration. A major challenge for HAD is the complex background of the input hyperspectral images (HSIs), resulting in anomalies confused in the background. On the other hand, most existing HAD methods require training a separate model for each HSI, resulting in poor generalization in practical applications. This paper starts the first attempt to study a new and generalizable background learning problem without labeled samples. We present a novel solution BSDM (background suppression diffusion model) for HAD, which can simultaneously learn latent background distributions and generalize to different datasets for suppressing complex background. It is featured in three aspects: (1) For the complex background of HSIs, we design pseudo-background noise and learn the potential background distribution in it with a diffusion model (DM). (2) For the generalizability problem, we apply a statistical offset module so that the BSDM adapts to datasets of different domains without labeling samples. (3) For achieving background suppression, we innovatively improve the inference process of DM by feeding the original HSIs into the denoising network, which removes the background as noise. Our work paves a new background suppression way for HAD that can improve HAD performance without the prerequisite of manually labeled data. Assessments and generalization experiments of four HAD methods on several real HSI datasets demonstrate the above three unique properties of the proposed method. Our project is available at https://github.com/majitao-xd/BSDM-HAD. Jitao Ma, Weiying Xie, Xueshuang Xiang, Yunsong Li 0001, Leyuan Fang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | LoME: LoRA-Driven Multimodal Extractor for RGB-X Vision TasksabstractRGB-X multimodal vision tasks present a highly promising approach to enhancing model performance in complex visual conditions. Existing multimodal frameworks are based on either the symmetric parallel network of feature fusion or the shared network of input fusion. However, parallel networks suffer from uncontrollable parameters and imbalanced optimization across modal branches, while shared networks often lead to a lack of diversity in gradient optimization. To address these challenges, we propose the LoRA-driven Multimodal Extractor (LoME), following a comprehensive analysis of existing multimodal frameworks. The low-rank properties of modal adapters for LoME ensure controllable growth in model parameters as the number of modalities increases. The dynamic parameter fusion between adapters and the shared feature extractor decouples gradient optimization directions, effectively mitigating imbalances caused by multimodal data biases while preserving complementary features. Moreover, we employ a training strategy based on dynamic rank allocation to reduce computational overhead and enhance modal diversity expression. We validate the effectiveness and generalizability of LoME across three multimodal vision tasks. LoME achieves superior performance compared to previous state-of-the-art methods on multiple datasets. For example, on the DroneVehicle dataset, our method achieves a 10.4% improvement in accuracy compared to the SOTA method, while the parameter overhead is reduced to 23% of the previous network (44.63M). The code has been open-sourced at https://github.com/zyszxhy/LoME. Weiying Xie, Tianlin Hui, Daixun Li, Jie Lei 0001, Yunsong Li 0001, Leyuan Fang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Aligning and Prompting Anything for Zero-Shot Generalized Anomaly DetectionabstractZero-shot generalized anomaly detection (ZGAD) plays a critical role in industrial automation and health screening. Recent studies have shown that ZGAD methods built on visual-language models (VLMs) like CLIP have excellent cross-domain detection performance. Different from other computer vision tasks, ZGAD needs to jointly optimize both image-level anomaly classification and pixel-level anomaly segmentation tasks for determining whether an image contains anomalies and detecting anomalous parts of an image, respectively, this leads to different granularity of the tasks. However, existing methods ignore this problem, processing these two tasks with one set of broad text prompts used to describe the whole image. This limits CLIP to align textual features with pixel-level visual features and impairs anomaly segmentation performance. Therefore, for precise visual-text alignment, in this paper we propose a novel fine-grained text prompts generation strategy. We then apply the broad text prompts and the generated fine-grained text prompts for visual-textual alignment in classification and segmentation tasks, respectively, accurately capturing normal and anomalous instances in images. We also introduce the Text Prompt Shunt (TPS) model, which performs joint learning by reconstruction the complementary and dependency relationships between the two tasks to enhance anomaly detection performance. This enables our method to focus on fine-grained segmentation of anomalous targets while ensuring accurate anomaly classification, and achieve pixel-level comprehensible CLIP for the first time in the ZGAD task. Extensive experiments on 13 real-world anomaly detection datasets demonstrate that TPS achieves superior ZGAD performance across highly diverse datasets from industrial and medical domains. Jitao Ma, Weiying Xie, Hangyu Ye, Daixun Li, Leyuan Fang |
AAAI | 2 |
| 2025 | AdaGK-SGD: Adaptive Global Knowledge Guided Distributed Stochastic Gradient DescentabstractDistributed machine learning (DML) is promising for training large models on large datasets. In DML, multiple workers collaborate on the training of neural networks, significantly reducing the time required for neural network training. The efficiency of DML is heavily influenced by communication, making it crucial to balance the trade-off between communication cost and model performance in current research. Local methods are excellent at reducing communication costs, yet face degradation in accuracy and generalizability. Indeed, global knowledge is valuable for improving performance in local methods. However, the theoretical analysis of global knowledge validity is lacking, and global knowledge can currently only be used in the global aggregation of local methods due to communication limitations and staleness. To this end, in this paper, we establish the mechanism of global knowledge guidance and propose Adaptive Global Knowledge Guided Distributed Stochastic Gradient Descent (AdaGK-SGD) to extend the guidance of global knowledge to the whole distributed training process without any additional communication. Specifically, we define the maximum lifetime of global knowledge based on the mechanism, and establish a correlation between the maximum lifetime and the validity of global knowledge to circumvent the adverse effects of global knowledge staleness. The Maximum Lifetime of Global Knowledge module of our algorithm can be applied separately to other algorithms. In addition, considering the application, we provide a straightforward and efficient strategy for achieving the maximum lifetime adaptive setting. We establish the convergence rate of AdaGK-SGD for convex and non-convex scenarios. Numerically, we find that AdaGK-SGD can significantly improve the accuracy and generalizability of distributed algorithms compared with existing methods. Hangyu Ye, Weiying Xie, Yunsong Li 0001, Leyuan Fang |
AAAI | 2 |
| 2025 | FedCS: Coreset Selection for Federated LearningabstractFederated Learning (FL) is an emerging direction in distributed machine learning that enables jointly training a model without sharing the data. However, as the size of datasets grows exponentially, computational costs of FL increase. In this paper, we propose the first Coreset Selection criterion for Federated Learning (FedCS) by exploring the Distance Contrast (DC) in feature space. Our FedCS is inspired by the discovery that DC can indicate the intrinsic properties inherent to samples regardless of the networks. Based on the observation, we develop a method that is mathematically formulated to prune samples with high DC. The principle behind our pruning is that high DC samples either contain less information or represent rare extreme cases, thus removal of them can enhance the aggregation performance. Besides, we experimentally show that samples with low DC usually contain substantial information and reflect the common features of samples within their classes, such that they are suitable for constructing coreset. With only two time of linear-logarithmic complexity operation, FedCS leads to significant improvements over the methods using whole dataset in terms of computational costs, with similar accuracies. For example, on the CIFAR-10 dataset with Dirichlet coefficient α = 0.1, FedCS achieves 58.88% accuracy using only 44% of the entire dataset, whereas other methods require twice the data volume as FedCS for same performance. Chenhe Hao, Weiying Xie, Daixun Li, Hangyu Ye, Leyuan Fang, Yunsong Li 0001 |
CVPR | 2 |
| 2025 | Towards Long-Horizon Vision-Language-Action System: Reasoning, Acting and Memory
Daixun Li, Mingxiang Cao, Donglai Liu, Weiying Xie, Tianlin Hui, Lunkai Lin, Yunsong Li 0001 |
ICCV | 5 |
| 2025 | Allowing Oscillation Quantization: Overcoming Solution Space Limitation in Low Bit-Width Quantization
Weiying Xie, Zihan Meng, Jitao Ma, Wenjin Guo, Leyuan Fang, Yunsong Li 0001 |
ICCV | 1 |
| 2025 | FusionSAM: Visual Multi-Modal Learning with Segment Anything ModelabstractMultimodal image fusion and semantic segmentation are critical for autonomous driving. Despite advancements, current models often struggle with segmenting densely packed elements due to a lack of comprehensive fusion features for guidance during training. While the Segment Anything Model (SAM) allows precise control during fine-tuning through its flexible prompting encoder, its potential remains largely unexplored in the context of multimodal segmentation for natural images. In this paper, we introduce SAM into multimodal image segmentation for the first time, proposing a novel framework that combines Latent Space Token Generation (LSTG) and Fusion Mask Prompting (FMP) modules. This approach transforms the training methodology for multimodal segmentation from a traditional black-box approach to a controllable, prompt-based mechanism. Specifically, we obtain latent space features for both modalities through vector quantization and embed them into a cross-attention-based inter-domain fusion module to establish long-range dependencies between modalities. We then use these comprehensive fusion features as prompts to guide precise pixel-level segmentation. Extensive experiments on multiple public datasets demonstrate that our method significantly outperforms SAM and SAM2 in multimodal autonomous driving scenarios, achieving an average improvement of 4.1% over the state-of-the-art method in segmentation mIoU, and the performance is also optimized in other multi-modal visual scenes. Daixun Li, Weiying Xie, Mingxiang Cao, Yunke Wang, Leyuan Fang, Yunsong Li 0001, Chang Xu 0002 |
KDD (2) | 2 |
| 2025 | Uni-Sight: An E2E Vision-Language-Action System Unifying Multi-View Alignment and Multi-Modal FusionabstractVision-Language-Action (VLA) systems are crucial for autonomous decision-making in embodied intelligence. While current systems have advanced the instruction-following capabilities, their limited spatial perception often leads to suboptimal performance for mobile manipulation tasks in unstructured environments. To address this challenge, we propose Uni-Sight, an end-to-end VLA system for robust mobile manipulation. Uni-Sight unifies decision-making, perception, and control through joint training, enabling synchronized cross-component optimization. Within the system, we introduce Latent Feature Aligner (LFA) that ensures accurate target localization by aligning multi-view data. Specifically, we develop Domain Transfer Policy (DTP), a hierarchical policy constrained by LiDAR-guided spatial priors, which ensures 3D spatial understanding with limited visual coverage. Extensive experiments on 20 real-world mobile manipulation tasks demonstrate the high task success rate and robust execution performance of Uni-Sight. Our Uni-Sight achieves a 3.04× the success rate of existing methods, and exhibits superior generalization in both long-horizon and zero-shot scenes. Code and dataset are publicly available at https://github.com/trantor2nd/Uni-Sight. Daixun Li, Sibo He, Jiayun Tian, Weiying Xie, Mingxiang Cao, Donglai Liu, Tianlin Hui, Yunsong Li 0001 |
ACM Multimedia | 5 |
| 2025 | TF-ATM: Training-Free Adaptive Token Merging
Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Tianlin Hui, Jitao Ma, Leyuan Fang |
ACM Multimedia | 2 |
| 2025 | Hyperspectral anomaly detection with self-supervised anomaly prior
Yidan Liu, Kai Jiang 0001, Weiying Xie, Yunsong Li 0001, Leyuan Fang |
Neural Networks | 3 |
| 2025 | Exploring hyperspectral anomaly detection with human vision: A small target aware detector
Jitao Ma, Weiying Xie, Yunsong Li 0001 |
Neural Networks | 2 |
| 2025 | M³amba: CLIP-Driven Mamba Model for Multi-Modal Remote Sensing ClassificationabstractMulti-modal fusion holds great promise for integrating information from different modalities. However, due to a lack of consideration for modal consistency, existing multi-modal fusion methods in the field of remote sensing still face challenges of incomplete semantic information and low computational efficiency in their fusion designs. Inspired by the observation that the visual language pre-training model CLIP can effectively extract strong semantic information from visual features, we propose M3amba, a novel end-to-end CLIP-driven Mamba model for multi-modal fusion to address these challenges. Specifically, we introduce CLIP-driven modality-specific adapters in the fusion architecture to avoid the bias of understanding specific domains caused by direct inference, making the original CLIP encoder modality-specific perception. This unified framework enables minimal training to achieve a comprehensive semantic understanding of different modalities, thereby guiding cross-modal feature fusion. To further enhance the consistent association between modality mappings, a multi-modal Mamba fusion architecture with linear complexity and a cross-attention module Cross-SS2D are designed, which fully considers effective and efficient information interaction to achieve complete fusion. Extensive experiments have shown that M3amba has an average performance improvement of at least 5.98% compared with the state-of-the-art methods in multi-modal hyperspectral image classification tasks in the remote sensing field, while also demonstrating excellent training efficiency, achieving a double improvement in accuracy and efficiency. The code is released athttps://github.com/kaka-Cao/M3amba. Mingxiang Cao, Weiying Xie, Xin Zhang 0092, Kai Jiang 0001, Jie Lei 0001, Yunsong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | SeaDATE: Remedy Dual-Attention Transformer With Semantic Alignment via Contrast Learning for Multimodal Object DetectionabstractMultimodal object detection leverages diverse modal information to enhance the accuracy and robustness of detectors. Due to its ability to capture long-range dependencies, the Transformer model provides a powerful mechanism for integrating multimodal features during feature extraction. This capability significantly enhances the accuracy of multimodal object detection by addressing the limitations of local feature extraction inherent in traditional methods. However, current methods merely stack Transformer-guided fusion techniques without exploring their capability to extract features at various depth layers of network, thus limiting the improvements in detection performance. In this paper, we introduce an accurate and efficient multimodal object detection method named SeaDATE. Initially, we propose a novel dual attention Feature Fusion (DTF) module that, under Transformer’s guidance, integrates local and global information through a dual attention mechanism, strengthening the fusion of modal features from orthogonal perspectives using spatial and channel tokens. Meanwhile, our theoretical analysis and empirical validation demonstrate that the Transformer-guided fusion method, treating images as sequences of pixels for fusion, performs better on shallow features’ detail information compared to deep semantic information. To address this, we designed a contrastive learning (CL) module aimed at learning features of multimodal samples, remedying the shortcomings of Transformer-guided fusion in extracting deep semantic features, and effectively utilizing cross-modal information. Extensive experiments and ablation studies on the FLIR, LLVIP, and M3FD datasets have proven our method to be effective, achieving state-of-the-art detection performance. Shuhan Dong, Weiying Xie, Danian Yang, Yunsong Li 0001, Jiayuan Tian, Jie Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | ShiftQuant: Toward Accurate and Efficient Sub-8-bit Integer TrainingabstractNeural network training is a memory- and compute-intensive task. Quantization, which enables low-bitwidth formats in training, can significantly mitigate the workload. To reduce quantization error, recent methods have developed new data formats and additional pre-processing operations on quantizers. However, it remains quite challenging to achieve high accuracy and efficiency simultaneously. In this paper, we explore sub-8-bit integer training from its essence of gradient descent optimization. Our integer training framework includes two components: ShiftQuant to realize accurate gradient estimation, and L1 normalization to smoothen the loss landscape. ShiftQuant attains performance that approaches the theoretical upper bound of group quantization. Furthermore, it liberates group quantization from inefficient memory rearrangement. The L1 normalization facilitates the implementation of fully quantized normalization layers with impressive convergence accuracy. Our method frees sub-8-bit integer training from pre-processing and supports general devices. This framework achieves negligible accuracy loss across various neural networks and tasks (0.92% on 4-bit ResNets, 0.61% on 6-bit Transformers). The prototypical implementation of ShiftQuant achieves more than 1.85×/15.3% performance improvement on CPU/GPU compared to its FP16 counterparts, and 33.9% resource consumption reduction on FPGA than the FP16 counterparts. The proposed fully-quantized L1 normalization layers achieve more than 35.54% improvement in throughout on CPU compared to traditional L2 normalization layers. Moreover, theoretical analysis verifies the advancement of our method. Wenjin Guo, Donglai Liu, Weiying Xie, Yunsong Li 0001, Xuefei Ning, Zihan Meng, Shulin Zeng, Jie Lei 0001, Zhenman Fang, Yu Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Hyperspectral Object Tracking With Spectral Information PromptabstractHyperspectral videos contain a larger number of spectral bands, providing extensive spectral information and material identification capabilities. This advantage confers hyperspectral trackers to achieve superior performance in challenging tracking scenarios. However, the limited availability of hyperspectral training data and the inability of existing algorithms to fully exploit hyperspectral information restrict the tracking performance. To address this issue, a novel framework, Spectral Prompt-based Hyperspectral Object Tracking (SP-HST), is proposed. SP-HST leverages a RGB tracking network as the main branch for feature extraction and tracking, which accounts for more than 98% of the total parameters and remains frozen during the training procedure. Additionally, the Spectral Prompt Learning (SPL) branch, comprising multiple lightweight prompt blocks, is introduced to generate complementary spectral representations as the prompt. The prompts contain abundant spectral information from hyperspectral data, enhancing the discriminative ability of features within the main branch. Furthermore, the Complementary Weight Learning (CWL) is employed to calculate the importance of spectral information from different prompts, enabling the features for hyperspectral object tracking to contain more spectral information that is absent in the feature of the main branch. By utilizing the spectral information as prompt, the number of trainable parameters is less than 2% of that in the tracking network, and the convergence is reached in 12 training epoch. Extensive experiments demonstrate the superiority of SP-HST, achieving a new state-of-the-art tracking performance, 71.3% of the AUC score on the HOTC dataset and 96.7% of the DP@20P score on the IMEC25 dataset. The code will be released at https://github.com/lgao001/SP-HST. Gang He 0002, Langkun Chen, Weiying Xie, Yunsong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Semi-Mamba: Mamba-Driven Semi-Supervised Multimodal Remote Sensing Feature ClassificationabstractMamba architecture achieves the same performance as attention mechanisms with linear complexity, leading to significant progress in remote sensing land cover classification. However, existing Mamba methods rarely leverage the representational complementarity and consistency between different modalities, resulting in challenges such as incomplete fusion. To address these issues, we propose Semi-Mamba, a novel semi-supervised framework specifically designed for high-dimensional multi-modal data fusion. We introduce the Mamba Cross-Modality Fusion Module, which enables cross-modal learning of temporal features through state-space model interactions and smooth integration of input matrices, enhancing the fusion of richer feature representations. Additionally, to tackle the inherent difficulty of acquiring pixel-level annotations in remote sensing datasets, we introduce a multi-modal semi-supervised mechanism. This mechanism utilizes cross-modal supervision between different modalities to maximize data utilization and improve learning efficiency. It effectively enables joint training on both labeled and unlabeled data without relying on pseudo-labels. We integrate these innovations into a unified end-to-end framework. Compared to state-of-the-art CNN and Transformerbased architectures, our framework shows a significant improvement of over 3.12%, setting a new benchmark for semi-supervised multi-modal data fusion. The code has open sourced at https://github.com/LDXDU/Semi_Mamba_RS. Yunsong Li 0001, Daixun Li, Weiying Xie, Jitao Ma, Sibo He, Leyuan Fang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Dual-Depth Unified Joint Optimization: Adaptive Curvature-Based CompressionabstractModel compression methods such as pruning and quantization have been proposed to facilitate the deployment of convolutional neural networks (CNNs) on resource-constrained devices. Existing methods aim to combine the two for simultaneous improvement in compression ratio and runtime efficiency. However, most of the joint methods adopt linear tandem structures. Due to the lack of a unified framework, different optimization directions result in suboptimal solutions, especially when the compression ratio is extremely high. In this paper, we propose a novel adaptive curvature-based compression (ACC) method, which achieves a dual-depth unified joint optimization of pruning and quantization. In the first depth, we unify the pruning and quantization criteria using mean curvature, which leverages the discrete nature of image data and the continuum theory of differential geometry. In the second depth, we replace the traditional training process in the joint pruning-quantization method with curvature-aware knowledge distillation (CKD), unifying the two-stage approach into a simple but powerful parallel step. Our method is effective and interpretable by utilizing inherent properties to promote the understanding of information distribution and the importance of feature maps. Extensive experiments on multiple advanced benchmarks and diverse downstream task datasets have validated the superiority and generalizability of our ACC. Notably, we can achieve a 1.05% Top-1 accuracy improvement over the baseline under an extreme compression ratio of 454.55×, outperforming existing state-of-the-art (SOTA) methods. Yunsong Li 0001, Xin Zhang 0092, Weiying Xie, Daixun Li, Hangyu Ye, Leyuan Fang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Visual State Space Model With Graph-Based Feature Aggregation for No-Reference Image Quality AssessmentabstractInspired by the human visual system (HVS), no-reference image quality assessment (NR-IQA) has made significant progress without relying on perfect reference images. The HVS is primarily influenced by the combined effects of representational information with different receptive fields and attribute categories when capturing subjective perceived quality. However, existing methods only roughly or partially utilize representations of multi-dimensional information. Furthermore, current NR-IQA methods either rely on convolutional neural networks (CNNs) with limited local perception or depend on the computational complexity of vision transformers (ViTs). To make up for the shortcomings of these two architectures, an emerging visual state space model (VMamba) is introduced. Motivated by this, this paper presents a NR-IQA method via VIsual State space model with Graph-based feature Aggregation (VISGA). Specifically, we utilize a plain, pre-training-free, and feature-enhanced VMamba as the backbone. To align with the perceptual mechanisms of the HVS by effectively using features with different dimensional information, a graph convolutional network-based multi-receptive field and multi-level aggregation module is designed to deeply explore the correlations and interactions of multi-dimensional representations. Additionally, we propose a gated local enhancement module with patch-wise perception to enhance the local perception of VMamba. Extensive experiments conducted on seven databases demonstrate that VISGA achieves outstanding performance. Notably, our model remains state-of-the-art when training with very few parameters. The code is released athttps://github.com/xirihao/VISGA. Haozhi Shi, Weiying Xie, Yunsong Li 0001, Leyuan Fang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | MIFNet: Multi-Scale Interaction Fusion Network for Remote Sensing Image Change DetectionabstractChange Detection (CD) is a crucial and challenging task in remote sensing observations. Despite the remarkable progress driven by deep learning in remote sensing change detection, several challenges remain regarding global information representation and efficient interaction. The traditional Siamese network structure, which extracts features from bitemporal images using a weight-sharing network and generates a change map, but often neglects phase interaction information between images. Additionally, multi-scale feature fusion methods frequently use FPN-like structures, leading to lossy cross-layer information transmission and hindering the effective utilization of features. To address these issues, we propose a multi-scale interaction fusion network (MIFNet) that fuses bitemporal features at an early stage, using deep supervision techniques to guide early fusion features in obtaining abundant semantic representation of changes, also we construct a dual complementary attention module (DCA) to capture temporal information. Furthermore, we introduce a collection-allocation fusion mechanism, which is different from previous layer-by-layer fusion methods since it collects global information and embeds features at different levels to achieve effective cross-layer information transmission and promote global semantic feature representation. Extensive experiments demonstrate that our method achieves competitive results on the LEVIR-CD+ dataset, outperforming other advanced methods on both the LEVIR-CD and SYSU-CD datasets, with F1 improved by 0.96% and 0.61%, respectively, compared to the most advanced models. Weiying Xie, Wenjie Shao, Daixun Li, Yunsong Li 0001, Leyuan Fang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Hyperspectral Target Detection Based on Generative Self-Supervised Learning With Wavelet TransformabstractRecently, generative self-supervised learning (GSSL) has gained extensive attention in hyperspectral remote sensing. For the hyperspectral target detection (HTD) task, traditional GSSL-based algorithms usually require hyperspectral images (HSIs) as additional datasets for pretraining, which are relatively resource-intensive and time-consuming. To better interpret the spectral-spatial information of HSIs while alleviating the dependence on large-scale hyperspectral datasets, we develop a novel two-stage framework for HTD based on GSSL in this article. In the preprocessing for the input HSI, a dimensional transformation (DT) module and a coarse detection reference (CDR) module are constructed to produce feature patches as training samples for subsequent pretraining and fine-tuning. In the pretraining stage for spectral-spatial reconstruction, we construct an asymmetric autoencoder (AE) architecture which leverages the transformer blocks with long-range perception to extract generalized features and explore discriminative feature representations of the input HSI. Specifically, a dual-stream wavelet patch embedding (DWPE) module is proposed to integrate the wavelet transform (WT) mechanism with the convolutional neural networks (CNNs), which extracts robust spectral-spatial features by performing convolutional operations with different frequency components of WT. In the fine-tuning stage, a novel signature-constrained cross-entropy (SC-CE) loss function is proposed to constrain the network optimization. For the final detection, a pixel-level fusion based on coarse detection based pixel-level fusion (CDPF) module is employed after inference to further suppress the interference from background. Experimental results on six real HSIs demonstrate that the proposed method achieves superior detection performance while maintaining the generalization of the pretrained model. Shuai Wang 0057, Yunsong Li 0001, Weiying Xie, Kai Jiang 0001, Kailang Cao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | A Signature-Constrained Two-Stage Framework for Hyperspectral Target Detection Based on Generative Self-Supervised LearningabstractRecently, hyperspectral target detection (HTD) technique based on deep learning (DL) has been developed rapidly. However, existing algorithms show poor generalization across different hyperspectral images (HSIs), where repeated training and inference based on limited prior information are necessary. To liberate HTD from dependence on the quantity and the quality of training samples, this article proposes a signature-constrained two-stage framework for HTD (HTD-STF) based on generative self-supervised learning (GSSL). In the first stage for pre-training, to realize spectral-spatial reconstruction, we build an asymmetric autoencoder (AE) employing transformer blocks with long-range perception for generalized feature extraction. During pre-training, the spectral-spatial similarity loss is designed to improve the effect of reconstruction. In the second stage for fine-tuning and detection, the signature is utilized in preprocessing, training and inference, respectively. Specifically, the coarse sample mining and tiling strategy in preprocessing not only facilitates the framework in flexible input dimension, but also provides pseudo labels for end-to-end training. During training, we adopt the signature as guidance for feature-level fusion, which alleviates the impact of sample imbalance. After training, the final inference based on pixel-level fusion refines the original output. For ideal GSSL, the HyperMix-10K, a new large-scale hyperspectral dataset, has been constructed in this work, which contains numerous unlabeled HSIs captured in various scenes. Experimental results and analysis on real HSIs verify the effectiveness and generalization ability of the HTD-STF. Shuai Wang 0057, Yunsong Li 0001, Weiying Xie, Kai Jiang 0001, Kailang Cao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | MROD-YOLO: Multimodal Joint Representation for Small Object Detection in Remote Sensing Imagery via Multiscale Iterative Aggregation
Ruitao Lu, Dingwen Zhang, Weiying Xie, Shuang Su, Zhenyu Zhang 0028 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | RS-IML: Federated Intrinsic Mask Learning on Remote Sensing Image InterpretationabstractInterpreting remote sensing (RS) images plays a crucial role in numerous applications such as environmental monitoring, urban planning, agricultural management and disaster assessment. Nevertheless, remote sensing data is frequently dispersed among various organizations. Privacy concerns and data-sharing limitations make it difficult to utilize large-scale datasets within a centralized training framework. Federated learning (FL) provides a promising approach by facilitating collaborative model training across decentralized data sources, eliminating the need for data centralization. However, the application of FL in RS scenarios is challenging due to the resource-limited edge nodes cannot meet the high demand for computation and memory resources to train deep learning models. Neural network lightweighting techniques have the potential to enhance model efficiency, but existing methods still present significant challenges, such as reliance on initial training of dense models during lightweighting and potential performance degradation after lightweighting. To address these challenges, we propose RS-IML, a novel FL framework for RS Image Interpretation based on Intrinsic Mask Learning without training dense models. RS-IML comprises three key components. First, we introduce the intrinsic dimension of objective landscape that the neural network is projected onto a low-dimensional subnetwork to lightweight neural networks without training dense models. Second, we propose intrinsic dimension parameter averaging to aggregate inconsistent local intrinsic models while suppressing the adverse effects of non-intrinsic parameters between local models. Third, we fine-tune the specific parameters of local intrinsic models to mitigate global intrinsic noise for better performance. Extensive experiments demonstrate the effectiveness of our proposed RS-IML. It achieves a significant improvement in model efficiency during lightweighting compared to existing methods while obtaining superior accuracy. Hangyu Ye, Weiying Xie, Xin Zhang 0092, Yibing Lu, Yunsong Li 0001, Leyuan Fang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Change Detection Meets Frequency Learning: A Coarse-to-Fine Dual-Domain Detection NetworkabstractChange detection (CD) aims to distinguish the changed regions in bitemporal images under the same area. Accurate detection needs comprehensive and precise semantic information extraction. Naturally, multiscale feature extraction becomes a fashion. Recent methods implement multiscale feature extraction through complicated convolutional neural network (CNN) modules or Transformers. However, complicated modules are far from the practical demands. First, implementing multiscale feature extraction through multilevel stacking of modules leads to unnecessary computing. Not all regions in images need very detailed detection. Second, respective modules for different scale extractions make discrete resolutions on detection. We need a more continuous recognition of spatial resolutions. Third, inherent weakness in high-frequency learning of CNNs and Transformers degrades the detection of details. In this article, we develop a new multiscale architecture, cross-domain coarse-to-fine (CDC2F) network. Specifically, we first perform coarse detection in the spatial domain. Based on the coarse change map (CM), the bitemporal images will be segmented into blocks, which are then go through a filtering process to retain the mixed blocks. For the retained blocks, a fine detection will be conducted in the frequency domain. Finally, spatial-frequency features are fused to make the final detection. The proposed coarse-to-fine (C2F) strategy guarantees computational efficiency. The cross-domain architecture (spatial and frequency) provides a continuous scale recognition through discrete cosine transform (DCT). And the explicit frequency learning makes detailed detection come true. Extensive experiments on three widely used CD datasets [learning, vision, and remote sensing CD (LEVIR-CD), Wuhan University (WHU), and season-varying CD (SVCD)] show that CDC2F achieves state-of-the-art (SOTA) performance in both the evaluation metrics and the visual presentation, with few parameters and low computational complexity. The source code is available athttps://github.com/Beat992/CDC2F. Wenjin Guo, Yunsong Li 0001, Weiying Xie |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | HSLabeling: Toward Efficient Labeling for Large-Scale Remote Sensing Image Segmentation With Hybrid Sparse LabelingabstractDense pixel-wise labeling of large-scale remote sensing images (RSI) is very time-consuming, while sparse labels (i.e., points, scribbles, or blocks) can be an efficient way to reduce labeling costs. Most existing sparse label-based methods adopt only one type of label for image segmentation, which cannot reflect the complex land covers in the RSI for training the model, thus leading to inferior segmentation performance. We observe that land covers with different shapes and complexity can be optimally represented by different sparse labels. Inspired by this observation, we propose a novel sparse labeling framework, termed Hybrid Sparse Labeling (HSLabeling), for large-scale RSI segmentation. Our HSLabeling can adaptively select the optimal hybrid sparse labels for different land covers, according to labeling cost and segmentation contribution of different sparse labels. Specifically, we first propose a label segmentation contribution information estimation module that estimates the information of different sparse labels according to the diversity and shape of land covers. After that, we propose an Optimal Hybrid Labeling Strategy (OHLS) to assign optimal types of labels for different land covers. In the OHLS, label assignment is formulated as an optimization problem that trades off label segmentation contribution information and labeling cost. We employ the greedy algorithm to efficiently solve the optimization problem and adaptively assign labels for varied land covers. Extensive experiments on three large-scale RSI datasets have demonstrated that our HSLabeling achieves almost fully supervised performance with extremely low labeling costs. In addition, compared with the single type sparse label, HSLabeling can also utilize much lower labeling costs to obtain the same performance. The source code is available at https://github.com/linjiaxing99/HSLabeling. Jiaxing Lin, Zhen Yang 0026, Yinglong Yan, Pedram Ghamisi, Weiying Xie, Leyuan Fang |
IEEE Trans. Image Process. | 6 |
| 2025 | Distributed Deep Learning With Gradient Compression for Big Remote Sensing Image InterpretationabstractFast and reliable interpretation of high-dimensional hyperspectral images (HSIs) can provide great support to remote sensing-based Earth observations. Targets of interest in HSI can be detected using deep neural networks (DNNs) for background learning on an acquired image where the occurrence probability of background samples is much greater than that of targets, accounting for more than 95% of the whole scene. However, there is an increasing gap between theory and feasible application, because of the contradiction between massive hyperspectral data and resource-limited Internet of Things (IoT)/edge device hardware like satellite. To facilitate the deployment of hyperspectral target detection (HTD) in an edge computing environment, we introduce distributed background learning-a decentralized deep learning approach to meet the computing requirements of exploding high-dimensional data and larger DNNs. To address the communication bottleneck caused by gradient exchange during distributed learning, the proposed gradient compression solution, named gradient compression via centroid (GCC), uniquely compresses the most replaceable gradients with redundant information, thereby reducing communication overhead while maintaining accuracy. To illustrate the feasibility of the proposed method, we test it over two very large hyperspectral datasets with a total size of about 3.2 gigabytes (GBs) on a distributed system based on Ring All-reduce. We show that HTD based on distributed background learning outperforms those developed on a single node in terms of speed. Besides, the GCC compresses 50% gradients with only 0.01% loss of target detection accuracy to greatly reduce the communication overhead, surpassing existing gradient compression methods. It is expected that this framework will accelerate the introduction of distributed training on IoT/edge devices. Weiying Xie, Jitao Ma, Tianen Lu, Yunsong Li 0001, Jie Lei 0001, Leyuan Fang, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | MDFL: Multi-Domain Diffusion-Driven Feature LearningabstractHigh-dimensional images, known for their rich semantic information, are widely applied in remote sensing and other fields. The spatial information in these images reflects the object's texture features, while the spectral information reveals the potential spectral representations across different bands. Currently, the understanding of high-dimensional images remains limited to a single-domain perspective with performance degradation. Motivated by the masking texture effect observed in the human visual system, we present a multi-domain diffusion-driven feature learning network (MDFL) , a scheme to redefine the effective information domain that the model really focuses on. This method employs diffusion-based posterior sampling to explicitly consider joint information interactions between the high-dimensional manifold structures in the spectral, spatial, and frequency domains, thereby eliminating the influence of masking texture effects in visual models. Additionally, we introduce a feature reuse mechanism to gather deep and raw features of high-dimensional data. We demonstrate that MDFL significantly improves the feature extraction performance of high-dimensional data, thereby providing a powerful aid for revealing the intrinsic patterns and structures of such data. The experimental results on three multi-modal remote sensing datasets show that MDFL reaches an average overall accuracy of 98.25%, outperforming various state-of-the-art baseline schemes. Code available at https://github.com/LDXDU/MDFL-AAAI-24. Daixun Li, Weiying Xie, Yunsong Li 0001 |
AAAI | 2 |
| 2024 | JointSQ: Joint Sparsification-Quantization for Distributed LearningabstractGradient sparsification and quantization offer a promising prospect to alleviate the communication overhead problem in distributed learning. However, direct combination of the two results in suboptimal solutions, due to the fact that sparsification and quantization haven't been learned together. In this paper, we propose Joint Sparsification-Quantization (JointSQ) inspired by the discovery that sparsification can be treated as 0-bit quantization, regardless of architectures. Specifically, we mathematically formu-late JointSQ as a mixed-precision quantization problem, expanding the solution space. It can be solved by the designed MCKP-Greedy algorithm. Theoretical analysis demon-strates the minimal compression noise of JointSQ, and ex-tensive experiments on various network architectures, including CNN, RNN, and Transformer, also validate this point. Under the introduction of computation overhead consistent with or even lower than previous methods, JointSQ achieves a compression ratio of 1000× on different models while maintaining near-lossless accuracy and brings 1.4× to 2.9× speedup over existing methods. Weiying Xie, Jitao Ma, Yunsong Li 0001, Jie Lei 0001, Donglai Liu, Leyuan Fang |
CVPR | 1 |
| 2024 | Spanning Training Progress: Temporal Dual-Depth Scoring (TDDS) for Enhanced Dataset PruningabstractDataset pruning aims to construct a coreset capable of achieving performance comparable to the original, full dataset. Most existing dataset pruning methods rely on snapshot-based criteria to identify representative samples, often resulting in poor generalization across various pruning and cross-architecture scenarios. Recent studies have addressed this issue by expanding the scope of training dynamics considered, including factors such as forgetting event and probability change, typically using an averaging approach. However, these works struggle to integrate a broader range of training dynamics without overlooking well-generalized samples, which may not be sufficiently highlighted in an averaging manner. In this study, we propose a novel dataset pruning method termed as Temporal Dual-Depth Scoring (TDDS), to tackle this problem. TDDS utilizes a dual-depth strategy to achieve a balance between incorporating extensive training dynamics and identifying representative samples for dataset pruning. In the first depth, we estimate the series of each sample's individual contributions spanning the training progress, ensuring comprehensive integration of training dynamics. In the second depth, we focus on the variability of the sample-wise contributions identified in the first depth to highlight well- generalized samples. Extensive experiments conducted on CIFAR and ImageNet datasets verify the superiority of TDDS over previous SOTA methods. Specifically on CIFAR-100, our method achieves 54.51% accuracy with only 10% training data, surpassing baselines methods by more than 12.69%. Our codes are available at https://github.com/zhangxin-xd/Dataset-Pruning-TDDS. Xin Zhang 0092, Jiawei Du 0002, Yunsong Li 0001, Weiying Xie, Joey Tianyi Zhou |
CVPR | 4 |
| 2024 | DA-BEV: Unsupervised Domain Adaptation for Bird's Eye View Perception
Kai Jiang 0001, Jiaxing Huang 0001, Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Ling Shao 0001, Shijian Lu |
ECCV (82) | 3 |
| 2024 | E4SA: An Ultra-Efficient Systolic Array Architecture for 4-Bit Convolutional Neural NetworksabstractMany studies have demonstrated that 4-bit precision quantization can achieve comparable accuracy to floating-point DNNs, sparking significant interest in efficiently accelerating compressed DNNs, especially 4-bit convolutions, on edge devices. However, we observe that conventional systolic array (SA) architectures designed for DNNs cannot fully exploit the advantages of high DSP computational density offered by 4-bit DSP packing. Although state-of-the-art FPGA-based SA architectures (e.g., AutoSA) exhibit flexibility in accommodating 4-bit DSP packing, they suffer from resource consumption and data supply latency issues, especially when adapting to various convolution spatial sizes. This work introduces a customizable and ultra-efficient SA architectural template for 4-bit convolution, called E4SA. First, we propose a fine-grained row-temporal weight stationary dataflow that aligns with the specific requirements of 4-bit DSP full packing (4bF packing). Based on this, we design a cost-effective SA unit (SAU) composed of 4bF-packing-based processing elements (PEs) to enhance computational efficiency. This includes column-shared packed-data splitters and shift-register-based feature-map/weight fetchers to ensure continuous data supply, all of which are locally interconnected via more cost-effective registers. In addition, we develop a two-level hierarchy SA that decomposes the original large SA into parallel 4×4 SAU sets, which not only allows multiple PEs in the same column to share data splitting and reorganization logic and thus reducing the LUT overhead, but also maintains near-theoretical latency across various convolutional spatial sizes. Experimental results demonstrate that E4SA achieves up to 576.6 GOPS with 13.8× higher GOPS/DSP efficiency and 51.6× higher GOPS/kLUTs efficiency compared to 4-bit AutoSA-based design. Geng Yang 0001, Jie Lei 0001, Zhenman Fang, Junrong Zhang 0002, Weiying Xie, Yunsong Li 0001 |
FPGA | 6 |
| 2024 | SA4: A Comprehensive Analysis and Optimization of Systolic Array Architecture for 4-bit ConvolutionsabstractMany studies have demonstrated that 4-bit precision quantization can maintain accuracy levels comparable to those of floating-point deep neural networks (DNNs). Thus, it has sparked a keen interest in the efficient acceleration of such compressed DNNs, especially 4-bit convolutions, on edge devices. However, we observe that conventional systolic array (SA) architectures, widely adopted for DNN acceleration, fail to fully exploit the high computational density benefits of 4 -bit DSP packing. In this paper, we conduct the first comprehensive analysis of the integration of modern DSP packing techniques (specifically, 4-bit fully DSP packing) into the 4-bit systolic array design for convolutions. First, we introduce a row-temporal weight stationary 4-bit SA dataflow that complements the loop execution order inherent in 4-bit fully DSP packing in conventional SAs, which is called BaseSA. Next, we analyze the performance and resource efficiency of BaseSA, and identify two inefficiencies in the integration: 1) excessive LUT resource utilization that constraints the overall SA size, and 2) large latency gap to the theoretical optimum, due to various stalls in data supplies. To overcome these obstacles, we propose SA4: an HLS-based, customizable, and ultra-efficient hierarchical $\underline{\text { SA}}$ architecture optimized for 4 -bit convolutions. The core unit in SA4 is a delicately designed cost-effective SA unit (SAU), which 1) replaces the costly buffer-based data suppliers for activations and weights with shift-register-based ones, 2) replaces LUT-intensive FIFO connections between SA PEs (processing elements) with registers, and 3) replaces the finite state machines (FSM) and data unpacking logic inside each PE with a global FSM inside each SAU and a data splitter shared by a column of PEs. While such an SAU can only support a small spatial size for an SA due to its delicate design, we further scale it out using an array of SAUs. Experimental results show that our proposed SA4 achieves 1153.2 GOPS on the AMD-Xilinx Ultra96-V2 FPGA, with a $13.8 \times$ increase in GOPS/DSP efficiency and a $49 \times$ increase in GOPS/kLUTs efficiency compared to a straightforward SA and 4-bit DSP packing integration. Our SA4 project is open sourced here: https://github.com/Michaela1224/SA4. Geng Yang 0001, Jie Lei 0001, Zhenman Fang, Junrong Zhang 0002, Weiying Xie, Yunsong Li 0001 |
FPL | 6 |
| 2024 | SDA: Low-Bit Stable Diffusion Acceleration on Edge FPGAsabstractThis paper introduces SDA, the first effort to adapt the expensive stable diffusion (SD) model for edge FPGA deployment. First, we apply quantization-aware training to quantize its weights to 4 -bit and activations to 8 -bit ($W 4 A 8$) with a negligible accuracy loss. Based on that, we propose a high-performance hybrid systolic array (hybridSA) architecture that natively executes convolution and attention operators across varying quantization bit-widths (e.g., $W 4 A 8$ and all 8 -bit $Q K^{T} V$ in attention). To improve computational efficiency, hybridSA integrates diverse DSP packing techniques into hybrid weightstationary and output-stationary dataflows that are optimized for convolution and attention. It also supports flexible dataflow transitions to address the distinct demands of its output sequence by subsequent nonlinear operators. Moreover, we observe that nonlinear operators become the new performance bottleneck after the acceleration of convolution and attention, and offload them onto the FPGA as well. To reduce the latency of each nonlinear operator, we pipeline its own execution at a fine granularity. To minimize the resource utilization of nonlinear operators, we carefully balance their execution with hybridSA in a coarse-grained pipeline. Experimental results demonstrate that our low-bit ($W 4 A 8$) SDA accelerator on the embedded AMDXilinx ZCU102 FPGA achieves a speedup of $97.3 \times$ (which takes about $\mathrm{2 . 1}$ minutes for one SD inference), compared to the original SD-v1.5 model on the ARM Cortex-A53 CPU (which takes about 3.5 hours for one SD inference). Our SDA project is open sourced here: https://github.com/Michaela1224/SDA_code. Geng Yang 0001, Yanyue Xie, Zhong Jia Xue, Sung-En Chang, Yanyu Li, Peiyan Dong, Jie Lei 0001, Weiying Xie, Yanzhi Wang 0001, Xue Lin 0001, Zhenman Fang |
FPL | 8 |
| 2024 | Adaptive Hierarchical Aggregation for Federated Object DetectionabstractIn practical object detection scenarios, distributed data and stringent privacy protections significantly limit the feasibility of traditional centralized training methods. Federated learning (FL) emerges as a promising solution to this dilemma. Nonetheless, the issue of data heterogeneity introduces distinct challenges to federated object detection, evident in diminished object perception, classification and localization abilities. In response, we introduce a task-driven federated learning methodology, dubbed Adaptive Hierarchical Aggregation (FedAHA), tailored to overcome these obstacles. Our algorithm unfolds in two strategic phases from shallow-to-deep layers: (1) Structure-aware Aggregation (SAA) aligns feature extractors during the aggregation phase, thus bolstering the global model's object perception capabilities; (2) Convex Semantic Calibration (CSC) leverages convex function theory to average semantic features instead of model parameters, enhancing the global model's classification and localization precision. We demonstrate experimentally and theoretically the effectiveness of the proposed two modules respectively. Our method consistently outperforming the state-of-the-art methods across multiple valuable application scenarios from 2.26% to 7.61%. Moreover, we build a real FL system using Raspberry Pis to demonstrate that our approach achieves a good trade-off between performance and efficiency. Ruofan Jia, Weiying Xie, Jie Lei 0001, Yunsong Li 0001 |
ACM Multimedia | 2 |
| 2024 | FedSLS: Exploring Federated Aggregation in Saliency Latent SpaceabstractFederated Learning (FL) is an emerging direction in distributed machine learning that enables jointly training a global model without sharing data with server. However, data heterogeneity biases the parameter aggregation at the server, leading to slower convergence and poorer accuracy of the global model. To cope with this, most of the existing works involve enforcing regularization in local optimization or improving the model aggregation scheme at the server. Though effective, they lack a deep understanding of cross-client features. In this paper, we propose a saliency latent space feature aggregation method (FedSLS) across federated clients. By Guided BackPropagation (GBP), we transform deep models into powerful and flexible visual fidelity encoders, applicable to general state inputs across different image domains, and achieve powerful aggregation in the form of saliency latent features. Notably, since GBP is label-insensitive, it is sufficient to capture saliency features only once on each client. Experimental results demonstrate that FedSLS leads to significant improvements over the state-of-the-arts in terms of accuracies, especially in highly heterogeneous settings. For example, on CIFAR-10 dataset, FedSLS achieves 63.43% accuracy within the strongly heterogeneous environment α=0.05, which is 6% to 23% higher than other baselines. Hengyi Wang, Weiying Xie, Jitao Ma, Daixun Li, Yunsong Li 0001 |
ACM Multimedia | 2 |
| 2024 | Adaptive Pruning of Channel Spatial Dependability in Convolutional Neural NetworksabstractDeep Convolutional Neural Networks (CNNs) have demonstrated excellent performance in various multimedia application scenarios. However, complex models often require significant computational resources and energy costs. Therefore, CNN compression is crucial for addressing deployment challenges of multimedia application on resource constrained edge devices. However, existing CNN channel pruning strategies primarily focus on the "weights" or "activations" of the model, overlooking its "interpretability" information. In this paper, we explore CNN pruning strategies from the perspective of model interpretability. We model the correspondence between channel feature maps and interpretable visual perception based on class saliency maps, aiming to assess the contribution of each channel to the desired output. Additionally, we utilize Discrete Wavelet Transform (DWT) to capture the global features and structure of class saliency maps. Based on this, we propose a Channel Spatial Dependability (CSD) metric, evaluating the importance and contribution of channels in a bidirectional manner to guide model pruning. And we dynamically adjust the pruning rate of each layer based on performance changes, in order to achieve more accurate and efficient adaptive pruning. Our method achieves significant results across a range of different networks and datasets. For instance, we achieved a 51.3% pruning on the ResNet-56 model while maintaining an accuracy of 94.16%, outperforming feature-map or other State-of-the-Art (SOTA). Weiying Xie, Mei Yuan, Jitao Ma, Yunsong Li 0001 |
ACM Multimedia | 1 |
| 2024 | Domain Adaptation for Large-Vocabulary Object DetectorsabstractLarge-vocabulary object detectors (LVDs) aim to detect objects of many categories, which learn super objectness features and can locate objects accurately while applied to various downstream data. However, LVDs often struggle in recognizing the located objects due to domain discrepancy in data distribution and object vocabulary. At the other end, recent vision-language foundation models such as CLIP demonstrate superior open-vocabulary recognition capability.
This paper presents KGD, a Knowledge Graph Distillation technique that exploits the implicit knowledge graphs (KG) in CLIP for effectively adapting LVDs to various downstream domains.
KGD consists of two consecutive stages: 1) KG extraction that employs CLIP to encode downstream domain data as nodes and their feature distances as edges, constructing KG that inherits the rich semantic relations in CLIP explicitly;
and 2) KG encapsulation that transfers the extracted KG into LVDs to enable accurate cross-domain object classification.
In addition, KGD can extract both visual and textual KG independently, providing complementary vision and language knowledge for object localization and object classification in detection tasks over various downstream domains.
Experiments over multiple widely adopted detection benchmarks show that KGD outperforms the state-of-the-art consistently by large margins.
Codes will be released. Kai Jiang 0001, Jiaxing Huang 0001, Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Ling Shao 0001, Shijian Lu |
NeurIPS | 3 |
| 2024 | E2E-MFD: Towards End-to-End Synchronous Multimodal Fusion DetectionabstractMultimodal image fusion and object detection are crucial for autonomous driving. While current methods have advanced the fusion of texture details and semantic information, their complex training processes hinder broader applications. Addressing this challenge, we introduce E2E-MFD, a novel end-to-end algorithm for multimodal fusion detection. E2E-MFD streamlines the process, achieving high performance with a single training phase. It employs synchronous joint optimization across components to avoid suboptimal solutions associated to individual tasks. Furthermore, it implements a comprehensive optimization strategy in the gradient matrix for shared parameters, ensuring convergence to an optimal fusion detection configuration. Our extensive testing on multiple public datasets reveals E2E-MFD's superior capabilities, showcasing not only visually appealing image fusion but also impressive detection outcomes, such as a 3.9\% and 2.0\% $\text{mAP}_{50}$ increase on horizontal object detection dataset M3FD and oriented object detection dataset DroneVehicle, respectively, compared to state-of-the-art approaches. Mingxiang Cao, Weiying Xie, Jie Lei 0001, Daixun Li, Wenbo Huang 0001, Yunsong Li 0001 |
NeurIPS | 3 |
| 2024 | Hyperspectral band selection via region-wise latent feature fusion and graph filter embedded subspace clustering
Minhui Wang, Chang Tang, Weiying Xie, Xianju Li, Jiangfeng Xu |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | FedDiff: Diffusion Model Driven Federated Learning for Multi-Modal and Multi-ClientsabstractWith the rapid development of imaging sensor technology in the field of remote sensing, multi-modal remote sensing data fusion has emerged as a crucial research direction for land cover classification tasks. While diffusion models have made great progress in generative models and image classification tasks, existing models primarily focus on single-modality and single-client control, that is, the diffusion process is driven by a single modal in a single computing node. To facilitate the secure fusion of heterogeneous data from clients, it is necessary to enable distributed multi-modal control, such as merging the hyperspectral data of organization A and the LiDAR data of organization B privately on each base station client. In this study, we propose a multi-modal collaborative diffusion federated learning framework called FedDiff. Our framework establishes a dual-branch diffusion model feature extraction setup, where the two modal data are inputted into separate branches of the encoder. Our key insight is that diffusion models driven by different modalities are inherently complementary in terms of potential denoising steps on which bilateral connections can be built. Considering the challenge of private and efficient communication between multiple clients, we embed the diffusion model into the federated learning communication structure, and introduce a lightweight communication module. Qualitative and quantitative experiments validate the superiority of our framework in terms of image quality and conditional consistency. To the best of our knowledge, this is the first instance of deploying a diffusion model into a federated learning framework, achieving optimal both privacy protection and performance for heterogeneous data. Our FedDiff surpasses existing methods in terms of performance on three multi-modal datasets, achieving a classification average accuracy of 96.77% while reducing the communication cost. Daixun Li, Weiying Xie, Yibing Lu, Yunsong Li 0001, Leyuan Fang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Markov-PQ: Joint Pruning-Quantization via Learnable Markov ChainabstractVarious network compression methods, such as pruning and quantization, have been proposed to synergistically reduce resource requirements. However, existing joint compression works are based on black-box optimization and do not interpret the interaction mechanism between these two compression techniques, leading to a slow and unstable convergence of compression strategy. To address this issue, we present Markov-PQ, the first interpretable pruning-quantization co-compression framework using a Markov Chain. In Markov-PQ, the joint strategy search is modeled as a Markov Chain and decoupled with Bayes Rule into pruning and quantization strategy searching. Specifically, the quantization state accounts for the co-compression state from the last time and is updated by a learnable transition probability matrix. To ensure differentiability, we design a forward-hard and backward-soft quantization. The pruning state is influenced not only by the last co-compression state but also by the concurrent quantization state. In addition, to perceive the current layer-wise bit sensitivity and alleviate the long-tail problem, a complexity-aware regularizer is devised to re-evaluate the filter importance. Extensive experiments demonstrate the superiority of Markov-PQ. For example, with an accuracy loss of only 0.33%, we can achieve a$56.12\times $acceleration for ResNet-18 on ImageNet2012. Yunsong Li 0001, Xin Zhang 0092, Weiying Xie, Leyuan Fang, Jiawei Du 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Multimodal Informative ViT: Information Aggregation and Distribution for Hyperspectral and LiDAR ClassificationabstractIn multimodal land cover classification (MLCC), a common challenge is the redundancy in data distribution, where task-irrelevant information from multiple modalities can hinder the effective integration of their unique features. To tackle this, we introduce the Multimodal Informative Vit (MIVit), a system with an innovative information aggregate-distributing mechanism. This approach redefines redundancy levels and integrates performance-aware elements into the fused representation, facilitating the learning of semantics in both forward and backward directions. MIVit stands out by significantly reducing redundancy in the empirical distribution of each modality’s separate and fused features. It employs oriented attention fusion (OAF) for extracting shallow local shape features across modalities in horizontal and vertical dimensions, and a Transformer feature extractor for extracting deep global features through long-range attention. We also propose an information aggregation constraint (IAC) based on mutual information, designed to remove redundant information and preserve complementary information within embedded features. Additionally, the information distribution flow (IDF) in MIVit enhances performance-awareness by distributing global classification information across different modalities’ feature maps. This architecture also addresses missing modality challenges with lightweight independent modality classifiers, reducing the computational load typically associated with Transformers. Our results show that MIVit’s bidirectional aggregate-distributing mechanism between modalities is highly effective, achieving an average overall accuracy of 95.56% across three multimodal datasets. This performance surpasses current state-of-the-art methods in MLCC. The code for MIVit is accessible at https://github.com/icey-zhang/MIViT. Jie Lei 0001, Weiying Xie, Geng Yang 0001, Daixun Li, Yunsong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | MMIF: Interpretable Hyperspectral and Multispectral Image Fusion via Maximum Mutual InformationabstractFusion-based hyperspectral image (HSI) super-resolution (SR) aims to recover high-resolution HSI (hr-HSI) from its two degraded modalities, that is, low-resolution HSI (lr-HSI) and high-resolution multispectral image (hr-MSI). The resulting HSI-SR image can be explained from two viewpoints, an inverse problem solution or a spatial–spectral information fusion product. Recent methods focusing on the former point are limited by demands of accurate degradation parameters and precise prior assumptions. On the contrary, the other research line avoids these drawbacks and offers a more flexible design. However, recent methods implicitly handle information fusion. The interaction between lr-HSI and hr-MSI only lies in the extracted feature domain, not directly on the data. Moreover, the proposed modules in recent methods promote information fusion from the perspective of deep learning and neglect the specialty of the HSI domain, leading to weak interpretability and poor reliability in practice. Considering the essence of the HSI-SR problem and the inherent property of HSIs, in this article, we propose a maximum mutual information (MMI) strategy. At first, we model the HSI-SR problem in a new variance inference (VI) architecture. This new VI model simulates the physical process of HSI pairs imaging and provides convenience for the MMI strategy. Then, we insert the MMI strategy in the VI model. The MMI strategy promotes information fusion with quantitative constraints on the information interaction between lr-HSI and hr-MSI. Finally, we implement the VI architecture with a neural network. Benefitting from the MMI strategy, a simple network structure can achieve efficient fusion performance, which indicates that MMI frees DL-based HSI-SR methods from the complicated structure design. The experimental results on synthetic and real datasets demonstrate the superiority of our method to state-of-the-art methods in terms of effectiveness, generalization, and interpretability. Yunsong Li 0001, Wen-jin Guo, Weiying Xie, Tao Jiang 0031, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | FedFusion: Manifold-Driven Federated Learning for Multi-Satellite and Multi-Modality FusionabstractMulti-Satellite, multi-modality in-orbit fusion is a challenging task as it explores the fusion representation of complex high-dimensional data under limited computational resources. Deep neural networks can reveal the underlying distribution of multimodal remote sensing data, but the in-orbit fusion of multimodal data is more difficult because of the limitations of different sensor imaging characteristics, especially when the multimodal data follow nonindependent identically distribution (Non-IID) distributions. To address this problem while maintaining classification performance, this article proposes a manifold-driven multi-modality fusion framework, FedFusion, which randomly samples local data on each client to jointly estimate the prominent manifold structure of shallow features of each client and explicitly compresses the feature matrices into a low-rank subspace through cascading and additive approaches, which is used as the feature input of the subsequent classifier. Considering the physical space limitations of the satellite constellation, we developed a multimodal federated learning (FL) module designed specifically for manifold data in a deep latent space. This module achieves iterative updating of the subnetwork parameters of each client through global weighted averaging, constructing a framework that can represent compact representations of each client. The proposed framework surpasses existing methods in terms of performance on three multimodal datasets, achieving a classification average accuracy of 94.35% while compressing communication costs by a factor of 4. Furthermore, extensive numerical evaluations of real-world satellite images were conducted on the orbiting edge computing architecture based on Jetson TX2 industrial modules, which demonstrated that FedFusion significantly reduced training time by 48.4 min (15.18%) while optimizing accuracy. The codes will be available at:https://github.com/LDXDU/FedFusion. Daixun Li, Weiying Xie, Yunsong Li 0001, Leyuan Fang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Diffusion Models Meet Remote Sensing: Principles, Methods, and PerspectivesabstractAs a newly emerging advance in deep generative models, diffusion models have achieved state-of-the-art results in many fields, including computer vision, natural language processing, and molecule design. The remote sensing (RS) community has also noticed the powerful ability of diffusion models and quickly applied them to a variety of tasks for image processing. Given the rapid increase in research on diffusion models in the field of RS, it is necessary to conduct a comprehensive review of existing diffusion model-based RS papers, to help researchers recognize the potential of diffusion models and provide some directions for further exploration. Specifically, this article first introduces the theoretical background of diffusion models, and then systematically reviews the applications of diffusion models in RS, including image generation, enhancement, and interpretation. Finally, the limitations of existing RS diffusion models and worthy research directions for further exploration are discussed and summarized. Yidan Liu, Jun Yue 0004, Shaobo Xia, Pedram Ghamisi, Weiying Xie, Leyuan Fang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | SwiMDiff: Scene-Wide Matching Contrastive Learning With Diffusion Constraint for Remote Sensing ImageabstractWith recent advancements in aerospace technology, the volume of unlabeled remote sensing image (RSI) data has increased dramatically. Effectively leveraging this data through self-supervised learning (SSL) is vital in the field of remote sensing. However, current methodologies, particularly contrastive learning (CL), a leading SSL method, encounter specific challenges in this domain. Firstly, CL often mistakenly identifies geographically adjacent samples with similar semantic content as negative pairs, leading to confusion during model training. Secondly, as an instance-level discriminative task, it tends to neglect the essential fine-grained features and complex details inherent in unstructured RSIs. To overcome these obstacles, we introduce SwiMDiff, a novel self-supervised pre-training framework designed for RSIs. SwiMDiff employs a scene-wide matching approach that effectively recalibrates labels to recognize data from the same scene as false negatives. This adjustment makes CL more applicable to the nuances of remote sensing. Additionally, SwiMDiff seamlessly integrates CL with a diffusion model. Through the implementation of pixel-level diffusion constraints, we enhance the encoder’s ability to capture both the global semantic information and the fine-grained features of the images more comprehensively. Our proposed framework significantly enriches the information available for downstream tasks in remote sensing. Demonstrating exceptional performance in change detection and land-cover classification tasks, SwiMDiff proves its substantial utility and value in the field of remote sensing. Jiayuan Tian, Jie Lei 0001, Weiying Xie, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Ebbinghaus-Curve Guided Low-Rank Component-Induced Attention for Multisource Remote Sensing ClassificationabstractThe integration of multisource remote sensing (RS) data is crucial in land use and land cover (LULC) studies, offering numerous applications. Using diverse data sources enhances the accuracy of land cover classification. However, due to differences in imaging mechanisms, existing methods face challenges in capturing complex local and global relationships. Moreover, current multimodal fusion approaches often fail to efficiently preserve heterogeneous data, leading to the overfusion of redundant features. To address these challenges, we propose Ebbinghaus-curve guided multisource RS classification network (ECNet). This framework maximizes the benefits of convolutional operators for local feature representation and leverages Transformer architecture for learning long-distance dependencies. Inspired by the forgetting strategy of human brain neurons, we extend the concept of information loss to address feature preservation issues in neural networks. ECNet effectively preserves essential features while eliminating redundancy in high-dimensional manifold structures, thus mitigating overfitting caused by redundant multimodal features. Extensive experiments on four publicly available datasets demonstrate the competitiveness of ECNet in classification tasks. Notably, on the Houston2013 dataset, ECNet achieves an impressive overall accuracy (OA) of 96.63%, surpassing various state-of-the-art baseline approaches. The code is available athttps://github.com/lyb10087/ECNetfor the sake of reproducibility. Weiying Xie, Yibing Lu, Daixun Li, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | RS-DGC: Exploring Neighborhood Statistics for Dynamic Gradient Compression on Remote Sensing Image InterpretationabstractDistributed deep learning has recently been attracting more attention in remote sensing (RS) applications due to the challenges posed by the increased amount of open data that are produced daily by Earth observation programs. However, the high communication costs of sending model updates among multiple nodes are a significant bottleneck for scalable distributed learning. Gradient sparsification has been validated as an effective gradient compression (GC) technique for reducing communication costs and thus accelerating the training speed. Existing state-of-the-art gradient sparsification methods are mostly based on the “larger-absolute-more-important” criterion, ignoring the importance of small gradients, which is generally observed to affect the performance. Inspired by informative representation of manifold structures from neighborhood information, we propose a simple yet effective dynamic gradient compression scheme leveraging neighborhood statistics indicator for RS image interpretation, termed RS-DGC. We first enhance the interdependence between gradients by introducing the gradient neighborhood to reduce the effect of random noise. The key component of RS-DGC is a Neighborhood Statistical Indicator (NSI), which can quantify the importance of gradients within a specified neighborhood on each node to sparsify the local gradients before gradient transmission in each iteration. Further, a layer-wise dynamic compression scheme is proposed to track the importance changes of each layer in real time. Extensive downstream tasks validate the superiority of our method in terms of intelligent interpretation of RS images. For example, we achieve an accuracy improvement of 0.51% with more than 50× communication compression on the NWPU-RESISC45 dataset using VGG-19 network. To the best of our knowledge, this is the first gradient compression method designed for RS images and downstream tasks, achieving a successful trade-off between high compression ratio and performance. Weiying Xie, Jitao Ma, Daixun Li, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | When Vectorization Meets Change DetectionabstractIn long-term Earth observation, change detection (CD) is a crucial and intricate task with applications spanning diverse fields, including land resource planning and natural disaster monitoring. Most existing CD approaches typically output segmentation results in raster format. However, raster format results suffer from higher memory usage, poorer shape accuracy, magnified distortions, and challenges in topological editing. To address the issues of raster format, we propose a novel end-to-end change vectorization network (CVNet), which is the first attempt to extract changes using vector format. The CVNet directly learns the vector components of changed objects and uses them to construct vectors. Specifically, since the vectorization of CD faces the inherent imbalance between changed and unchanged samples, we first introduce the Change-Collector to collect the changed regions and combine them into more compact samples. Next, the vector components learning model (VCLM) is introduced to capture the fundamental components for constructing the vectors, including change maps, junction positions, and segmentation masks. Finally, the changed instances obtained from the masks are used to divide and connect junctions to generate the vector output. To verify the effectiveness of the proposed framework, we construct two building change vectorization datasets by modifying the WHU-CD and LEVIR-CD benchmarks. Experimental results demonstrate that the CVNet outperforms the existing postprocess vectorization methods in terms of the visual effect and all evaluation metrics. The dataset and source code will be made publicly available athttps://github.com/yyyyll0ss/CVNet. Yinglong Yan, Jun Yue 0004, Jiaxing Lin, Weiying Xie, Leyuan Fang |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Distribution-Aware Interactive Attention Network and Large-Scale Cloud Recognition Benchmark on FY-4A Satellite ImageabstractAccurate cloud recognition and warning are crucial for various applications, including in-flight support, weather forecasting, and climate research. However, recent deep learning algorithms have predominantly focused on detecting cloud regions in satellite imagery, with insufficient attention to the specificity required for accurate cloud recognition. This limitation inspired us to develop the novel FY-4A-Himawari-8 (FYH) dataset, which includes nine distinct cloud categories and uses precise domain adaptation methods to align 70419 image-label pairs (including 110000 train/5500 test$100\times 100$size images) in terms of projection, temporal resolution, and spatial resolution, thereby facilitating the training of supervised deep learning networks. Given the complexity and diversity of cloud formations, we have thoroughly analyzed the challenges inherent to cloud recognition tasks, examining the intricate characteristics and distribution of the data. To effectively address these challenges, we designed a distribution-aware interactive-attention network (DIAnet), which preserves pixel-level details through a high-resolution branch and a parallel multiresolution cross-branch. We also integrated a distribution-aware loss (DAL) to mitigate the imbalance across cloud categories. An interactive attention module (IAM) further enhances the robustness of feature extraction combined with spatial and channel information. Empirical evaluations on the FYH dataset demonstrate that our method outperforms other cloud recognition networks, achieving superior performance in terms of mean intersection over union (mIoU). The code for implementing DIAnet is available athttps://github.com/icey-zhang/DIAnet. Jie Lei 0001, Weiying Xie, Kai Jiang 0001, Xin Zhang 0092, Mingxiang Cao, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | SemiRS-COC: Semi-Supervised Classification for Complex Remote Sensing Scenes With Cross-Object ConsistencyabstractSemi-supervised learning (SSL), which aims to learn with limited labeled data and massive amounts of unlabeled data, offers a promising approach to exploit the massive amounts of satellite Earth observation images. The fundamental concept underlying most state-of-the-art SSL methods involves generating pseudo-labels for unlabeled data based on image-level predictions. However, complex remote sensing (RS) scene images frequently encounter challenges, such as interference from multiple background objects and significant intra-class differences, resulting in unreliable pseudo-labels. In this paper, we propose the SemiRS-COC, a novel semi-supervised classification method for complex RS scenes. Inspired by the idea that neighboring objects in feature space should share consistent semantic labels, SemiRS-COC utilizes the similarity between foreground objects in RS images to generate reliable object-level pseudo-labels, effectively addressing the issues of multiple background objects and significant intra-class differences in complex RS images. Specifically, we first design a Local Self-Learning Object Perception (LSLOP) mechanism, which transforms multiple background objects interference of RS images into usable annotation information, enhancing the model's object perception capability. Furthermore, we present a Cross-Object Consistency Pseudo-Labeling (COCPL) strategy, which generates reliable object-level pseudo-labels by comparing the similarity of foreground objects across different RS images, effectively handling significant intra-class differences. Extensive experiments demonstrate that our proposed method achieves excellent performance compared to state-of-the-art methods on three widely-adopted RS datasets. Jun Yue 0004, Weiying Xie, Leyuan Fang |
IEEE Trans. Image Process. | 4 |
| 2024 | HTD-TS3: Weakly Supervised Hyperspectral Target Detection Based on Transformer via Spectral-Spatial SimilarityabstractAs an advanced technique in remote sensing, hyperspectral target detection (HTD) is widely concerned in civilian and military applications. However, the limitation of prior and heterogeneous backgrounds makes HTD models sensitive to data corruption under various interference from the environment. In this article, a novel united HTD framework based on the concept of transformer is proposed to extract [HTD based on transformer via spectral-spatial similarity (HTD-TS3)] under weak supervision, which opens up more flexible ways to study HTD. For the first time, the transformer mechanism is introduced into the HTD task to extract spectral and spatial features in a unified optimization procedure. By modeling long-range dependence among spectra, it realizes spectral-spatial joint inference based on long-range context, which addresses the issues of insufficient utilization of spatial information. To provide samples for weakly supervised learning (WSL), the coarse sample selection and spectral sequence construction in an efficient way are proposed, which makes full use of limited prior information. Finally, an exponential constrained nonlinear function is adopted to acquire pixel-level prediction via combining discriminative spectral-spatial features and coarse spatial information. Experiments on real hyperspectral images (HSIs) captured by different sensors at various scenes verify the effectiveness and efficiency of HTD-TS3. Weiying Xie, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Block-Wise Partner Learning for Model CompressionabstractDespite the great potential of convolutional neural networks (CNNs) in various tasks, the resource-hungry nature greatly hinders their wide deployment in cost-sensitive and low-powered scenarios, especially applications in remote sensing. Existing model pruning approaches, implemented by a "subtraction" operation, impose a performance ceiling on the slimmed model. Self-knowledge distillation (Self-KD) resorts to auxiliary networks that are only active in the training phase for performance improvement. However, the knowledge is holistic and crude, and the learning-based knowledge transfer is mediate and lossy. Here, we propose a novel model-compression method, termed block-wise partner learning (BPL), which comprises "extension" and "fusion" operations and liberates the compressed model from the bondage of baseline. Different from the Self-KD, the proposed BPL creates a partner for each block for performance enhancement in training. For the model to absorb more diverse information, a diversity loss (DL) is designed to evaluate the difference between the original block and the partner. Besides, the partner is fused equivalently instead of being discarded directly. After training, we can simply adopt the fused compressed model that contains the enhancement information of partners but with fewer parameters and less inference cost. As validated using the UC Merced land-use, NWPU-RESISC45, and RSD46-WHU datasets, the BPL demonstrates superiority over other compared model-compression approaches. For example, it attains a substantial floating-point operations (FLOPs) reduction of 73.97% with only 0.24 accuracy (ACC.) loss for ResNet-50 on the UC Merced land-use dataset. The code is available at https://github.com/zhangxin-xd/BPL. Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Kai Jiang 0001, Leyuan Fang, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | HyBNN: Quantifying and Optimizing Hardware Efficiency of Binary Neural NetworksabstractBinary neural network (BNN), where both the weight and the activation values are represented with one bit, provides an attractive alternative to deploy highly efficient deep learning inference on resource-constrained edge devices. However, our investigation reveals that, to achieve satisfactory accuracy gains, state-of-the-art (SOTA) BNNs, such as FracBNN and ReActNet, usually have to incorporate various auxiliary floating-point components and increase the model size, which in turn degrades the hardware performance efficiency. In this article, we aim to quantify such hardware inefficiency in SOTA BNNs and further mitigate it with negligible accuracy loss. First, we observe that the auxiliary floating-point (AFP) components consume an average of 93% DSPs, 46% LUTs, and 62% FFs, among the entire BNN accelerator resource utilization. To mitigate such overhead, we propose a novel algorithm-hardware co-design, called FuseBNN , to fuse those AFP operators without hurting the accuracy. On average, FuseBNN reduces AFP resource utilization to 59% DSPs, 13% LUTs, and 16% FFs. Second, SOTA BNNs often use the compact MobileNetV1 as the backbone network but have to replace the lightweight 3 × 3 depth-wise convolution (DWC) with the 3 × 3 standard convolution (SC, e.g., in ReActNet and our ReActNet-adapted BaseBNN) or even more complex fractional 3 × 3 SC (e.g., in FracBNN) to bridge the accuracy gap. As a result, the model parameter size is significantly increased and becomes 2.25× larger than that of the 4-bit direct quantization with the original DWC (4-Bit-Net); the number of multiply-accumulate operations is also significantly increased so that the overall LUT resource usage of BaseBNN is almost the same as that of 4-Bit-Net. To address this issue, we propose HyBNN , where we binarize depth-wise separation convolution (DSC) blocks for the first time to decrease the model size and incorporate 4-bit DSC blocks to compensate for the accuracy loss. For the ship detection task in synthetic aperture radar imagery on the AMD-Xilinx ZCU102 FPGA, HyBNN achieves a detection accuracy of 94.8% and a detection speed of 615 frames per second (FPS), which is 6.8× faster than FuseBNN+ (94.9% accuracy) and 2.7× faster than 4-Bit-Net (95.9% accuracy). For image classification on the CIFAR-10 dataset on the AMD-Xilinx Ultra96-V2 FPGA, HyBNN achieves 1.5× speedup and 0.7% better accuracy over SOTA FracBNN. Geng Yang 0001, Jie Lei 0001, Zhenman Fang, Yunsong Li 0001, Weiying Xie |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2023 | Toward Stable, Interpretable, and Lightweight Hyperspectral Super-ResolutionabstractFor real applications, existing HSI-SR methods are not only limited to unstable performance under unknown scenarios but also suffer from high computation consumption. In this paper, we develop a new coordination optimization framework for stable, interpretable, and lightweight HSI-SR. Specifically, we create a positive cycle between fusion and degradation estimation under a new probabilistic framework. The estimated degradation is applied to fusion as guidance for a degradation-aware HSI-SR. Under the framework, we establish an explicit degradation estimation method to tackle the indeterminacy and unstable performance caused by the black-box simulation in previous methods. Considering the interpretability in fusion, we integrate spectral mixing prior into the fusion process, which can be easily realized by a tiny autoencoder, leading to a dramatic release of the computation burden. Based on the spectral mixing prior, we then develop a partial fine-tune strategy to reduce the computation cost further. Comprehensive experiments demonstrate the superiority of our method against the state-of-the-arts under synthetic and real datasets. For instance, we achieve a 2.3 dB promotion on PSNR with$120\times$model size reduction and$4300 \times$FLOPs reduction under the CAVE dataset. Code is available in https://github.com/WenjinGuo/DAEM. Wen-jin Guo, Weiying Xie, Kai Jiang 0001, Yunsong Li 0001, Jie Lei 0001, Leyuan Fang |
CVPR | 2 |
| 2023 | HyBNN: Quantifying and Optimizing Hardware Efficiency of Binary Neural NetworksabstractBinary neural network (BNN) has recently presented a promising opportunity for deep learning inferences on resource-constrained edge devices. Using extreme data precision, i.e., 1-bit weight and 1-bit activation, BNN not only significantly reduces the network memory footprint, but also trades massive multiply-accumulate operations for much cheaper logical XNOR and population count operations. However, our investigation reveals that, to achieve satisfactory accuracy gains, state-of-the-art (SOTA) BNNs, such as FracBNN [4] and ReActNet [1], usually have to incorporate various auxiliary floating-point ($AFP$) components and increase the model size, which in turn degrades the hardware performance efficiency. Geng Yang 0001, Jie Lei 0001, Zhenman Fang, Yunsong Li 0001, Weiying Xie |
FCCM | 6 |
| 2023 | Weakly supervised adversarial learning via latent space for hyperspectral target detection
Weiying Xie, Yunsong Li 0001, Kai Jiang 0001, Jie Lei 0001, Qian Du 0001 |
Pattern Recognit. | 2 |
| 2023 | Filter Pruning via Learned Representation Median in the Frequency DomainabstractIn this article, we propose a novel filter pruning method for deep learning networks by calculating the learned representation median (RM) in frequency domain (LRMF). In contrast to the existing filter pruning methods that remove relatively unimportant filters in the spatial domain, our newly proposed approach emphasizes the removal of absolutely unimportant filters in the frequency domain. Through extensive experiments, we observed that the criterion for "relative unimportance" cannot be generalized well and that the discrete cosine transform (DCT) domain can eliminate redundancy and emphasize low-frequency representation, which is consistent with the human visual system. Based on these important observations, our LRMF calculates the learned RM in the frequency domain and removes its corresponding filter, since it is absolutely unimportant at each layer. Thanks to this, the time-consuming fine-tuning process is not required in LRMF. The results show that LRMF outperforms state-of-the-art pruning methods. For example, with ResNet110 on CIFAR-10, it achieves a 52.3% FLOPs reduction with an improvement of 0.04% in Top-1 accuracy. With VGG16 on CIFAR-100, it reduces FLOPs by 35.9% while increasing accuracy by 0.5%. On ImageNet, ResNet18 and ResNet50 are accelerated by 53.3% and 52.7% with only 1.76% and 0.8% accuracy loss, respectively. The code is based on PyTorch and is available at https://github.com/zhangxin-xd/LRMF. Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001 |
IEEE Trans. Cybern. | 2 |
| 2023 | A Transformer-Based Network for Hyperspectral Object TrackingabstractWith the abundant spectral information, the hyperspectral images could be benefit for tracking the target in various application scenarios. Most of the predominant hyperspectral object tracking methods were based on transferred red–green–blue (RGB) object tracking networks, since the lack of training data. Different strategies of processing the hyperspectral images to adapt the transferred RGB object tracking networks have been exploited. However, the existing strategies led to lose the spectral information or the interaction information between bands and have shown limited performances. In this article, a novel transformer-based hyperspectral object tracking algorithm (Trans-HST) is proposed to make advantages of the spectral information with transformer modules. In Trans-HST, the cross-band groups of feature enhancement (CBFE) is introduced to reduce the negative effects of the interaction information loss. To address the problem of spectral information loss, the transformer-based deep features’ fusion (TDFF) fuses the deep features corresponding to different groups of bands in the hyperspectral images and integrates the deep feature corresponding to the original hyperspectral images into the fused features. Experiments on the commonly used hyperspectral object tracking dataset have been applied to verify the effectiveness of the two proposed modules, and they indicate the superior performance of the Trans-HST comparing with other RGB and hyperspectral trackers. Langkun Chen, Pan Liu 0009, Weiying Xie, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | CBFF-Net: A New Framework for Efficient and Accurate Hyperspectral Object TrackingabstractVisual object tracking is a fundamental task in computer vision, and thrived in recent decades. With the development of snapshot hyperspectral sensors, efforts have been made to exploit tracking the object with hyperspectral (HS) videos to overcome the inherent limitation of RGB images. Existing HS tracking algorithms extract the deep features from image data separately, which break the interaction information between bands. Therefore, the discrimination ability of HS trackers is limited and the efficiency of the existing HS algorithms is low. In this paper, a novel algorithm (CBFF-Net) is proposed for HS object tracking to improve the discrimination ability and reduce the computational complexity. Specifically, the backbone and head network are implemented with modules of a transferred RGB object tracking network to carry out the HS target tracking task while maintaining the discrimination ability learned from RGB data. Moreover, a bi-directional multiple deep feature fusion (BMDFF) module is proposed to fuse the features extracted from different bands of the HS images, and a cross-band group attention (CBGA) module is introduced to learn interaction information across bands of the HS images. Experiments results indicate the superiority in performance of CBFF-Net, and it runs at 24 frames per second. Pan Liu 0009, Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | A Semantic Transferred Priori for Hyperspectral Target Detection With Spatial-Spectral AssociationabstractHyperspectral target detection is a crucial application that encompasses military, environmental, and civil needs. Target detection algorithms that have prior knowledge often assume a fixed laboratory target spectrum, which can differ significantly from the test image in the scene. This discrepancy can be attributed to various factors such as atmospheric conditions and sensor internal effects, resulting in decreased detection accuracy. To address this challenge, this article introduces a novel method for detecting hyperspectral image (HSI) targets with certain spatial information, referred to as the semantic transferred priori for hyperspectral target detection with spatial–spectral association (SSAD). Considering that the spatial textures of the HSI remain relatively constant compared to the spectral features, we propose to extract a unique and precise target spectrum from each image data via target detection in its spatial domain. Specifically, employing transfer learning, we designed a semantic segmentation network adapted for HSIs to discriminate the spatial areas of targets and then aggregated a customized target spectrum with those spectral pixels localized. With the extracted target spectrum, spectral dimensional target detection is performed subsequently by the constrained energy minimization (CEM) detector. The final detection results are obtained by combining an attention generator module to aggregate target features and deep stacked feature fusion (DSFF) module to hierarchically reduce the false alarm rate. Experiments demonstrate that our proposed method achieves higher detection accuracy and superior visual performance compared to the other benchmark methods. Jie Lei 0001, Simin Xu, Weiying Xie, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | A Model-Driven Deep Mixture Network for Robust Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) aims to identify samples with unknown atypical spectra from the background. Deep learning (DL)-based methods, particularly autoencoders (AEs), have proven effective in uncovering the underlying profiles for HAD. However, in real-world applications of hyperspectral images (HSIs), complex background land-covers and anomaly corruptions are common, leading to two issues: 1) A low-dimensional manifold characterized by DL-based HAD methods can only reveal a few underlying variation factors of the background distribution and cannot capture the complex structures behind land-covers of all categories. 2) DL-based HAD methods trained on anomaly-contaminated HSIs tend to overfit specific anomalies, resulting in poor background characterization. To tackle these issues, this study presents a novel and robust framework for HAD called Model-Driven Deep Mixture Network (MDMN) that combines the strengths of model-driven and data-driven approaches while emphasizing interpretability. By assuming that the background, consisting of various land-covers, arises from a mixture of low-dimensional manifolds, the MDMN incorporates a novel deep mixture module to comprehensively characterize the background. This module utilizes a low-dimensional manifold learned by an AE to represent a specific category of background land-covers. To mitigate the impact of anomaly corruptions, the MDMN incorporates a convex relaxation of a sparse constraint, which helps prevent overfitting anomalies. Extensive experimental results demonstrate that the proposed MDMN offers more satisfactory and robust detection performance. Yunsong Li 0001, Kai Jiang 0001, Weiying Xie, Jie Lei 0001, Xin Zhang 0092, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | HTDFormer: Hyperspectral Target Detection Based on Transformer With Distributed LearningabstractIn recent years, many hyperspectral target detection (HTD) methods based on advanced techniques have been proposed and achieved good results. However, the large amount of data produced by satellites and airborne remote sensing instruments has posed new challenges for efficient target detection of massive hyperspectral images (HSIs). In this paper, we propose a new weakly supervised HTD framework based on a transformer with distributed learning (HTDFormer), which capitalizes on the parallel processing capabilities of multiple workers to efficiently handle large-scale HSIs. Specifically, the HTDFormer framework effectively integrates both spectral and spatial features within a unified optimization procedure via the transformer mechanism. A flexible sample augmentation approach is proposed to overcome the limitations of inadequate well-labeled training instances and meet the requirements of the transformer. To facilitate model training, we introduce the concept of distributed deep learning (DDL) into HTDFormer by leveraging a ring all-reduce (RAR) decentralized architecture, which embeds distributed learning into an HTD framework for the first time. Furthermore, the large-batch training strategy and the gradient compression strategy are employed to enable large-scale distributed processing and reduce communication costs, respectively. Finally, an exponentially constrained nonlinear function is adopted to acquire pixel-level prediction via spectral-spatial fusion. Experimental results demonstrate that the proposed framework achieves promising performance with regard to the increasing scale of real HSIs. Yunsong Li 0001, Weiying Xie |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Spatial and Spectral Structure Preserved Self-Representation for Unsupervised Hyperspectral Band SelectionabstractAs an effective manner to reduce data redundancy and processing inconvenience, hyperspectral band selection aims to select a subset of informative and discriminative bands from the original data cube. Although a large number of approaches have been proposed and obtained great success, they still face at least two issues. Firstly, most of the previous methods only consider the redundancy between neighbor bands, while the global information has been ignored. Secondly, each band is often treated as a whole and reshaped to a feature vector without considering the spatial structure of different regions. In this paper, in order to address these issues, we propose a spatial and spectral structure preserved self-representation model for unsupervised hyperspectral band selection without using any label information, referred to as S4P briefly. Different from previous methods that stretch each band into a feature vector, the first principal component of the original hyperspectral cube is segmented into different superpixels, which can reflect the spatial structure of homogeneous regions. Then each band can be represented by a superpixel level feature vector and the self-representation model is utilized to learn the spectral correlation of different bands. In addition, an adaptive and weighted multiple graph fusion term is designed to generate a unified similarity graph between different superpixels, which is used to capture the spatial structure in the self-representation space. Finally, anl2,1-norm is imposed on the self-representation coefficient matrix to measure the band importance. We design an alternative update scheme to optimize the resultant problem, the self-representation coefficient matrix and the superpixel-wise similarity graph can boost each other during the updating process to obtain optimal results. Extensive experiments with detailed analysis of three public datasets are conducted to validate the superiority of the proposed S4P when compared with other state-of-the-art competitors. Chang Tang, Jun Wang 0118, Xinwang Liu 0002, Weiying Xie, Xianju Li, Xinzhong Zhu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Co-Compression via Superior Gene for Remote Sensing Scene ClassificationabstractConvolutional neural networks (CNNs) have been successfully employed in remote sensing image classification because of their robust feature representation for different visual tasks and powerful graphics processing units (GPUs). The attendant problem is that high computational cost and high memory footprint hindering the application of CNNs for remote sensing applications in resource- and time-sensitive situations. Based on practical deployment requirements, we pioneer a pruning-quantization joint learning model compression method for remote sensing image classification, called co-compression via superior gene (CC-SG). An enhanced evolution algorithm (EEA) is adopted as the agent to search a “superior gene,” and immediately following, a director receives the “superior gene” and gives a compression mask and a resource constraint feedback to the agent. The network is eventually compressed and fine-tuned according to the optimal compression mask. Specifically, we introduce gene age and progressive shrinkage mutation rate to EEA and design a fitness function that balances accuracy and resource constraints. As validated using the UC Merced land-use and NWPU-RESISC45 datasets, the proposed CC-SG demonstrated superiority over other compared model compression approaches. For example, CC-SG attained substantial bit operations (BOPs) compression ratio of 40.04 with 0.956% accuracy increase for VGG-16 on UC Merced land-use dataset and 40.00 with 0.203% accuracy increase for ResNet-56 on NWPU-RESISC45 dataset. The code is available athttps://github.com/fanxxxxyi/CC-SG. Weiying Xie, Xiaoyi Fan 0002, Xin Zhang 0092, Yunsong Li 0001, Min Sheng, Leyuan Fang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | SuperYOLO: Super Resolution Assisted Object Detection in Multimodal Remote Sensing ImageryabstractAccurately and timely detecting multiscale small objects that contain tens of pixels from remote sensing images (RSI) remains challenging. Most of the existing solutions primarily design complex deep neural networks to learn strong feature representations for objects separated from the background, which often results in a heavy computation burden. In this article, we propose an accurate yet fast object detection method for RSI, named SuperYOLO, which fuses multimodal data and performs high-resolution (HR) object detection on multiscale objects by utilizing the assisted super resolution (SR) learning and considering both the detection accuracy and computation cost. First, we utilize a symmetric compact multimodal fusion (MF) to extract supplementary information from various data for improving small object detection in RSI. Furthermore, we design a simple and flexible SR branch to learn HR feature representations that can discriminate small objects from vast backgrounds with low-resolution (LR) input, thus further improving the detection accuracy. Moreover, to avoid introducing additional computation, the SR branch is discarded in the inference stage, and the computation of the network model is reduced due to the LR input. Experimental results show that, on the widely used VEDAI RS dataset, SuperYOLO achieves an accuracy of 75.09% (in terms of$\text {mA}{{\text {P}}_{{50}}}$), which is more than 10% higher than the SOTA large models, such as YOLOv5l, YOLOv5x, and RS designed YOLOrs. Meanwhile, the parameter size and GFLOPs of SuperYOLO are about$18\times $and$3.8\times $less than YOLOv5x. Our proposed model shows a favorable accuracy–speed tradeoff compared to the state-of-the-art models. The code will be open-sourced athttps://github.com/icey-zhang/SuperYOLO. Jie Lei 0001, Weiying Xie, Zhenman Fang, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Guided Hybrid Quantization for Object Detection in Remote Sensing Imagery via One-to-One Self-TeachingabstractDeep convolutional neural networks (CNNs) have improved remote sensing image analysis, but their high computational demands may limit their deployment on low-end devices with limited resources, such as intelligent satellites and unmanned aerial vehicles. Considering the computation complexity, we propose a Guided Hybrid Quantization with One-to-one Self-Teaching (GHOST) framework. More concretely, we first design a structure called guided quantization self-distillation (GQSD), an innovative idea for realizing a lightweight model through the synergy of quantization and distillation. The training process of the quantization model is guided by its full-precision model, which is time-saving and cost-saving without preparing a huge pre-trained model in advance. Second, we put forward a hybrid quantization (HQ) module that automatically acquires the optimal bit-width by imposing a threshold constraint on the distribution distance between the center point and samples in the weight search space, aiming to retain more shallow detail information that is advantageous for small object detection. Third, to improve information transformation, we propose a one-to-one self-teaching (OST) module to give the student network the ability to self-judgment. A switch control machine (SCM) builds a bridge between the student and teacher networks in the same location to help the teacher reduce wrong guidance and impart vital knowledge about objects without vast background information to the student. This distillation method allows a model to learn from itself and gain substantial improvement without any additional supervision. Extensive experiments on a multimodal dataset (VEDAI) and single-modality datasets (DOTA, NWPU, and DIOR) show that object detection based on GHOST outperforms the existing detectors. The tiny parameters (<9.7 MB) and Bit-Operations (BOPs) (<2158 G) compared with any remote sensing-based, lightweight, or distillation-based algorithms demonstrate the superiority in the lightweight design domain. Our code and model will be released at https://github.com/icey-zhang/GHOST. Jie Lei 0001, Weiying Xie, Yunsong Li 0001, Geng Yang 0001, Xiuping Jia |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | REAF: Remembering Enhancement and Entropy-Based Asymptotic Forgetting for Filter PruningabstractNeurologically, filter pruning is a procedure of forgetting and remembering recovering. Prevailing methods directly forget less important information from an unrobust baseline at first and expect to minimize the performance sacrifice. However, unsaturated base remembering imposes a ceiling on the slimmed model leading to suboptimal performance. And significantly forgetting at first would cause unrecoverable information loss. Here, we design a novel filter pruning paradigm termed Remembering Enhancement and Entropy-based Asymptotic Forgetting (REAF). Inspired by robustness theory, we first enhance remembering by over-parameterizing baseline with fusible compensatory convolutions which liberates pruned model from the bondage of baseline at no inference cost. Then the collateral implication between original and compensatory filters necessitates a bilateral-collaborated pruning criterion. Specifically, only when the filter has the largest intra-branch distance and its compensatory counterpart has the strongest remembering enhancement power, they are preserved. Further, Ebbinghaus curve-based asymptotic forgetting is proposed to protect the pruned model from unstable learning. The number of pruned filters is increasing asymptotically in the training procedure, which enables the remembering of pretrained weights gradually to be concentrated in the remaining filters. Extensive experiments demonstrate the superiority of REAF over many state-of-the-art (SOTA) methods. For example, REAF removes 47.55% FLOPs and 42.98% parameters of ResNet-50 only with 0.98% TOP-1 accuracy loss on ImageNet. The code is available at https://github.com/zhangxin-xd/REAF. Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Kai Jiang 0001, Leyuan Fang |
IEEE Trans. Image Process. | 2 |
| 2023 | Deep Hybrid 2-D-3-D CNN Based on Dual Second-Order Attention With Camera Spectral Sensitivity Prior for Spectral Super-ResolutionabstractA largely ignored fact in spectral super-resolution (SSR) is that the subsistent mapping methods neglect the auxiliary prior of camera spectral sensitivity (CSS) and only pay attention to wider or deeper network framework design while ignoring to excavate the spatial and spectral dependencies among intermediate layers, hence constraining representational capability of convolutional neural networks (CNNs). To conquer these drawbacks, we propose a novel deep hybrid 2-D-3-D CNN based on dual second-order attention with CSS prior (HSACS), which can excavate sufficient spatial-spectral context information. Specifically, dual second-order attention embedded in the residual block for more powerful spatial-spectral feature representation and relation learning is composed of a brand new trainable 2-D second-order channel attention (SCA) or 3-D second-order band attention (SBA) and a structure tensor attention (STA). Concretely, the band and channel attention modules are developed to adaptively recalibrate the band-wise and interchannel features via employing second-order band or channel feature statistics for more discriminative representations. Besides, the STA is promoted to rebuild the significant high-frequency spatial details for enough spatial feature extraction. Moreover, the CSS is first employed as a superior prior to avoid its effect of SSR quality, on the strength of which the resolved RGB can be calculated naturally through the super-reconstructed hyperspectral image (HSI); then, the final loss consists of the discrepancies of RGB and the HSI as a finer constraint. Experimental results demonstrate the superiority and progressiveness of the presented approach in terms of quantitative metrics and visual effect over SOTA SSR methods. Jiaojiao Li 0001, Chaoxiong Wu, Rui Song 0003, Yunsong Li 0001, Weiying Xie, Lihuo He, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Transcoded Video Restoration by Temporal Spatial Auxiliary NetworkabstractIn most video platforms, such as Youtube, Kwai, and TikTok, the played videos usually have undergone multiple video encodings such as hardware encoding by recording devices, software encoding by video editing apps, and single/multiple video transcoding by video application servers. Previous works in compressed video restoration typically assume the compression artifacts are caused by one-time encoding. Thus, the derived solution usually does not work very well in practice. In this paper, we propose a new method, temporal spatial auxiliary network (TSAN), for transcoded video restoration. Our method considers the unique traits between video encoding and transcoding, and we consider the initial shallow encoded videos as the intermediate labels to assist the network to conduct self-supervised attention training. In addition, we employ adjacent multi-frame information and propose the temporal deformable alignment and pyramidal spatial fusion for transcoded video restoration. The experimental results demonstrate that the performance of the proposed method is superior to that of the previous techniques. The code is available at https://github.com/icecherylXuli/TSAN. Li Xu 0008, Gang He 0002, Jinjia Zhou, Jie Lei 0001, Weiying Xie, Yunsong Li 0001, Yu-Wing Tai |
AAAI | 5 |
| 2022 | HTD-VIT: Spectral-Spatial Joint Hyperspectral Target Detection with Vision TransformerabstractIn hyperspectral images (HSIs), spatial context provides complementary information to abundant spectral features. In this paper, a united spectral-spatial framework named HTD-ViT based on vision transformer (ViT) is proposed for HTD tasks. The HTD-ViT leverages the ViT to learn discriminative spectral-spatial features of each pixel and its neighboring pixels. Meanwhile, the spectral-spatial sequence construction operation uses spectrums in the cross region centered on the selected pixel to produce the corresponding spectral-spatial sequence for ViT processing. Furthermore, the spectral-spatial sample selection procedure based on coarse detection addresses the issue of lacking well-labeled training instances in the HTD tasks. Finally, the spectral-spatial pixel-level detection combines the discriminative feature from the spectral and the spatial domains to suppress the background. In contrast to traditional spatial-spectral feature extraction methods that stack the original spectral feature with spatial neighborhood information directly, joint spectral-spatial inference in HTD-ViT can effectively discover the underlying contextual and structure information in HSIs. Experiments on real HSIs verify the effectiveness of HTD-ViT, which takes full advantage of both the variable spectral and spatial features. Weiying Xie, Yunsong Li 0001, Qian Du 0001 |
IGARSS | 2 |
| 2022 | Interlayer Restoration Deep Neural Network for Scalable High Efficiency Video CodingabstractThis paper applies an interlayer restoration deep neural network (IRDNN) for scalable high efficiency video coding (SHVC) to improve visual quality and coding efficiency. It is the first time to combine deep neural network (DNN) and SHVC. Considering the coding architecture of SHVC, we elaborate a multi-frame and multi-layer neural network to restore the interlayer of SHVC by utilizing both the adjacent reconstructed frames of the base layer (BL) and enhancement layer (EL). Moreover, we analyze the temporal motion relationship of frames in one layer and the compression degradation relationship of frames between different layers, and propose the synergistic mechanism of motion restoration and compression restoration in our IRDNN. The network can generate an interlayer with higher quality serving for the EL coding and thus enhance the coding efficiency. A large-scale and various-quality-degradation dataset is self-made for the task of interlayer restoration of SHVC. The experimental results show that with our implementation on SHVC, the EL Bj$\phi $ntegaard delta bit-rate (BD-BR) reduction is 9.291% and 6.007% in signal-to-noise ratio scalability and spatial scalability, respectively. The code is available athttps://github.com/icecherylXuli/IRDNN. Gang He 0002, Li Xu 0008, Jie Lei 0001, Weiying Xie, Yunsong Li 0001, Yibo Fan, Jinjia Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | E2E-LIADE: End-to-End Local Invariant Autoencoding Density Estimation Model for Anomaly Target Detection in Hyperspectral ImageabstractHyperspectral anomaly target detection (also known as hyperspectral anomaly detection (HAD)] is a technique aiming to identify samples with atypical spectra. Although some density estimation-based methods have been developed, they may suffer from two issues: 1) separated two-stage optimization with inconsistent objective functions makes the representation learning model fail to dig out characterization customized for HAD and 2) incapability of learning a low-dimensional representation that preserves the inherent information from the original high-dimensional spectral space. To address these problems, we propose a novel end-to-end local invariant autoencoding density estimation (E2E-LIADE) model. To satisfy the assumption on the manifold, the E2E-LIADE introduces a local invariant autoencoder (LIA) to capture the intrinsic low-dimensional manifold embedded in the original space. Augmented low-dimensional representation (ALDR) can be generated by concatenating the local invariant constrained by a graph regularizer and the reconstruction error. In particular, an end-to-end (E2E) multidistance measure, including mean-squared error (MSE) and orthogonal projection divergence (OPD), is imposed on the LIA with respect to hyperspectral data. More important, E2E-LIADE simultaneously optimizes the ALDR of the LIA and a density estimation network in an E2E manner to avoid the model being trapped in a local optimum, resulting in an energy map in which each pixel represents a negative log likelihood for the spectrum. Finally, a postprocessing procedure is conducted on the energy map to suppress the background. The experimental results demonstrate that compared to the state of the art, the proposed E2E-LIADE offers more satisfactory performance. Kai Jiang 0001, Weiying Xie, Jie Lei 0001, Zan Li 0001, Yunsong Li 0001, Tao Jiang 0031, Qian Du 0001 |
IEEE Trans. Cybern. | 2 |
| 2022 | Fusion of Hyperspectral and Panchromatic Images Using Generative Adversarial Network and Image SegmentationabstractHyperspectral (HS) image fusion aims at integrating a panchromatic (PAN) image and an HS image, featuring the fused image with the spatial quality of the former and the spectral diversity of the latter. The classic fusion algorithm generally includes three consecutive procedures that are upsampling, detail extraction, and detail injection. In this article, we propose an HS and PAN image fusion method based on generative adversarial network and local estimation of injection gain. Instead of upsampling the HS image by classical interpolation techniques, a generative adversarial super-resolution network (GASN) is designed to obtain the interpolated HS image in the fusion framework. GASN establishes a spectral-information-based discriminator to conduct adversarial learning with the generator, so as to preserve the spectral information of the low-resolution HS image. An image segmentation-based injection gain estimation (ISGE) algorithm is subsequently proposed for HS and PAN images fusion. The injection gain is estimated over image segments obtained by a binary partition tree approach to improve the fusion performance. The proposed GASN and ISGE are implemented into two credible global estimation pansharpening methods, and experimental results prove the performance improvement of the proposed method. The proposed method is also compared with existing state-of-the-art methods, and experiments on several public databases demonstrate that the proposed method is competitive or superior to the state-of-the-art fusion methods. Wenqian Dong, Jiahui Qu, Weiying Xie, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Boundary Extraction Constrained Siamese Network for Remote Sensing Image Change DetectionabstractChange detection (CD) is crucial to the understanding of relationships and interactions among multitemporal high-resolution remote sensing (RS) images. However, various inherent attributes of images have different impacts on CD judgment. How to effectively use helpful information to improve the performance of CD is still a challenge. In this article, we present a boundary extraction constrained Siamese network (BESNet) to dig out the efficacy of boundary information. BESNet is a joint learning network in which a novel multiscale boundary extraction (MSBE) module is embedded. In this way, traditional and deep learning techniques are leveraged to learn together to maximize their respective strengths through cooperation. In particular, a new boundary extraction constrained (BEC) loss function combined with a contractive loss function is used to optimize the BESNet. Considering the interaction between various extracted features, a channel-shuffle fusion strategy is developed to exploit their complementary advantages between features. Our experiments show that the proposed BESNet can significantly improve the CD performance and generate more complete and clearer object boundaries. Experiments conducted on two real datasets over different scenes demonstrate its state-of-the-art performance. Jie Lei 0001, Yijie Gu, Weiying Xie, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Sparse Coding-Inspired GAN for Hyperspectral Anomaly Detection in Weakly Supervised LearningabstractAnomaly detection (AD) from hyperspectral images (HSIs) is of great importance in both space exploration and Earth observations. However, the challenges caused by insufficient datasets, no labels, and noise corruption substantially downgrade the accuracy of detection. To solve these problems, this article proposes a sparse coding (SC)-inspired generative adversarial network (GAN) for weakly supervised hyperspectral AD (HAD), named sparseHAD. It can learn a discriminative latent reconstruction with small errors for background pixels and large errors for anomalous ones. First, a background-category searching step is built to alleviate the difficulty of data annotation. Then, an SC-inspired regularized network is integrated into an end-to-end GAN to form a weakly supervised spectral mapping model consisting of two encoders, a decoder, and a discriminator. This model not only makes the network more robust and interpretable experimentally and theoretically but also develops a new SC-inspired path for HAD. Subsequently, the proposed sparseHAD detects anomalies in a latent space rather than the original space, which also contributes to its noise robustness. Quantitative assessments and experiments over real HSIs demonstrate the unique promise of the proposed sparseHAD. The code, data, and trained models are available athttps://github.com/JiangThea/HAD. Yunsong Li 0001, Tao Jiang 0031, Weiying Xie, Jie Lei 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Structure-Guided Feature Transform Hybrid Residual Network for Remote Sensing Object DetectionabstractObject detection in remote sensing imagery (RSI) is a fundamental task for Earth monitoring. Objects captured from the bird’s eye view perspective in RSI can appear as multiscale in arbitrary orientations, most of which are small and dense. In specific, vehicles or ships only occupy a dozen pixels in the image, but are surrounded by roads and seas, which occupy thousands of pixels and comprise overwhelmingly dominant of all pixels. Although a large number of common object detection methods have been proposed, most of them cannot detect small and dense objects accurately because none of them has paid enough attention to the unique characteristic of RSI. In this work, we propose a novel structure-guided feature transform hybrid residual (SGFTHR) network, which can conquer the low performance of detection of objects at different scales, especially for small and dense objects, in an anchor-free manner. The structure-guided feature transform (SGFT) module is promoted to extract discriminative structural information and guide this information into high-level contextual feature maps, preventing the important low-level spatial and structural information from being lost when the network goes deeper. Furthermore, the hybrid residual (HR) module is embedded in the backbone to acquire multiscale features in a novel hybrid hierarchical residual-like manner. Extensive experiments are performed on the HRRSD and NWPU VHR-10 datasets to evaluate the performance of the SGFTHR network, which demonstrates that our SGFTHR network achieves state-of-the-art detection accuracy with high efficiency and robustness. Specifically, 4.12% improvements in mean average precision (mAP) on the HRRSD dataset compared with baseline powerfully demonstrate the effectiveness and superiority of the SGFTHR network. Jiaojiao Li 0001, Huanqing Zhang, Rui Song 0003, Weiying Xie, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Dual-Frequency Autoencoder for Anomaly Detection in Transformed Hyperspectral ImageryabstractHyperspectral anomaly detection (HAD) is a challenging task since samples are unavailable for training. Although unsupervised learning methods have been developed, they often train the model using an original hyperspectral image (HSI) and require retraining on different HSIs, which may limit the feasibility of HAD methods in practical applications. To tackle this problem, we propose a dual-frequency autoencoder (DFAE) detection model in which the original HSI is transformed into high-frequency components (HFCs) and low-frequency components (LFCs) before detection. A novel spectral rectification is first proposed to alleviate the spectral variation problem and generate the LFCs of HSI. Meanwhile, the HFCs are extracted by the Laplacian operator. Subsequently, the proposed DFAE model is learned to detect anomalies from the LFCs and HFCs in parallel. Finally, the learned model is well-generalized for anomaly detection from other hyperspectral datasets. While breaking the dilemma of limited generalization in the sample-free HAD task, the proposed DFAE can enhance the background–anomaly separability, providing a better performance gain. Experiments on real datasets demonstrate that the DFAE method exhibits competitive performance compared with other advanced HAD methods. Yidan Liu, Weiying Xie, Yunsong Li 0001, Zan Li 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Multilevel Encoder-Decoder Attention Network for Change Detection in Hyperspectral ImagesabstractConvolutional neural networks (CNNs) have attracted much attention in change detection (CD) for their superior feature learning ability. However, most of the existing CNN-based CD methods adopt an early- or late-fusion strategy to fuse low-level spatial details or high-level semantic information. So far, the impact of multilevel fusion strategy across multitemporal hyperspectral (HS) images, and its application to CD, remains unexplored. In this article, we propose a multilevel encoder–decoder attention network (ML-EDAN), which allows the network to make full use of the hierarchical features for CD in HS images. A two-stream encoder–decoder framework is taken as the backbone to exploit and fuse the hierarchical features from all the convolutional layers of multitemporal HS images. Within the encoder–decoder, a contextual-information-guided attention module is developed to yield more effective spatial–spectral feature transfer in the network. After fully obtaining the multilevel hierarchical features, the long short-term memory (LSTM) subnetwork is devised to analyze temporal dependence between multitemporal images. Moreover, the proposed ML-EDAN is trained in an end-to-end manner with a new joint loss function considering both reconstruction error and pixelwise classification error. The experiments are conducted on three datasets, demonstrating the effectiveness of the proposed ML-EDAN in HS CD in comparison with widely accepted state-of-the-art methods. Jiahui Qu, Shaoxiong Hou, Wenqian Dong, Yunsong Li 0001, Weiying Xie |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | MSSL: Hyperspectral and Panchromatic Images Fusion via Multiresolution Spatial-Spectral Feature Learning NetworksabstractThe fusion of hyperspectral (HS) and panchromatic (PAN) images aims to generate a fused HS image that combines spectral information of the HS image with spatial information of the PAN image. In this article, we propose a multiresolution spatial–spectral feature learning (MSSL) framework for fusing HS and PAN images. The proposed MSSL transforms the existing deep and complex network into several simple and shallow subnetworks to simplify the feature learning process. MSSL upsamples the HS image while downsamples the PAN image and designs multiresolution 3-D convolutional autoencoder (CAEs) networks with a spectral constraint to learn complete spatial–spectral features of the HS image. MSSL designs multiresolution 2-D CAEs with spatial constraint to extract spatial features of the PAN image, with a low computational cost. In order to effectively generate the pansharpened HS image with high spatial and spectral fidelity, a multiresolution residual network is presented to reconstruct the HS image from the extracted spatial–spectral features. Extensive experiments are conducted on three widely used remote sensing data sets in comparison with state-of-the-art HS image fusion methods, demonstrating the superiority of the proposed MSSL method. Code is available athttps://github.com/Jiahuiqu/MSSL. Jiahui Qu, Yanzi Shi, Weiying Xie, Yunsong Li 0001, Xianyun Wu, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Spectral Distribution-Aware Estimation Network for Hyperspectral Anomaly DetectionabstractRecently developed deep learning-based hyperspectral anomaly detection (HAD) methods typically include two steps where the deep feature extraction is not designed specifically for the HAD task. In this article, we propose a spectral distribution-aware estimation network (SDEN) that does not conduct feature extraction and anomaly detection in two separate steps but instead learns both jointly to estimate anomalies directly in an end-to-end manner without postprocessing. The unified framework can ensure that the extracted features serve better for anomaly detection. To preserve the distribution of hyperspectral images (HSIs) during dimensionality reduction, the SDEN introduces a spectral distribution (SD)-aware module imposed with a local-invariant constraint. More specifically, we adopt Markov chain Monte Carlo (MCMC) that enables the SD module to better estimate the distribution of the complex HSIs. Considering the powerful representation capability of Gaussian mixture model (GMM), the SDEN leverages it to establish an estimation module in the deep latent space where the anomaly resides in low density while the background not. We demonstrate that the SDEN yields competitive and highly promising results in comparison with the anomaly detection benchmarks. Weiying Xie, Shuran Fan, Jiahui Qu, Xianyun Wu, Yanli Lu, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Algorithm/Hardware Codesign for Real-Time On-Satellite CNN-Based Ship Detection in SAR ImageryabstractRecently, the convolutional neural network (CNN)-based approach for on-satellite ship detection in synthetic aperture radar (SAR) images has received increasing attention since it does not rely on predefined imagery features and distributions that are required in conventional detection methods. To achieve high detection accuracy, most of the existing CNN-based methods leverage complex off-the-shelf CNN models for optical imagery. Unfortunately, this usually leads to expensive computational cost, which is hard to process in real time using resource-constrained devices deployed in the harsh satellite environment. In this article, we propose OSCAR-RT, the first end-to-end algorithm/hardware codesign framework for real-time on-satellite CNN-based SAR ship detection, which can simultaneously produce an accurate and hardware-friendly CNN model and an ultraefficient field-programmable gate array (FPGA)-based hardware accelerator that can be deployed on satellites. With the real-time on-satellite processing speed in mind, we start from a state-of-the-art compact CNN model for optical imagery. To eliminate the sharp decrease in the detection accuracy for SAR imagery, we analyze the discrepancy between the SAR domain and optical domain and propose to adapt the model by adjusting the output feature size to better detect relatively smaller objects in SAR imagery. To improve the detection speed, we propose to develop a fully pipelined interlayer streaming accelerator architecture, where all the layers of the CNN model can be concurrently processed using on-chip FPGA resources. To achieve this architecture, we first propose a hardware-guided, progressive, and structural pruning strategy, which is guided by our modeled hardware metrics and applies state-of-the-art coarse-grained and fine-grained filter pruning as well as mixed-precision quantization techniques. Moreover, to improve the reusability and portability of the hardware accelerator design, we develop a library of highly optimized CNN components in high-level synthesis, together with their performance and resource models. Finally, we map the pruned CNN model onto these hardware library components in a fully pipelined interlayer streaming fashion, by adjusting their parallelism factors to balance the execution of each layer and fit into the resource constraint. Experimental results using the adapted MobileNetV1, MobileNetV2, and SqueezeNet models on the widely used SAR ship detection dataset (SSDD) demonstrate the effectiveness of OSCAR-RT; for the MobileNetV1 model, it achieves an average precision of 94%, a detection speed of 652 frames/s on the Xilinx VC709 FPGA evaluation board while consuming about 5.8-W power. Geng Yang 0001, Jie Lei 0001, Weiying Xie, Zhenman Fang, Yunsong Li 0001, Xin Zhang 0092 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Rank-Aware Generative Adversarial Network for Hyperspectral Band SelectionabstractTraditional clustering-based band selection (BS) methods treat each band as individuals, and selection is conducted by enlarging the difference between clusters, which leads to the loss of band interaction and information saliency evaluation. In this article, we propose a BS method named rank-aware generative adversarial network (R-GAN) to address these problems. First, centralized reference feature extraction (FE) with GAN aids R-GAN to combine interpretability and interband relevance. Then, the reference feature is refined with the saliency estimation provided by the rank-aware strategy. According to data characteristics, there are two versions of rank computation including tensor and matrix. Finally, the structural similarity index measurement (SSIM) maps the saliency to the original data space to obtain the final BS result. Extensive comparison experiments with popular existing BS approaches on five hyperspectral images (HSIs) datasets show that the proposed R-GAN can address spectral saliency effectively and select more informative band subsets, which outperforms other competitors for both detection and classification tasks. For example, on the SD-1 dataset, the ten bands selected by R-GAN achieve 0.982 ± 0.003 with an improvement of 13.7% in the area under the curve (AUC) value of anomaly detection performance. The peaked accuracy surpasses the baseline by 0.46% for the classification on the PaviaU dataset. Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001, Geng Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Weakly Supervised Discriminative Learning With Spectral Constrained Generative Adversarial Network for Hyperspectral Anomaly DetectionabstractAnomaly detection (AD) using hyperspectral images (HSIs) is of great interest for deep space exploration and Earth observations. This article proposes a weakly supervised discriminative learning with a spectral constrained generative adversarial network (GAN) for hyperspectral anomaly detection (HAD), called weaklyAD. It can enhance the discrimination between anomaly and background with background homogenization and anomaly saliency in cases where anomalous samples are limited and sensitive to the background. A novel probability-based category thresholding is first proposed to label coarse samples in preparation for weakly supervised learning. Subsequently, a discriminative reconstruction model is learned by the proposed network in a weakly supervised fashion. The proposed network has an end-to-end architecture, which not only includes an encoder, a decoder, a latent layer discriminator, and a spectral discriminator competitively but also contains a novel Kullback-Leibler (KL) divergence-based orthogonal projection divergence (OPD) spectral constraint. Finally, the well-learned network is used to reconstruct HSIs captured by the same sensor. Our work paves a new weakly supervised way for HAD, which intends to match the performance of supervised methods without the prerequisite of manually labeled data. Assessments and generalization experiments over real HSIs demonstrate the unique promise of such a proposed approach. Tao Jiang 0031, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | LREN: Low-Rank Embedded Network for Sample-Free Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) is a challenging task because it explores the intrinsic structure of complex high-dimensional signals without any samples at training time. Deep neural networks (DNNs) can dig out the underlying distribution of hyperspectral data but are limited by the labeling of large-scale hyperspectral datasets, especially the low spatial resolution of hyperspectral data, which makes labeling more difficult. To tackle this problem while ensuring the detection performance, we present an unsupervised low-rank embedded network (LREN) in this paper. LREN is a joint learning network in which the latent representation is specifically designed for HAD, rather than merely as a feature input for the detector. And it searches the lowest rank representation based on a representative and discriminative dictionary in the deep latent space to estimate the residual efficiently. Considering the physically mixing properties in hyperspectral imaging, we develop a trainable density estimation module based on Gaussian mixture model (GMM) in the deep latent space to construct a dictionary that can better characterize the complex hyperspectral images (HSIs). The closed-form solution of the proposed low-rank learner surpasses existing approaches on four real hyperspectral datasets with different anomalies. We argue that this unified framework paves a novel way to combine feature extraction and anomaly estimation-based methods for HAD, which intends to learn the underlying representation tailored for HAD without the prerequisite of manually labeled data. Code available at https://github.com/xdjiangkai/LREN. Kai Jiang 0001, Weiying Xie, Jie Lei 0001, Tao Jiang 0031, Yunsong Li 0001 |
AAAI | 2 |
| 2021 | PTGAN: A Proposal-Weighted Two-Stage GAN with Attention for Hyperspectral Target DetectionabstractIn this paper, a proposal-weighted two-stage generative adversarial network (GAN) with attention mechanism is proposed for hyperspectral target detection (HTD). PTGAN leverages GAN to estimate spectral background distribution and realize mapping from the latent space to the spectral space. Meanwhile, PTGAN conducts the reversed mapping through latent-spectral-latent and spectral-latent-spectral learning. On this basis, PTGAN implements accurate reconstruction of background spectrum via latent space. Therefore, targets of interest can be detected through larger pixel-level reconstruction error. In particular, the variance attention module is designed to make full use of global information among spectral bands to selectively emphasize channel-wise spectral features. Furthermore, a proposal-weighted strategy in a two-stage manner reduces the false alarm of detection by refining the previous detection proposal. Finally, exponential nonlinear fusion combines the discriminative feature from two stages to suppress the background. Extensive experiments on two real hyperspectral images (HSIs) verify the effectiveness of PTGAN. Weiying Xie, Yunsong Li 0001, Kai Jiang 0001, Jie Lei 0001, Qian Du 0001 |
IGARSS | 2 |
| 2021 | Spectral mapping with adversarial learning for unsupervised hyperspectral change detection
Jie Lei 0001, Meiqi Li, Weiying Xie, Yunsong Li 0001, Xiuping Jia |
Neurocomputing | 3 |
| 2021 | A Specially Optimized One-Stage Network for Object Detection in Remote Sensing ImagesabstractWith great significance in military and civilian applications, detecting indistinguishable small objects in wide-scale remote sensing images is still a challenging topic. In this letter, we propose a specially optimized one-stage network (SOON) focusing on extracting spatial information of high-resolution images by understanding and analyzing the combination of feature and semantic information of small objects. The SOON model consists of feature enhancement, multiscale detection, and feature fusion. The first part is implemented by constructing a receptive field enhancement (RFE) module and incorporating it into the network's specific parts where the information of small objects mainly exists. The second part is achieved by four detectors with different sensitivities, which access to the fused and enhanced features to enable the network to make full use of features in different scales. The third part consolidates the high-level and low-level features by adopting upsampling, concatenation, and convolution operations to build a feature pyramid structure, which explicitly yields strong feature representation and semantic information. In addition, we introduce the soft-nonmaximum suppression to preserve accurate bounding boxes in the postprocessing stage for densely arranged objects. Note that the split and merge strategy and the multiscale training strategy are employed. Extensive experiments and thorough analysis are performed on the NorthWestern Polytechnical University Very-High-Resolution (NWPU VHR)-10-v2 data set and the airplane, car and ship (ACS) data set as compared with several state-of-the-art methods. The satisfactory performance in experiments verifies the effectiveness of the design and optimization. Yunsong Li 0001, Jie Lei 0001, Weiying Xie |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2021 | Self-spectral learning with GAN based spectral-spatial target detection for hyperspectral image
Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Xiuping Jia |
Neural Networks | 1 |
| 2021 | Parallel and Distributed Computing for Anomaly Detection From Hyperspectral Remote Sensing ImageryabstractAnomaly detection from remote sensing images is to detect pixels whose spectral signatures are different from their background. Anomalies are often man-made targets. With such target signatures being unknown, anomaly detection has many important applications, such as water quality monitoring, crop stress surveying, and law enforcement-related uses, where prior information of targets is often unavailable. The key to success is accurate background modeling. Anomaly detection from remote sensing images is challenging because spatial coverage is very large and the background is highly heterogeneous. For pixel-based anomaly detection, computing cost in background modeling and a spatial-convolution-type detection process is very expensive. Thus, parallel and distributed computing is critical in reducing execution time, which can fit the need for real-time or near real-time detection from airborne and spaceborne platforms in support of immediate decision-making. This article reviews the recent advances in anomaly detection from hyperspectral remote sensing images and their implementation using parallel and distributed systems. The classical methods, i.e., the Reed-Xiaoli (RX) algorithm and its variants, including its real-time processing version, are illustrated in commodity graphic processing units (GPUs), cloud, and field-programmable gate array (FPGA) implementations. Practical issues and future development trends are also discussed. Qian Du 0001, Bo Tang 0011, Weiying Xie, Wei Li 0032 |
Proc. IEEE | 3 |
| 2021 | Dual feature extraction network for hyperspectral image analysis
Weiying Xie, Jie Lei 0001, Shuo Fang, Yunsong Li 0001, Xiuping Jia, Mingsuo Li |
Pattern Recognit. | 1 |
| 2021 | Weakly Supervised Low-Rank Representation for Hyperspectral Anomaly DetectionabstractIn this article, we propose a weakly supervised low-rank representation (WSLRR) method for hyperspectral anomaly detection (HAD), which formulates deep learning-based HAD into a low-lank optimization problem not only characterizing the complex and diverse background in real HSIs but also obtaining relatively strong supervision information. Different from the existing unsupervised and supervised methods, we first model the background in a weakly supervised manner, which achieves better performance without prior information and is not restrained by richly correct annotation. Considering reconstruction biases introduced by the weakly supervised estimation, LRR is an effective method for further exploring the intricate background structures. Instead of directly applying the conventional LRR approaches, a dictionary-based LRR, including both observed training data and hidden learned data drawn by the background estimation model, is proposed. Finally, the derived low-rank part and sparse part and the result of the initial detection work together to achieve anomaly detection. Comparative analyses validate that the proposed WSLRR method presents superior detection performance compared with the state-of-the-art methods. Weiying Xie, Xin Zhang 0092, Yunsong Li 0001, Jie Lei 0001, Jiaojiao Li 0001, Qian Du 0001 |
IEEE Trans. Cybern. | 1 |
| 2021 | Hybrid 2-D-3-D Deep Residual Attentional Network With Structure Tensor Constraints for Spectral Super-Resolution of RGB ImagesabstractRGB image spectral super-resolution (SSR) is a challenging task due to its serious ill-posedness, which aims at recovering a hyperspectral image (HSI) from a corresponding RGB image. In this article, we propose a novel hybrid 2-D-3-D deep residual attentional network (HDRAN) with structure tensor constraints, which can take fully advantage of the spatial-spectral context information in the reconstruction progress. Previous works improve the SSR performance only through stacking more layers to catch local spatial correlation neglecting the differences and interdependences among features, especially band features; different from them, our novel method focuses on the context information utilization. First, the proposed HDRAN consists of a 2D-RAN following by a 3D-RAN, where the 2D-RAN mainly focuses on extracting abundant spatial features, whereas the 3D-RAN mainly simulates the interband correlations. Then, we introduce 2-D channel attention and 3-D band attention mechanisms into the 2D-RAN and 3D-RAN, respectively, to adaptively recalibrate channelwise and bandwise feature responses for enhancing context features. Besides, since structure tensor represents structure and spatial information, we apply structure tensor constraint to further reconstruct more accurate high-frequency details during the training process. Experimental results demonstrate that our proposed method achieves the state-of-the-art performance in terms of mean relative absolute error (MRAE) and root mean square error (RMSE) on both the “clean” and “real world” tracks in the NTIRE 2018 Spectral Reconstruction Challenge. As for competitive ranking metric MRAE, our method separately achieves a 16.06% and 2.90% relative reduction on two tracks over the first place. Furthermore, we investigate HDRAN on the other two HSI benchmarks noted as the CAVE and Harvard data sets, also demonstrating better results than state-of-the-art methods. Jiaojiao Li 0001, Chaoxiong Wu, Rui Song 0003, Weiying Xie, Chiru Ge, Bo Li 0090, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | HPGAN: Hyperspectral Pansharpening Using 3-D Generative Adversarial NetworksabstractHyperspectral (HS) pansharpening, as a special case of the superresolution (SR) problem, is to obtain a high-resolution (HR) image from the fusion of an HR panchromatic (PAN) image and a low-resolution (LR) HS image. Though HS pansharpening based on deep learning has gained rapid development in recent years, it is still a challenging task because of the following requirements: 1) a unique model with the goal of fusing two images with different dimensions should enhance spatial resolution while preserving spectral information; 2) all the parameters should be adaptively trained without manual adjustment; and 3) a model with good generalization should overcome the sensitivity to different sensor data in reasonable computational complexity. To meet such requirements, we propose a unique HS pansharpening framework based on a 3-D generative adversarial network (HPGAN) in this article. The HPGAN induces the 3-D spectral-spatial generator network to reconstruct the HR HS image from the newly constructed 3-D PAN cube and the LR HS image. It searches for an optimal HR HS image by successive adversarial learning to fool the introduced PAN discriminator network. The loss function is specifically designed to comprehensively consider global constraint, spectral constraint, and spatial constraint. Besides, the proposed 3-D training in the high-frequency domain reduces the sensitivity to different sensor data and extends the generalization of HPGAN. Experimental results on data sets captured by different sensors illustrate that the proposed method can successfully enhance spatial resolution and preserve spectral information. Weiying Xie, Yuhang Cui, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001, Jiaojiao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Characterization of Background-Anomaly Separability With Generative Adversarial Network for Hyperspectral Anomaly DetectionabstractHyperspectral images (HSIs) have unique advantages in distinguishing subtle spectral differences of different materials. However, due to complex and diverse backgrounds, unknown prior knowledge, and imbalanced samples, it is challenging to separate background and anomaly. In this article, we present a novel characterization of background-anomaly separability with a generative adversarial network (BASGAN) for hyperspectral anomaly detection. The key contribution is the proposal to explicitly constrain the background and anomaly separability by characterizing background spectral samples while avoiding anomaly reconstruction. First, we use a class saliency map extraction algorithm to obtain pseudobackground and anomaly samples for adversarial training. To further mitigate the suffering of anomaly contamination in background distribution estimation, we introduce background-anomaly separability constrained loss function to enhance the reconstruction of the background while weakening the anomaly reconstruction in a semisupervised way. Additionally, a discriminator is induced into the latent space to make the encoded representation resemble Gaussian distribution during adversarial training. The other is adversarial training in the reconstruction space so that the background estimation can be improved. Experiments conducted on real data sets illustrate the superior background-anomaly separability of the proposed method. Jiaping Zhong, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Discriminative Semi-Supervised Generative Adversarial Network for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection has been facing great challenges in the field of deep learning due to high dimensions and limited samples. To address these challenges, a novel discriminative semi-supervised generative adversarial network (GAN) method with dual RX (Reed-Xiaoli), called semiDRX, is proposed in this paper. The main contribution of the proposed method is to learn a reconstruction of background homogenization and anomaly saliency through a semi-supervised GAN. To achieve this goal, firstly, the coarse RX detection is performed to obtain a background sample set with potential anomalous pixels being removed. Secondly, the obtained coarse background set learns more comprehensive background characteristics through the network. The original hyperspectral image (HSI) is fed into the learned network to obtain reconstructions with homogeneous backgrounds and salient anomalies. The refined detection results are generated by a second RX detector. Experiments on three HSIs over different scenes demonstrate its advancement and effectiveness. Tao Jiang 0031, Weiying Xie, Yunsong Li 0001, Qian Du 0001 |
IGARSS | 2 |
| 2020 | Unsupervised spectral mapping and feature selection for hyperspectral anomaly detection
Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Jiaojiao Li 0001, Xiuping Jia |
Neural Networks | 1 |
| 2020 | Discriminative Reconstruction Constrained Generative Adversarial Network for Hyperspectral Anomaly DetectionabstractThe rich and distinguishable spectral information in hyperspectral images (HSIs) makes it possible to capture anomalous samples [i.e., anomaly detection (AD)] that deviate from background samples. However, hyperspectral anomaly detection (HAD) faces various challenges due to high dimensionality, redundant information, and unlabeled and limited samples. To address these problems, this article proposes an unsupervised discriminative reconstruction constrained generative adversarial network for HAD (HADGAN). Our solution is mainly based on the assumption that the number of normal samples is much larger than the number of abnormal ones. The key contribution of this article is to learn a discriminative background reconstruction with anomaly targets being suppressed, which produces the initial detection image (i.e., the residual image between the original image and reconstructed image) with anomaly targets being highlighted and background samples being suppressed. To accomplish this goal, first, by using an autoencoder (AE) network and an adversarial latent discriminator, the latent feature layer learns normal background distribution and AE learns a background reconstruction as much as possible. Second, consistency enhanced representation and shrink constraints are added to the latent feature layer to ensure that anomaly samples are projected to similar positions as normal samples in the latent feature layer. Third, using an adversarial image feature corrector in the input space can guarantee the reliability of the generated samples. Finally, an energy-based spatial and distance-based spectral joint anomaly detector is applied in the residual map to generate the final detection map. Experiments conducted on several data sets over different scenes demonstrate its state-of-the-art performance. Tao Jiang 0031, Yunsong Li 0001, Weiying Xie, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Semisupervised Spectral Learning With Generative Adversarial Network for Hyperspectral Anomaly DetectionabstractLimited by the anomalous spectral vectors in unlabeled hyperspectral images (HSIs), anomaly detection methods based on background distribution estimation often suffer from the contamination of anomalies, which decreases the estimation accuracy and, thus, weakens the detection performance. To address this problem, we proposed a novel semisupervised spectral learning (SSL) for the hyperspectral anomaly detection framework based on the generative adversarial network (GAN). GAN is applied and developed to estimate the background distribution in a semisupervised manner and obtain an initial spectral feature because of its strong representational capability and adversarial training advantage. In the proposed framework, an initial spatial feature is generated via morphological attribute filtering. Finally, an exponential constrained nonlinear suppression fusion technique is adopted to suppress the background and combine the complementary information in different features to obtain a fused detection map. The performance of the proposed anomaly detection technique is evaluated on a series of HSIs. Experimental results demonstrate that our method can outperform state-of-the-art anomaly detection methods. Kai Jiang 0001, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Gang He 0002, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Discriminative Reconstruction for Hyperspectral Anomaly Detection With Spectral LearningabstractRecently, autoencoder (AE)-based anomaly detection has drawn considerable interest in hyperspectral image (HSI) analysis. In this article, we propose a novel discriminative reconstruction method for hyperspectral anomaly detection images with spectral learning (SLDR). The proposed algorithm has the following innovations. First, we use the spectral error map (SEM) to detect anomalies because the SEM can preferably reflect the spectral similarity of each pixel between the input and the reconstruction. Second, the loss function of the proposed SLDR model additionally introduces the spectral angle distance (SAD), which constrains the model to generate a reconstruction having greater spectral similarity to the input. Third, a constraint is imposed on the encoder, forcing it to generate latent variables that obey a unit Gaussian distribution, which helps the decoder to reconstruct a better background with respect to the input. Compared with the Reed-Xiaoli (RX), collaborative representation detection (CRD), attribute and edge-preserving filtering-based anomaly detection (AED) and adversarial autoencoder-based anomaly detection (AAE), through two real HSI data sets, the detection performance of the proposed SLDR method is found to be competitive. Jie Lei 0001, Shuo Fang, Weiying Xie, Yunsong Li 0001, Chein-I Chang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Spectral Adversarial Feature Learning for Anomaly Detection in Hyperspectral ImageryabstractTheoretically, hyperspectral images (HSIs) are capable of providing subtle spectral differences between different materials, but in fact, it is difficult to distinguish between background and anomalies because the samples of anomalous pixels in HSIs are limited and susceptible to background and noise. To explore the discriminant features, a spectral adversarial feature learning (SAFL) architecture is specially designed for hyperspectral anomaly detection in this article. In addition to reconstruction loss, SAFL also introduces spectral constraint loss and adversarial loss in the network with batch normalization to extract the intrinsic spectral features in deep latent space. To further reduce the false alarm rate, we present an iterative optimization approach by a weighted suppression function that depends on the contribution rate of each feature to the detection. In particular, the structure tensor matrix is adopted to adaptively calculate the contribution rate of each feature. Benefiting from these improvements, the proposed method is superior to the typical and state-of-the-art methods either in detection probability or false alarm rate. Weiying Xie, Baozhu Liu, Yunsong Li 0001, Jie Lei 0001, Chein-I Chang, Gang He 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Autoencoder and Adversarial-Learning-Based Semisupervised Background Estimation for Hyperspectral Anomaly DetectionabstractReliable detection of anomalies without any prior information is a critical yet challenging task in many applications, not least military and civilian fields. An intelligent anomaly detection system would use the material-specific spectral information in hyperspectral images (HSIs), thereby avoiding the loss of visually confusing objects. However, conventional hyperspectral anomaly detection methods are mainly achieved in an unsupervised way leading to limited performance due to lack of prior knowledge. In this article, we propose a novel autoencoder and adversarial-learning based semisupervised background estimation model (SBEM) that is trained only on the background spectral samples in order to accurately learn the background distribution. In particular, an unsupervised background searching method is firstly conducted on the original HSIs to search the background spectral samples. Our proposed SBEM consists of an encoder, a decoder, and a discriminator to thoroughly capture background distribution. Furthermore, jointly minimizing the reconstruction loss, spectral loss, and adversarial loss during training aids the model to learn the background distribution as required. Experiments on four real HSIs demonstrate that compared to the current state-of-the-art, the proposed framework yields higher detection capability and lower false alarm rate, which shows that it has a significant benefit in the tradeoff between detection accuracy and false alarm rate. Weiying Xie, Baozhu Liu, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Hyperspectral Band Selection for Spectral-Spatial Anomaly DetectionabstractOwing to significantly improved spectral resolution, a hyperspectral imaging sensor can now uncover many unknown subtle material substances. In many cases, anomalies are usually embedded in the background. To develop a means through which these anomalies may be detected and separated from the background, we propose a spectral-spatial anomaly detection method based on a selected band subset. To be specific, we constrain an unsupervised network by making full use of the underlying physical characteristics which are beneficial to hyperspectral anomaly detection. Based on that, a selection criterion is constructed to adaptively select a subset of bands that essentially contain discriminative and informative features between the anomaly and background in an unsupervised manner. Then, the selected bands are simultaneously inputted into the spatial detector and spectral detector. To overcome the deficiencies of detecting anomalies in only one aspect, an adaptive combination of spatial result and the spectral result is introduced. Finally, a simple and powerful iterative suppression is conducted on the initial detection map to further reduce false alarm rate while ensuring detection capability. Extensive empirical researches performed on eighteen publicly available hyperspectral images (HSIs) of different sizes over different scenes demonstrate that our proposed method can achieve an average detection capability of 0.99564, and the average false alarm rate is one order of magnitude lower than the second one. Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Chein-I Chang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Deep Latent Spectral Representation Learning-Based Hyperspectral Band Selection for Target DetectionabstractHyperspectral images (HSIs) can provide discriminative spectral signatures regarding the physical nature of different materials. It is this unique nature that makes HSIs to be of great interest in many fields. However, HSI application faces various challenges due to high dimensionality, redundant information, noisy bands, and insufficient samples. To address these problems, we propose an unsupervised band selection method based on deep latent spectral representation learning, called DLSRL, in this article. It imposes spectral consistency on deep latent space that resolves the issue of insufficient samples and spectral information lost in HSI interpretation. It pursues the low-dimensional optimal representation of the high-dimensional HSIs. In particular, an adaptive mapping relationship is constructed between the deep latent representation and the optimal subset to preserve physical significance optimally. Furthermore, a hierarchical optimization approach is introduced to achieve target detection with the selected subset. To verify the superiority of the proposed method, experiments have been conducted on four data sets captured by different sensors over different scenes. Comparative analyses validate that the proposed method presents superior performance in terms of high detection accuracy and low false alarm rate. Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | SRUN: Spectral Regularized Unsupervised Networks for Hyperspectral Target DetectionabstractThe high dimensionality of a hyperspectral image (HSI) provides the possibility of deeply capturing the underlying and intrinsic characteristics in spectra, such that targets embedded in the background can be detected. However, redundant information, deteriorated bands, and other interferences from background challenge the target detection problem. In this article, an effective feature extraction method based on unsupervised networks is proposed to mine intrinsic properties underlying HSIs. Our approach, called spectral regularized unsupervised networks (SRUN), imposes spectral regularization on autoencoder (AE) and variational AE (VAE) to emphasize spectral consistency, which is more suitable for characterizing spectral information of HSIs by hidden nodes than the original AE and VAE models. Then, we conduct a simple feature selection algorithm on the hidden nodes in the deepest code to select specific nodes that contain distinguishability between target and background, which is based on the spectral angular difference between a known target spectrum and spectra of other pixels in input. The selected nodes are further weighted adaptively to obtain a discriminative map depending on the observation that each selected node provides different contribution rates to target detection. Experimental results on several data sets illustrate that the proposed SRUN-based target detection algorithm is suitable for targets at the subpixel level and those with structural information. Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Qian Du 0001, Gang He 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Hyperspectral Pansharpening With Deep PriorsabstractHyperspectral (HS) image can describe subtle differences in the spectral signatures of materials, but it has low spatial resolution limited by the existing technical and budget constraints. In this paper, we propose a promising HS pansharpening method with deep priors (HPDP) to fuse a low-resolution (LR) HS image with a high-resolution (HR) panchromatic (PAN) image. Different from the existing methods, we redefine the spectral response function (SRF) based on the larger eigenvalue of structure tensor (ST) matrix for the first time that is more in line with the characteristics of HS imaging. Then, we introduce HFNet to capture deep residual mapping of high frequency across the upsampled HS image and the PAN image in a band-by-band manner. Specifically, the learned residual mapping of high frequency is injected into the structural transformed HS images, which are the extracted deep priors served as additional constraint in a Sylvester equation to estimate the final HR HS image. Comparative analyses validate that the proposed HPDP method presents the superior pansharpening performance by ensuring higher quality both in spatial and spectral domains for all types of data sets. In addition, the HFNet is trained in the high-frequency domain based on multispectral (MS) images, which overcomes the sensitivity of deep neural network (DNN) to data sets acquired by different sensors and the difficulty of insufficient training samples for HS pansharpening. Weiying Xie, Jie Lei 0001, Yuhang Cui, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | SOON: Specifically Optimized One-Stage Network for Object Detection in Remote Sensing ImageryabstractWith great significance in military and civilian applications, detecting indistinguishable small objects in wide-scale remote sensing images is still a challenging topic. In this work, we propose a specially optimized one-stage network (SOON) focusing on extracting spatial information of high-resolution images by understanding and analyzing the combination of feature and semantic information of small objects, which consists of feature enhancement, multi-scale detection, and feature fusion. The first part is implemented by constructing a receptive field enhancement (RFE) module and incorporating it into the specific parts of the network where the information of small objects mainly exists. The second part is achieved by four detectors with different sensitivities accessing to the fused and enhanced features, which enables the network to make full use of features in different scales. The third part consolidates the high-level and low-level features by adopting up-sampling, concatenation and convolution operations to build a feature pyramid structure, which explicitly yields strong feature representation and semantic information. In addition, we introduce the Soft-NMS to preserve accurate bounding boxes in the post-processing stage for densely arranged objects. Note that the split and merge strategy, as well as the multi-scale training strategy, are employed in this work. Extensive experiments and thorough analysis are performed on the NWPU VHR-10-v2 dataset and the ACS dataset as compared with several state-of-the-art methods, in which satisfactory performance verifies the effectiveness of the design and optimization. The code will be released for reproduction. Yunsong Li 0001, Jie Lei 0001, Weiying Xie |
ICTAI | 5 |
| 2019 | Spectral constraint adversarial autoencoders approach to feature representation in hyperspectral anomaly detection
Weiying Xie, Jie Lei 0001, Baozhu Liu, Yunsong Li 0001, Xiuping Jia |
Neural Networks | 1 |
| 2019 | High-quality spectral-spatial reconstruction using saliency detection and deep feature enhancement
Weiying Xie, Yanzi Shi, Yunsong Li 0001, Xiuping Jia, Jie Lei 0001 |
Pattern Recognit. | 1 |
| 2019 | Spectral-Spatial Feature Extraction for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection faces various levels of difficulty due to the high dimensionality of hyperspectral images (HSIs), redundant information, noisy bands, and the limited capability of utilizing spectral-spatial information. In this paper, we address these problems and propose a novel approach, called spectral-spatial feature extraction (SSFE), which is based on two main aspects. In the spectral domain, we assume that the anomalous pixels are rarely present and all (or most) of the samples around the anomalies belong to background (BKG). Using this fact, we introduce a suppression function to construct a discriminative feature space and utilize a deep brief network to learn spectral representation and abstraction automatically that are used as inputs to the Mahalanobis distance (MD)-based detector. In the spatial domain, the anomalies appear as a small area grouped by pixels with high correlation among them compared to BKG. Therefore, the objects appearing as a small area are extracted based on attribute filtering, and a guided filter is further employed for local smoothness. More specifically, we extract spatial features of anomalies only from one single band obtained by fusing all bands in the visible wavelength range. Finally, we detect anomalies by jointly considering the spectral and spatial detection results. Several experiments are performed, which show that our proposed method outperforms the state-of-the-art methods. Jie Lei 0001, Weiying Xie, Yunsong Li 0001, Chein-I Chang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | DDLPS: Detail-Based Deep Laplacian Pansharpening for Hyperspectral ImageryabstractIn this paper, we propose a new pansharpening method called detail-based deep Laplacian pansharpening (DDLPS) to improve the spatial resolution of hyperspectral imagery. This method includes three main components: upsampling, detail injection, and optimization. In particular, a deep Laplacian pyramid super-resolution network (LapSRN) improves the resolution of each band. Then, a guided image filter and a gain matrix are used to combine the spatial and spectral details with an optimization problem, which is formed to adaptively select an injection coefficient. The DDLPS method is compared with 11 state-of-the-art or traditional pansharpening approaches. The experimental results demonstrate the superiority of the DDLPS method in terms of both quantitative indices and visual appearance. In addition, the training of LapSRN is based on the data sets of traditional RGB images, which overcomes the practical difficulty of insufficient training samples for pansharpening. Kaiyan Li 0002, Weiying Xie, Qian Du 0001, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Structure Tensor and Guided Filtering-Based Algorithm for Hyperspectral Anomaly DetectionabstractAnomaly detection is one of the most important applications of hyperspectral imaging technology. It is a challenging task due to the high dimensionality of hyperspectral images (HSIs), redundant information, noisy bands, and the limited capability of utilizing spatial information. In this paper, we address these problems and propose a novel anomaly detection method in HSIs. Our approach, called structure tensor and guided filter (STGF)-based strategy for anomaly detection, is based on the characteristics of HSIs. First, a novel band selection algorithm is proposed to reduce dimension, remove noisy bands, and select bands with effective information. Second, the selected bands are decomposed into two parts according to the characteristics of anomalies that are usually in a small area. Followed by this step, the backgrounds are removed through a simple differential operation for each selected band. Considering that not all of the bands provide the same contributions to anomaly detection, we then fuse the differential maps by a novel adaptive weighting method to obtain an initial detection map. Finally, GF is conducted to rectify the previous map under the condition that the neighboring pixels usually have quite strong correlations with each other. Experiments have been conducted on real-scene remote sensing HSI. Comparative analyses validate that the proposed STGF method presents superior performance in terms of detection accuracy and computational time. Weiying Xie, Tao Jiang 0031, Yunsong Li 0001, Xiuping Jia, Jie Lei 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Hyperspectral Image Super-Resolution Using Deep Feature Matrix FactorizationabstractHyperspectral images (HSIs) can describe the subtle differences in the spectral signatures of materials. However, they have low spatial resolution due to various hardware limitations. Improving it via postprocess without an auxiliary high-resolution (HR) image still remains a challenging problem. In this paper, we address this problem and propose a new HSI super-resolution (SR) method. Our approach, called deep feature matrix factorization (DFMF), blends feature matrix extracted by a deep neural network (DNN) with nonnegative matrix factorization strategy for super-resolving real-scene HSI. The estimation of the HR HSI is formulated as a combination of latent spatial feature matrix and spectral feature matrix. In the DFMF model, the input low-resolution (LR) HSI is first partitioned into several subsets according to the correlation matrix, and the key band is selected from each subset. Then, the key band group is super-resolved by a DNN model, and the HR key band group is then used as a guide to carry out deep spatial feature matrix. Specifically, the input LR HSI with prototype reflectance spectral vectors of the scene will be preserved when super-resolving in a spatial domain. Thus, the nonnegative spectral and spatial feature matrices are extracted simultaneously from alternately factorizing the pair of LR HSI and the HR key band group. Finally, the HR HSI is obtained by the integration of the spectral and spatial feature matrices. Experiments have been conducted on real-scene remote sensing HSI. Comparative analyses validate that the proposed DFMF method presents a superior super-resolving performance, as it preserves spectral information better. Weiying Xie, Xiuping Jia, Yunsong Li 0001, Jie Lei 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Deep convolutional networks with residual learning for accurate spectral-spatial denoising
Weiying Xie, Yunsong Li 0001, Xiuping Jia |
Neurocomputing | 1 |
| 2018 | Efficient coarse-to-fine spectral rectification for hyperspectral image
Weiying Xie, Yunsong Li 0001, Weiping Zhou, Yuxuan Zheng |
Neurocomputing | 1 |
| 2018 | Trainable spectral difference learning with spatial starting for hyperspectral image denoising
Weiying Xie, Yunsong Li 0001, Jing Hu 0005, Duan-Yu Chen |
Neural Networks | 1 |
| 2017 | A spatial constraint and deep learning based hyperspectral image super-resolution methodabstractThe image super-resolution (SR) technique, which aims at reconstructing a high-resolution (HR) image from a single low-resolution (LR) image, is a classical problem in computer vision. Limited by the imaging hardware, the spatial resolution of a hyperspectral images (HSI) is usually very coarse. Meanwhile, the spectral information of the HSI is extremely important for its applications and cannot be severely distorted. This paper presents a spatial constraint (SCT) strategy with combination of a deep learning method for HSI SR. The SCT strategy restraints the LR HSI generated by the reconstructed HR HSI should be spatially close to the input LR HSI. The deep learning method learns an end-to-end mapping between the spectral difference of the LR HSI and that of the HR HSI. The mapping is represented as a deep convolutional neural network (CNN). The CNN learned spectral difference is utilized to super-resolve the LR HSI while preserve the important spectral information of the desired HR HSI. Experiments have been conducted on three databases that contains both indoor scenes and outdoor scenes. Comparative analyses have verified the effectiveness of the overall method. Jing Hu 0005, Yunsong Li 0001, Weiying Xie |
IGARSS | 4 |
| 2017 | Hyperspectral image super-resolution using deep convolutional neural network
Yunsong Li 0001, Jing Hu 0005, Weiying Xie, Jiaojiao Li 0001 |
Neurocomputing | 4 |
| 2017 | Hyperspectral Image Super-Resolution by Spectral Difference Learning and Spatial Error CorrectionabstractA hyperspectral image (HSI) super-resolution (SR) is a highly attractive topic in computer vision. However, most existed methods require an auxiliary high-resolution (HR) image with respect to the input low-resolution (LR) HSI. This limits the practicability of these HSI SR methods. Moreover, these methods often destroy the important spectral information. This letter presents a deep spectral difference convolutional neural network (SDCNN) with the combination of a spatial-error-correction (SEC) model for HSI SR. This method allows for full exploration of the spectral and spatial correlations, which achieves a good spatial information enhancement and spectral information preservation. In the proposed method, the key band is automatically selected and super-resolved with the boundary bands. Meanwhile, spectral difference mapping between the LR and HR HSIs can be learned by the SDCNN, and then be transformed according to the SEC model, which aims at correcting the spatial error while preserving the spectral information. The rest nonkey bands will be super-resolved under the guidance of the transformed spectral difference. Experimental results on synthesized and real-scenario HSIs suggest that the proposed method: (1) achieves comparable performance without requiring any auxiliary images of the same scene and (2) requires less computation time than the state-of-the-art methods. Jing Hu 0005, Yunsong Li 0001, Weiying Xie |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Hyperspectral Imagery Denoising by Deep Learning With Trainable Nonlinearity FunctionabstractHyperspectral images (HSIs) can describe subtle differences in the spectral signatures of objects, and thus they are effective in a wide array of applications. However, an HSI is inevitably contaminated with some unwanted components like noise resulting in spectral distortion, which significantly decreases the performance of postprocessing. In this letter, a deep stage convolutional neural network (CNN) with trainable nonlinearity functions is applied for the first time to remove noise in HSIs. Besides the fact that the weight and bias matrices are learned from cubic training clean-noisy HSI patches, the nonlinearity functions in each stage are also trainable, which differ from the conventional CNN with a fixed nonlinearity function. Compared with the state-of-the-art HSI denoising methods, the experimental results on both synthetic and real HSIs confirm that the proposed method can obtain a more effective and efficient performance. Weiying Xie, Yunsong Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Hyperspectral image reconstruction by deep convolutional neural network for classification
Yunsong Li 0001, Weiying Xie, Huaqing Li 0003 |
Pattern Recognit. | 2 |
| 2016 | Breast mass classification in digital mammography based on extreme learning machine
Weiying Xie, Yunsong Li 0001, Yide Ma |
Neurocomputing | 1 |
| 2014 | Plant recognition based on intersecting cortical modelabstractPlant recognition recently becomes more and more attractive in computer vision and pattern recognition. Although some researchers have proposed several methods, their accuracy is not satisfactory. Therefore, a novel method of plant recognition based on leaf image is proposed in the paper. Both shape and texture features are employed in the proposed method Texture feature is extracted by intersecting cortical model, and shape feature is obtained by the representation of center distance sequence. Support vector machine is employed for the classifier. The leaf image is preprocessed to get better quality for extracting features, and then entropy sequence and center distance sequence are obtained by intersecting cortical model and center distance transform, respectively. Redundant data of entropy sequence vector and center distance are reduced by principal component analysis. Finally, feature vector is imported into the classifier for classification. In order to evaluate the performance, several existing methods are used to compare with the proposed method and three leaf image datasets are taken as test samples. The experimental result shows the proposed method gets the better accuracy of recognition than other methods. Zhaobin Wang, Xiaoguang Sun, Yide Ma, Hongjuan Zhang, Yurun Ma, Weiying Xie, Yaonan Zhang |
IJCNN | 6 |