EDBT 2026 Demo / reviewers in the wild / expert
Jie Lei 0001
dblp:61/5501-1
· DBLP profile ↗
65ranked-venue papers
7as first author
46since 2021 · last 2026
0000-0003-0851-6565ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 22 · 1 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Systems, architecture and hardware · 10 · 2 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A2H-MAS: An Algorithm-to-HLS Multi-Agent System for Automated and Reliable FPGA ImplementationabstractBridging the gap between high-level algorithm development and efficient FPGA implementation remains a fundamental challenge, particularly in latency- and resource-constrained domains such as wireless communications. In practice, many signal processing and communication algorithms are first developed and validated at the software level, where rapid iteration and numerical flexibility are critical, making MATLAB a de facto environment for algorithm prototyping and verification. To deploy these algorithms on FPGA platforms, High-Level Synthesis (HLS) is commonly adopted as an intermediate step that enables hardware generation from high-level descriptions. Although HLS significantly improves productivity compared with register-transfer level design, translating MATLAB models into high-quality HLS implementations still requires substantial manual effort and deep domain expertise. Designers must carefully restructure computation patterns, redesign dataflows, and tune synthesis parameters to meet strict performance and resource constraints. Recent advances in large language models (LLMs) suggest a promising direction for automating this translation process. However, current LLM-based approaches remain unreliable in complex hardware design workflows, as they suffer from hallucinations, forgetting, limited domain expertise, and often overlook key performance metrics such as latency, throughput, and resource utilization. Jie Lei 0001, Ruofan Jia, Jian (Andrew) Zhang, Hao Zhang 0082 |
FPGA | 1 |
| 2026 | HLS-Based Algorithm-Hardware Co-Design of MIMO-OFDM Receiver for Tactical Jamming SuppressionabstractThis paper presents an HLS-based algorithm-hardware co-design methodology and complete FPGA hardware accelerator for multi-user massive MIMO-OFDM receivers operating in contested tactical environments with jamming suppression capabilities. We develop a systematic bi-directional co-design methodology using high-level synthesis (HLS) where algorithms provide functional verification constraints while hardware synthesis feedback drives algorithmic complexity reduction, enabling efficient transformation from signal processing algorithms to optimized circuit implementations on the Xilinx ZCU111 radio frequency system-on-chip (RFSoC) platform. The primary contributions include: 1) hardware architecture innovations featuring QR decomposition-based synchronization achieving 330 MHz post-route operation and optimized frequency-domain minimum mean square error (FD-MMSE) equalization with systematic loop restructuring for data dependency removal in substitution modules, reducing hardware resources by 56% lookup tables (LUTs), 58% flip-flops (FFs), and 67% digital signal processing (DSP) blocks while maintaining real-time throughput; 2) systematic HLS-based co-design framework enabling automated architecture exploration with hardware-oriented algorithm adaptations including silent-period frame structure and Cholesky-based decision-feedback equalization; 3) complete system integration validated through field trials demonstrating 8 dB jamming suppression improvement with reliable spatial division multiple access (SDMA) communications. The presented methodology provides insights applicable to broader signal processing systems with stringent real-time constraints. Jian (Andrew) Zhang, Jie Lei 0001, Hao Zhang 0082, Anh Tuyen Le, Kin-Ping Hui, Damien Phillips, Asanka Kekirigoda, Alan Allwright |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2026 | LoME: LoRA-Driven Multimodal Extractor for RGB-X Vision TasksabstractRGB-X multimodal vision tasks present a highly promising approach to enhancing model performance in complex visual conditions. Existing multimodal frameworks are based on either the symmetric parallel network of feature fusion or the shared network of input fusion. However, parallel networks suffer from uncontrollable parameters and imbalanced optimization across modal branches, while shared networks often lead to a lack of diversity in gradient optimization. To address these challenges, we propose the LoRA-driven Multimodal Extractor (LoME), following a comprehensive analysis of existing multimodal frameworks. The low-rank properties of modal adapters for LoME ensure controllable growth in model parameters as the number of modalities increases. The dynamic parameter fusion between adapters and the shared feature extractor decouples gradient optimization directions, effectively mitigating imbalances caused by multimodal data biases while preserving complementary features. Moreover, we employ a training strategy based on dynamic rank allocation to reduce computational overhead and enhance modal diversity expression. We validate the effectiveness and generalizability of LoME across three multimodal vision tasks. LoME achieves superior performance compared to previous state-of-the-art methods on multiple datasets. For example, on the DroneVehicle dataset, our method achieves a 10.4% improvement in accuracy compared to the SOTA method, while the parameter overhead is reduced to 23% of the previous network (44.63M). The code has been open-sourced at https://github.com/zyszxhy/LoME. Weiying Xie, Tianlin Hui, Daixun Li, Jie Lei 0001, Yunsong Li 0001, Leyuan Fang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | M³amba: CLIP-Driven Mamba Model for Multi-Modal Remote Sensing ClassificationabstractMulti-modal fusion holds great promise for integrating information from different modalities. However, due to a lack of consideration for modal consistency, existing multi-modal fusion methods in the field of remote sensing still face challenges of incomplete semantic information and low computational efficiency in their fusion designs. Inspired by the observation that the visual language pre-training model CLIP can effectively extract strong semantic information from visual features, we propose M3amba, a novel end-to-end CLIP-driven Mamba model for multi-modal fusion to address these challenges. Specifically, we introduce CLIP-driven modality-specific adapters in the fusion architecture to avoid the bias of understanding specific domains caused by direct inference, making the original CLIP encoder modality-specific perception. This unified framework enables minimal training to achieve a comprehensive semantic understanding of different modalities, thereby guiding cross-modal feature fusion. To further enhance the consistent association between modality mappings, a multi-modal Mamba fusion architecture with linear complexity and a cross-attention module Cross-SS2D are designed, which fully considers effective and efficient information interaction to achieve complete fusion. Extensive experiments have shown that M3amba has an average performance improvement of at least 5.98% compared with the state-of-the-art methods in multi-modal hyperspectral image classification tasks in the remote sensing field, while also demonstrating excellent training efficiency, achieving a double improvement in accuracy and efficiency. The code is released athttps://github.com/kaka-Cao/M3amba. Mingxiang Cao, Weiying Xie, Xin Zhang 0092, Kai Jiang 0001, Jie Lei 0001, Yunsong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | SeaDATE: Remedy Dual-Attention Transformer With Semantic Alignment via Contrast Learning for Multimodal Object DetectionabstractMultimodal object detection leverages diverse modal information to enhance the accuracy and robustness of detectors. Due to its ability to capture long-range dependencies, the Transformer model provides a powerful mechanism for integrating multimodal features during feature extraction. This capability significantly enhances the accuracy of multimodal object detection by addressing the limitations of local feature extraction inherent in traditional methods. However, current methods merely stack Transformer-guided fusion techniques without exploring their capability to extract features at various depth layers of network, thus limiting the improvements in detection performance. In this paper, we introduce an accurate and efficient multimodal object detection method named SeaDATE. Initially, we propose a novel dual attention Feature Fusion (DTF) module that, under Transformer’s guidance, integrates local and global information through a dual attention mechanism, strengthening the fusion of modal features from orthogonal perspectives using spatial and channel tokens. Meanwhile, our theoretical analysis and empirical validation demonstrate that the Transformer-guided fusion method, treating images as sequences of pixels for fusion, performs better on shallow features’ detail information compared to deep semantic information. To address this, we designed a contrastive learning (CL) module aimed at learning features of multimodal samples, remedying the shortcomings of Transformer-guided fusion in extracting deep semantic features, and effectively utilizing cross-modal information. Extensive experiments and ablation studies on the FLIR, LLVIP, and M3FD datasets have proven our method to be effective, achieving state-of-the-art detection performance. Shuhan Dong, Weiying Xie, Danian Yang, Yunsong Li 0001, Jiayuan Tian, Jie Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | ShiftQuant: Toward Accurate and Efficient Sub-8-bit Integer TrainingabstractNeural network training is a memory- and compute-intensive task. Quantization, which enables low-bitwidth formats in training, can significantly mitigate the workload. To reduce quantization error, recent methods have developed new data formats and additional pre-processing operations on quantizers. However, it remains quite challenging to achieve high accuracy and efficiency simultaneously. In this paper, we explore sub-8-bit integer training from its essence of gradient descent optimization. Our integer training framework includes two components: ShiftQuant to realize accurate gradient estimation, and L1 normalization to smoothen the loss landscape. ShiftQuant attains performance that approaches the theoretical upper bound of group quantization. Furthermore, it liberates group quantization from inefficient memory rearrangement. The L1 normalization facilitates the implementation of fully quantized normalization layers with impressive convergence accuracy. Our method frees sub-8-bit integer training from pre-processing and supports general devices. This framework achieves negligible accuracy loss across various neural networks and tasks (0.92% on 4-bit ResNets, 0.61% on 6-bit Transformers). The prototypical implementation of ShiftQuant achieves more than 1.85×/15.3% performance improvement on CPU/GPU compared to its FP16 counterparts, and 33.9% resource consumption reduction on FPGA than the FP16 counterparts. The proposed fully-quantized L1 normalization layers achieve more than 35.54% improvement in throughout on CPU compared to traditional L2 normalization layers. Moreover, theoretical analysis verifies the advancement of our method. Wenjin Guo, Donglai Liu, Weiying Xie, Yunsong Li 0001, Xuefei Ning, Zihan Meng, Shulin Zeng, Jie Lei 0001, Zhenman Fang, Yu Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Distributed Deep Learning With Gradient Compression for Big Remote Sensing Image InterpretationabstractFast and reliable interpretation of high-dimensional hyperspectral images (HSIs) can provide great support to remote sensing-based Earth observations. Targets of interest in HSI can be detected using deep neural networks (DNNs) for background learning on an acquired image where the occurrence probability of background samples is much greater than that of targets, accounting for more than 95% of the whole scene. However, there is an increasing gap between theory and feasible application, because of the contradiction between massive hyperspectral data and resource-limited Internet of Things (IoT)/edge device hardware like satellite. To facilitate the deployment of hyperspectral target detection (HTD) in an edge computing environment, we introduce distributed background learning-a decentralized deep learning approach to meet the computing requirements of exploding high-dimensional data and larger DNNs. To address the communication bottleneck caused by gradient exchange during distributed learning, the proposed gradient compression solution, named gradient compression via centroid (GCC), uniquely compresses the most replaceable gradients with redundant information, thereby reducing communication overhead while maintaining accuracy. To illustrate the feasibility of the proposed method, we test it over two very large hyperspectral datasets with a total size of about 3.2 gigabytes (GBs) on a distributed system based on Ring All-reduce. We show that HTD based on distributed background learning outperforms those developed on a single node in terms of speed. Besides, the GCC compresses 50% gradients with only 0.01% loss of target detection accuracy to greatly reduce the communication overhead, surpassing existing gradient compression methods. It is expected that this framework will accelerate the introduction of distributed training on IoT/edge devices. Weiying Xie, Jitao Ma, Tianen Lu, Yunsong Li 0001, Jie Lei 0001, Leyuan Fang, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | JointSQ: Joint Sparsification-Quantization for Distributed LearningabstractGradient sparsification and quantization offer a promising prospect to alleviate the communication overhead problem in distributed learning. However, direct combination of the two results in suboptimal solutions, due to the fact that sparsification and quantization haven't been learned together. In this paper, we propose Joint Sparsification-Quantization (JointSQ) inspired by the discovery that sparsification can be treated as 0-bit quantization, regardless of architectures. Specifically, we mathematically formu-late JointSQ as a mixed-precision quantization problem, expanding the solution space. It can be solved by the designed MCKP-Greedy algorithm. Theoretical analysis demon-strates the minimal compression noise of JointSQ, and ex-tensive experiments on various network architectures, including CNN, RNN, and Transformer, also validate this point. Under the introduction of computation overhead consistent with or even lower than previous methods, JointSQ achieves a compression ratio of 1000× on different models while maintaining near-lossless accuracy and brings 1.4× to 2.9× speedup over existing methods. Weiying Xie, Jitao Ma, Yunsong Li 0001, Jie Lei 0001, Donglai Liu, Leyuan Fang |
CVPR | 5 |
| 2024 | DA-BEV: Unsupervised Domain Adaptation for Bird's Eye View Perception
Kai Jiang 0001, Jiaxing Huang 0001, Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Ling Shao 0001, Shijian Lu |
ECCV (82) | 4 |
| 2024 | E4SA: An Ultra-Efficient Systolic Array Architecture for 4-Bit Convolutional Neural NetworksabstractMany studies have demonstrated that 4-bit precision quantization can achieve comparable accuracy to floating-point DNNs, sparking significant interest in efficiently accelerating compressed DNNs, especially 4-bit convolutions, on edge devices. However, we observe that conventional systolic array (SA) architectures designed for DNNs cannot fully exploit the advantages of high DSP computational density offered by 4-bit DSP packing. Although state-of-the-art FPGA-based SA architectures (e.g., AutoSA) exhibit flexibility in accommodating 4-bit DSP packing, they suffer from resource consumption and data supply latency issues, especially when adapting to various convolution spatial sizes. This work introduces a customizable and ultra-efficient SA architectural template for 4-bit convolution, called E4SA. First, we propose a fine-grained row-temporal weight stationary dataflow that aligns with the specific requirements of 4-bit DSP full packing (4bF packing). Based on this, we design a cost-effective SA unit (SAU) composed of 4bF-packing-based processing elements (PEs) to enhance computational efficiency. This includes column-shared packed-data splitters and shift-register-based feature-map/weight fetchers to ensure continuous data supply, all of which are locally interconnected via more cost-effective registers. In addition, we develop a two-level hierarchy SA that decomposes the original large SA into parallel 4×4 SAU sets, which not only allows multiple PEs in the same column to share data splitting and reorganization logic and thus reducing the LUT overhead, but also maintains near-theoretical latency across various convolutional spatial sizes. Experimental results demonstrate that E4SA achieves up to 576.6 GOPS with 13.8× higher GOPS/DSP efficiency and 51.6× higher GOPS/kLUTs efficiency compared to 4-bit AutoSA-based design. Geng Yang 0001, Jie Lei 0001, Zhenman Fang, Junrong Zhang 0002, Weiying Xie, Yunsong Li 0001 |
FPGA | 2 |
| 2024 | SA4: A Comprehensive Analysis and Optimization of Systolic Array Architecture for 4-bit ConvolutionsabstractMany studies have demonstrated that 4-bit precision quantization can maintain accuracy levels comparable to those of floating-point deep neural networks (DNNs). Thus, it has sparked a keen interest in the efficient acceleration of such compressed DNNs, especially 4-bit convolutions, on edge devices. However, we observe that conventional systolic array (SA) architectures, widely adopted for DNN acceleration, fail to fully exploit the high computational density benefits of 4 -bit DSP packing. In this paper, we conduct the first comprehensive analysis of the integration of modern DSP packing techniques (specifically, 4-bit fully DSP packing) into the 4-bit systolic array design for convolutions. First, we introduce a row-temporal weight stationary 4-bit SA dataflow that complements the loop execution order inherent in 4-bit fully DSP packing in conventional SAs, which is called BaseSA. Next, we analyze the performance and resource efficiency of BaseSA, and identify two inefficiencies in the integration: 1) excessive LUT resource utilization that constraints the overall SA size, and 2) large latency gap to the theoretical optimum, due to various stalls in data supplies. To overcome these obstacles, we propose SA4: an HLS-based, customizable, and ultra-efficient hierarchical $\underline{\text { SA}}$ architecture optimized for 4 -bit convolutions. The core unit in SA4 is a delicately designed cost-effective SA unit (SAU), which 1) replaces the costly buffer-based data suppliers for activations and weights with shift-register-based ones, 2) replaces LUT-intensive FIFO connections between SA PEs (processing elements) with registers, and 3) replaces the finite state machines (FSM) and data unpacking logic inside each PE with a global FSM inside each SAU and a data splitter shared by a column of PEs. While such an SAU can only support a small spatial size for an SA due to its delicate design, we further scale it out using an array of SAUs. Experimental results show that our proposed SA4 achieves 1153.2 GOPS on the AMD-Xilinx Ultra96-V2 FPGA, with a $13.8 \times$ increase in GOPS/DSP efficiency and a $49 \times$ increase in GOPS/kLUTs efficiency compared to a straightforward SA and 4-bit DSP packing integration. Our SA4 project is open sourced here: https://github.com/Michaela1224/SA4. Geng Yang 0001, Jie Lei 0001, Zhenman Fang, Junrong Zhang 0002, Weiying Xie, Yunsong Li 0001 |
FPL | 2 |
| 2024 | SDA: Low-Bit Stable Diffusion Acceleration on Edge FPGAsabstractThis paper introduces SDA, the first effort to adapt the expensive stable diffusion (SD) model for edge FPGA deployment. First, we apply quantization-aware training to quantize its weights to 4 -bit and activations to 8 -bit ($W 4 A 8$) with a negligible accuracy loss. Based on that, we propose a high-performance hybrid systolic array (hybridSA) architecture that natively executes convolution and attention operators across varying quantization bit-widths (e.g., $W 4 A 8$ and all 8 -bit $Q K^{T} V$ in attention). To improve computational efficiency, hybridSA integrates diverse DSP packing techniques into hybrid weightstationary and output-stationary dataflows that are optimized for convolution and attention. It also supports flexible dataflow transitions to address the distinct demands of its output sequence by subsequent nonlinear operators. Moreover, we observe that nonlinear operators become the new performance bottleneck after the acceleration of convolution and attention, and offload them onto the FPGA as well. To reduce the latency of each nonlinear operator, we pipeline its own execution at a fine granularity. To minimize the resource utilization of nonlinear operators, we carefully balance their execution with hybridSA in a coarse-grained pipeline. Experimental results demonstrate that our low-bit ($W 4 A 8$) SDA accelerator on the embedded AMDXilinx ZCU102 FPGA achieves a speedup of $97.3 \times$ (which takes about $\mathrm{2 . 1}$ minutes for one SD inference), compared to the original SD-v1.5 model on the ARM Cortex-A53 CPU (which takes about 3.5 hours for one SD inference). Our SDA project is open sourced here: https://github.com/Michaela1224/SDA_code. Geng Yang 0001, Yanyue Xie, Zhong Jia Xue, Sung-En Chang, Yanyu Li, Peiyan Dong, Jie Lei 0001, Weiying Xie, Yanzhi Wang 0001, Xue Lin 0001, Zhenman Fang |
FPL | 7 |
| 2024 | Adaptive Hierarchical Aggregation for Federated Object DetectionabstractIn practical object detection scenarios, distributed data and stringent privacy protections significantly limit the feasibility of traditional centralized training methods. Federated learning (FL) emerges as a promising solution to this dilemma. Nonetheless, the issue of data heterogeneity introduces distinct challenges to federated object detection, evident in diminished object perception, classification and localization abilities. In response, we introduce a task-driven federated learning methodology, dubbed Adaptive Hierarchical Aggregation (FedAHA), tailored to overcome these obstacles. Our algorithm unfolds in two strategic phases from shallow-to-deep layers: (1) Structure-aware Aggregation (SAA) aligns feature extractors during the aggregation phase, thus bolstering the global model's object perception capabilities; (2) Convex Semantic Calibration (CSC) leverages convex function theory to average semantic features instead of model parameters, enhancing the global model's classification and localization precision. We demonstrate experimentally and theoretically the effectiveness of the proposed two modules respectively. Our method consistently outperforming the state-of-the-art methods across multiple valuable application scenarios from 2.26% to 7.61%. Moreover, we build a real FL system using Raspberry Pis to demonstrate that our approach achieves a good trade-off between performance and efficiency. Ruofan Jia, Weiying Xie, Jie Lei 0001, Yunsong Li 0001 |
ACM Multimedia | 3 |
| 2024 | Domain Adaptation for Large-Vocabulary Object DetectorsabstractLarge-vocabulary object detectors (LVDs) aim to detect objects of many categories, which learn super objectness features and can locate objects accurately while applied to various downstream data. However, LVDs often struggle in recognizing the located objects due to domain discrepancy in data distribution and object vocabulary. At the other end, recent vision-language foundation models such as CLIP demonstrate superior open-vocabulary recognition capability.
This paper presents KGD, a Knowledge Graph Distillation technique that exploits the implicit knowledge graphs (KG) in CLIP for effectively adapting LVDs to various downstream domains.
KGD consists of two consecutive stages: 1) KG extraction that employs CLIP to encode downstream domain data as nodes and their feature distances as edges, constructing KG that inherits the rich semantic relations in CLIP explicitly;
and 2) KG encapsulation that transfers the extracted KG into LVDs to enable accurate cross-domain object classification.
In addition, KGD can extract both visual and textual KG independently, providing complementary vision and language knowledge for object localization and object classification in detection tasks over various downstream domains.
Experiments over multiple widely adopted detection benchmarks show that KGD outperforms the state-of-the-art consistently by large margins.
Codes will be released. Kai Jiang 0001, Jiaxing Huang 0001, Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Ling Shao 0001, Shijian Lu |
NeurIPS | 4 |
| 2024 | E2E-MFD: Towards End-to-End Synchronous Multimodal Fusion DetectionabstractMultimodal image fusion and object detection are crucial for autonomous driving. While current methods have advanced the fusion of texture details and semantic information, their complex training processes hinder broader applications. Addressing this challenge, we introduce E2E-MFD, a novel end-to-end algorithm for multimodal fusion detection. E2E-MFD streamlines the process, achieving high performance with a single training phase. It employs synchronous joint optimization across components to avoid suboptimal solutions associated to individual tasks. Furthermore, it implements a comprehensive optimization strategy in the gradient matrix for shared parameters, ensuring convergence to an optimal fusion detection configuration. Our extensive testing on multiple public datasets reveals E2E-MFD's superior capabilities, showcasing not only visually appealing image fusion but also impressive detection outcomes, such as a 3.9\% and 2.0\% $\text{mAP}_{50}$ increase on horizontal object detection dataset M3FD and oriented object detection dataset DroneVehicle, respectively, compared to state-of-the-art approaches. Mingxiang Cao, Weiying Xie, Jie Lei 0001, Daixun Li, Wenbo Huang 0001, Yunsong Li 0001 |
NeurIPS | 4 |
| 2024 | Multimodal Informative ViT: Information Aggregation and Distribution for Hyperspectral and LiDAR ClassificationabstractIn multimodal land cover classification (MLCC), a common challenge is the redundancy in data distribution, where task-irrelevant information from multiple modalities can hinder the effective integration of their unique features. To tackle this, we introduce the Multimodal Informative Vit (MIVit), a system with an innovative information aggregate-distributing mechanism. This approach redefines redundancy levels and integrates performance-aware elements into the fused representation, facilitating the learning of semantics in both forward and backward directions. MIVit stands out by significantly reducing redundancy in the empirical distribution of each modality’s separate and fused features. It employs oriented attention fusion (OAF) for extracting shallow local shape features across modalities in horizontal and vertical dimensions, and a Transformer feature extractor for extracting deep global features through long-range attention. We also propose an information aggregation constraint (IAC) based on mutual information, designed to remove redundant information and preserve complementary information within embedded features. Additionally, the information distribution flow (IDF) in MIVit enhances performance-awareness by distributing global classification information across different modalities’ feature maps. This architecture also addresses missing modality challenges with lightweight independent modality classifiers, reducing the computational load typically associated with Transformers. Our results show that MIVit’s bidirectional aggregate-distributing mechanism between modalities is highly effective, achieving an average overall accuracy of 95.56% across three multimodal datasets. This performance surpasses current state-of-the-art methods in MLCC. The code for MIVit is accessible at https://github.com/icey-zhang/MIViT. Jie Lei 0001, Weiying Xie, Geng Yang 0001, Daixun Li, Yunsong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | SwiMDiff: Scene-Wide Matching Contrastive Learning With Diffusion Constraint for Remote Sensing ImageabstractWith recent advancements in aerospace technology, the volume of unlabeled remote sensing image (RSI) data has increased dramatically. Effectively leveraging this data through self-supervised learning (SSL) is vital in the field of remote sensing. However, current methodologies, particularly contrastive learning (CL), a leading SSL method, encounter specific challenges in this domain. Firstly, CL often mistakenly identifies geographically adjacent samples with similar semantic content as negative pairs, leading to confusion during model training. Secondly, as an instance-level discriminative task, it tends to neglect the essential fine-grained features and complex details inherent in unstructured RSIs. To overcome these obstacles, we introduce SwiMDiff, a novel self-supervised pre-training framework designed for RSIs. SwiMDiff employs a scene-wide matching approach that effectively recalibrates labels to recognize data from the same scene as false negatives. This adjustment makes CL more applicable to the nuances of remote sensing. Additionally, SwiMDiff seamlessly integrates CL with a diffusion model. Through the implementation of pixel-level diffusion constraints, we enhance the encoder’s ability to capture both the global semantic information and the fine-grained features of the images more comprehensively. Our proposed framework significantly enriches the information available for downstream tasks in remote sensing. Demonstrating exceptional performance in change detection and land-cover classification tasks, SwiMDiff proves its substantial utility and value in the field of remote sensing. Jiayuan Tian, Jie Lei 0001, Weiying Xie, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Distribution-Aware Interactive Attention Network and Large-Scale Cloud Recognition Benchmark on FY-4A Satellite ImageabstractAccurate cloud recognition and warning are crucial for various applications, including in-flight support, weather forecasting, and climate research. However, recent deep learning algorithms have predominantly focused on detecting cloud regions in satellite imagery, with insufficient attention to the specificity required for accurate cloud recognition. This limitation inspired us to develop the novel FY-4A-Himawari-8 (FYH) dataset, which includes nine distinct cloud categories and uses precise domain adaptation methods to align 70419 image-label pairs (including 110000 train/5500 test$100\times 100$size images) in terms of projection, temporal resolution, and spatial resolution, thereby facilitating the training of supervised deep learning networks. Given the complexity and diversity of cloud formations, we have thoroughly analyzed the challenges inherent to cloud recognition tasks, examining the intricate characteristics and distribution of the data. To effectively address these challenges, we designed a distribution-aware interactive-attention network (DIAnet), which preserves pixel-level details through a high-resolution branch and a parallel multiresolution cross-branch. We also integrated a distribution-aware loss (DAL) to mitigate the imbalance across cloud categories. An interactive attention module (IAM) further enhances the robustness of feature extraction combined with spatial and channel information. Empirical evaluations on the FYH dataset demonstrate that our method outperforms other cloud recognition networks, achieving superior performance in terms of mean intersection over union (mIoU). The code for implementing DIAnet is available athttps://github.com/icey-zhang/DIAnet. Jie Lei 0001, Weiying Xie, Kai Jiang 0001, Xin Zhang 0092, Mingxiang Cao, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Block-Wise Partner Learning for Model CompressionabstractDespite the great potential of convolutional neural networks (CNNs) in various tasks, the resource-hungry nature greatly hinders their wide deployment in cost-sensitive and low-powered scenarios, especially applications in remote sensing. Existing model pruning approaches, implemented by a "subtraction" operation, impose a performance ceiling on the slimmed model. Self-knowledge distillation (Self-KD) resorts to auxiliary networks that are only active in the training phase for performance improvement. However, the knowledge is holistic and crude, and the learning-based knowledge transfer is mediate and lossy. Here, we propose a novel model-compression method, termed block-wise partner learning (BPL), which comprises "extension" and "fusion" operations and liberates the compressed model from the bondage of baseline. Different from the Self-KD, the proposed BPL creates a partner for each block for performance enhancement in training. For the model to absorb more diverse information, a diversity loss (DL) is designed to evaluate the difference between the original block and the partner. Besides, the partner is fused equivalently instead of being discarded directly. After training, we can simply adopt the fused compressed model that contains the enhancement information of partners but with fewer parameters and less inference cost. As validated using the UC Merced land-use, NWPU-RESISC45, and RSD46-WHU datasets, the BPL demonstrates superiority over other compared model-compression approaches. For example, it attains a substantial floating-point operations (FLOPs) reduction of 73.97% with only 0.24 accuracy (ACC.) loss for ResNet-50 on the UC Merced land-use dataset. The code is available at https://github.com/zhangxin-xd/BPL. Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Kai Jiang 0001, Leyuan Fang, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | HyBNN: Quantifying and Optimizing Hardware Efficiency of Binary Neural NetworksabstractBinary neural network (BNN), where both the weight and the activation values are represented with one bit, provides an attractive alternative to deploy highly efficient deep learning inference on resource-constrained edge devices. However, our investigation reveals that, to achieve satisfactory accuracy gains, state-of-the-art (SOTA) BNNs, such as FracBNN and ReActNet, usually have to incorporate various auxiliary floating-point components and increase the model size, which in turn degrades the hardware performance efficiency. In this article, we aim to quantify such hardware inefficiency in SOTA BNNs and further mitigate it with negligible accuracy loss. First, we observe that the auxiliary floating-point (AFP) components consume an average of 93% DSPs, 46% LUTs, and 62% FFs, among the entire BNN accelerator resource utilization. To mitigate such overhead, we propose a novel algorithm-hardware co-design, called FuseBNN , to fuse those AFP operators without hurting the accuracy. On average, FuseBNN reduces AFP resource utilization to 59% DSPs, 13% LUTs, and 16% FFs. Second, SOTA BNNs often use the compact MobileNetV1 as the backbone network but have to replace the lightweight 3 × 3 depth-wise convolution (DWC) with the 3 × 3 standard convolution (SC, e.g., in ReActNet and our ReActNet-adapted BaseBNN) or even more complex fractional 3 × 3 SC (e.g., in FracBNN) to bridge the accuracy gap. As a result, the model parameter size is significantly increased and becomes 2.25× larger than that of the 4-bit direct quantization with the original DWC (4-Bit-Net); the number of multiply-accumulate operations is also significantly increased so that the overall LUT resource usage of BaseBNN is almost the same as that of 4-Bit-Net. To address this issue, we propose HyBNN , where we binarize depth-wise separation convolution (DSC) blocks for the first time to decrease the model size and incorporate 4-bit DSC blocks to compensate for the accuracy loss. For the ship detection task in synthetic aperture radar imagery on the AMD-Xilinx ZCU102 FPGA, HyBNN achieves a detection accuracy of 94.8% and a detection speed of 615 frames per second (FPS), which is 6.8× faster than FuseBNN+ (94.9% accuracy) and 2.7× faster than 4-Bit-Net (95.9% accuracy). For image classification on the CIFAR-10 dataset on the AMD-Xilinx Ultra96-V2 FPGA, HyBNN achieves 1.5× speedup and 0.7% better accuracy over SOTA FracBNN. Geng Yang 0001, Jie Lei 0001, Zhenman Fang, Yunsong Li 0001, Weiying Xie |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2023 | Toward Stable, Interpretable, and Lightweight Hyperspectral Super-ResolutionabstractFor real applications, existing HSI-SR methods are not only limited to unstable performance under unknown scenarios but also suffer from high computation consumption. In this paper, we develop a new coordination optimization framework for stable, interpretable, and lightweight HSI-SR. Specifically, we create a positive cycle between fusion and degradation estimation under a new probabilistic framework. The estimated degradation is applied to fusion as guidance for a degradation-aware HSI-SR. Under the framework, we establish an explicit degradation estimation method to tackle the indeterminacy and unstable performance caused by the black-box simulation in previous methods. Considering the interpretability in fusion, we integrate spectral mixing prior into the fusion process, which can be easily realized by a tiny autoencoder, leading to a dramatic release of the computation burden. Based on the spectral mixing prior, we then develop a partial fine-tune strategy to reduce the computation cost further. Comprehensive experiments demonstrate the superiority of our method against the state-of-the-arts under synthetic and real datasets. For instance, we achieve a 2.3 dB promotion on PSNR with$120\times$model size reduction and$4300 \times$FLOPs reduction under the CAVE dataset. Code is available in https://github.com/WenjinGuo/DAEM. Wen-jin Guo, Weiying Xie, Kai Jiang 0001, Yunsong Li 0001, Jie Lei 0001, Leyuan Fang |
CVPR | 5 |
| 2023 | HyBNN: Quantifying and Optimizing Hardware Efficiency of Binary Neural NetworksabstractBinary neural network (BNN) has recently presented a promising opportunity for deep learning inferences on resource-constrained edge devices. Using extreme data precision, i.e., 1-bit weight and 1-bit activation, BNN not only significantly reduces the network memory footprint, but also trades massive multiply-accumulate operations for much cheaper logical XNOR and population count operations. However, our investigation reveals that, to achieve satisfactory accuracy gains, state-of-the-art (SOTA) BNNs, such as FracBNN [4] and ReActNet [1], usually have to incorporate various auxiliary floating-point ($AFP$) components and increase the model size, which in turn degrades the hardware performance efficiency. Geng Yang 0001, Jie Lei 0001, Zhenman Fang, Yunsong Li 0001, Weiying Xie |
FCCM | 2 |
| 2023 | Weakly supervised adversarial learning via latent space for hyperspectral target detection
Weiying Xie, Yunsong Li 0001, Kai Jiang 0001, Jie Lei 0001, Qian Du 0001 |
Pattern Recognit. | 5 |
| 2023 | Filter Pruning via Learned Representation Median in the Frequency DomainabstractIn this article, we propose a novel filter pruning method for deep learning networks by calculating the learned representation median (RM) in frequency domain (LRMF). In contrast to the existing filter pruning methods that remove relatively unimportant filters in the spatial domain, our newly proposed approach emphasizes the removal of absolutely unimportant filters in the frequency domain. Through extensive experiments, we observed that the criterion for "relative unimportance" cannot be generalized well and that the discrete cosine transform (DCT) domain can eliminate redundancy and emphasize low-frequency representation, which is consistent with the human visual system. Based on these important observations, our LRMF calculates the learned RM in the frequency domain and removes its corresponding filter, since it is absolutely unimportant at each layer. Thanks to this, the time-consuming fine-tuning process is not required in LRMF. The results show that LRMF outperforms state-of-the-art pruning methods. For example, with ResNet110 on CIFAR-10, it achieves a 52.3% FLOPs reduction with an improvement of 0.04% in Top-1 accuracy. With VGG16 on CIFAR-100, it reduces FLOPs by 35.9% while increasing accuracy by 0.5%. On ImageNet, ResNet18 and ResNet50 are accelerated by 53.3% and 52.7% with only 1.76% and 0.8% accuracy loss, respectively. The code is based on PyTorch and is available at https://github.com/zhangxin-xd/LRMF. Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001 |
IEEE Trans. Cybern. | 4 |
| 2023 | CBFF-Net: A New Framework for Efficient and Accurate Hyperspectral Object TrackingabstractVisual object tracking is a fundamental task in computer vision, and thrived in recent decades. With the development of snapshot hyperspectral sensors, efforts have been made to exploit tracking the object with hyperspectral (HS) videos to overcome the inherent limitation of RGB images. Existing HS tracking algorithms extract the deep features from image data separately, which break the interaction information between bands. Therefore, the discrimination ability of HS trackers is limited and the efficiency of the existing HS algorithms is low. In this paper, a novel algorithm (CBFF-Net) is proposed for HS object tracking to improve the discrimination ability and reduce the computational complexity. Specifically, the backbone and head network are implemented with modules of a transferred RGB object tracking network to carry out the HS target tracking task while maintaining the discrimination ability learned from RGB data. Moreover, a bi-directional multiple deep feature fusion (BMDFF) module is proposed to fuse the features extracted from different bands of the HS images, and a cross-band group attention (CBGA) module is introduced to learn interaction information across bands of the HS images. Experiments results indicate the superiority in performance of CBFF-Net, and it runs at 24 frames per second. Pan Liu 0009, Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | A Semantic Transferred Priori for Hyperspectral Target Detection With Spatial-Spectral AssociationabstractHyperspectral target detection is a crucial application that encompasses military, environmental, and civil needs. Target detection algorithms that have prior knowledge often assume a fixed laboratory target spectrum, which can differ significantly from the test image in the scene. This discrepancy can be attributed to various factors such as atmospheric conditions and sensor internal effects, resulting in decreased detection accuracy. To address this challenge, this article introduces a novel method for detecting hyperspectral image (HSI) targets with certain spatial information, referred to as the semantic transferred priori for hyperspectral target detection with spatial–spectral association (SSAD). Considering that the spatial textures of the HSI remain relatively constant compared to the spectral features, we propose to extract a unique and precise target spectrum from each image data via target detection in its spatial domain. Specifically, employing transfer learning, we designed a semantic segmentation network adapted for HSIs to discriminate the spatial areas of targets and then aggregated a customized target spectrum with those spectral pixels localized. With the extracted target spectrum, spectral dimensional target detection is performed subsequently by the constrained energy minimization (CEM) detector. The final detection results are obtained by combining an attention generator module to aggregate target features and deep stacked feature fusion (DSFF) module to hierarchically reduce the false alarm rate. Experiments demonstrate that our proposed method achieves higher detection accuracy and superior visual performance compared to the other benchmark methods. Jie Lei 0001, Simin Xu, Weiying Xie, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | A Model-Driven Deep Mixture Network for Robust Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) aims to identify samples with unknown atypical spectra from the background. Deep learning (DL)-based methods, particularly autoencoders (AEs), have proven effective in uncovering the underlying profiles for HAD. However, in real-world applications of hyperspectral images (HSIs), complex background land-covers and anomaly corruptions are common, leading to two issues: 1) A low-dimensional manifold characterized by DL-based HAD methods can only reveal a few underlying variation factors of the background distribution and cannot capture the complex structures behind land-covers of all categories. 2) DL-based HAD methods trained on anomaly-contaminated HSIs tend to overfit specific anomalies, resulting in poor background characterization. To tackle these issues, this study presents a novel and robust framework for HAD called Model-Driven Deep Mixture Network (MDMN) that combines the strengths of model-driven and data-driven approaches while emphasizing interpretability. By assuming that the background, consisting of various land-covers, arises from a mixture of low-dimensional manifolds, the MDMN incorporates a novel deep mixture module to comprehensively characterize the background. This module utilizes a low-dimensional manifold learned by an AE to represent a specific category of background land-covers. To mitigate the impact of anomaly corruptions, the MDMN incorporates a convex relaxation of a sparse constraint, which helps prevent overfitting anomalies. Extensive experimental results demonstrate that the proposed MDMN offers more satisfactory and robust detection performance. Yunsong Li 0001, Kai Jiang 0001, Weiying Xie, Jie Lei 0001, Xin Zhang 0092, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | SuperYOLO: Super Resolution Assisted Object Detection in Multimodal Remote Sensing ImageryabstractAccurately and timely detecting multiscale small objects that contain tens of pixels from remote sensing images (RSI) remains challenging. Most of the existing solutions primarily design complex deep neural networks to learn strong feature representations for objects separated from the background, which often results in a heavy computation burden. In this article, we propose an accurate yet fast object detection method for RSI, named SuperYOLO, which fuses multimodal data and performs high-resolution (HR) object detection on multiscale objects by utilizing the assisted super resolution (SR) learning and considering both the detection accuracy and computation cost. First, we utilize a symmetric compact multimodal fusion (MF) to extract supplementary information from various data for improving small object detection in RSI. Furthermore, we design a simple and flexible SR branch to learn HR feature representations that can discriminate small objects from vast backgrounds with low-resolution (LR) input, thus further improving the detection accuracy. Moreover, to avoid introducing additional computation, the SR branch is discarded in the inference stage, and the computation of the network model is reduced due to the LR input. Experimental results show that, on the widely used VEDAI RS dataset, SuperYOLO achieves an accuracy of 75.09% (in terms of$\text {mA}{{\text {P}}_{{50}}}$), which is more than 10% higher than the SOTA large models, such as YOLOv5l, YOLOv5x, and RS designed YOLOrs. Meanwhile, the parameter size and GFLOPs of SuperYOLO are about$18\times $and$3.8\times $less than YOLOv5x. Our proposed model shows a favorable accuracy–speed tradeoff compared to the state-of-the-art models. The code will be open-sourced athttps://github.com/icey-zhang/SuperYOLO. Jie Lei 0001, Weiying Xie, Zhenman Fang, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Guided Hybrid Quantization for Object Detection in Remote Sensing Imagery via One-to-One Self-TeachingabstractDeep convolutional neural networks (CNNs) have improved remote sensing image analysis, but their high computational demands may limit their deployment on low-end devices with limited resources, such as intelligent satellites and unmanned aerial vehicles. Considering the computation complexity, we propose a Guided Hybrid Quantization with One-to-one Self-Teaching (GHOST) framework. More concretely, we first design a structure called guided quantization self-distillation (GQSD), an innovative idea for realizing a lightweight model through the synergy of quantization and distillation. The training process of the quantization model is guided by its full-precision model, which is time-saving and cost-saving without preparing a huge pre-trained model in advance. Second, we put forward a hybrid quantization (HQ) module that automatically acquires the optimal bit-width by imposing a threshold constraint on the distribution distance between the center point and samples in the weight search space, aiming to retain more shallow detail information that is advantageous for small object detection. Third, to improve information transformation, we propose a one-to-one self-teaching (OST) module to give the student network the ability to self-judgment. A switch control machine (SCM) builds a bridge between the student and teacher networks in the same location to help the teacher reduce wrong guidance and impart vital knowledge about objects without vast background information to the student. This distillation method allows a model to learn from itself and gain substantial improvement without any additional supervision. Extensive experiments on a multimodal dataset (VEDAI) and single-modality datasets (DOTA, NWPU, and DIOR) show that object detection based on GHOST outperforms the existing detectors. The tiny parameters (<9.7 MB) and Bit-Operations (BOPs) (<2158 G) compared with any remote sensing-based, lightweight, or distillation-based algorithms demonstrate the superiority in the lightweight design domain. Our code and model will be released at https://github.com/icey-zhang/GHOST. Jie Lei 0001, Weiying Xie, Yunsong Li 0001, Geng Yang 0001, Xiuping Jia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Transcoded Video Restoration by Temporal Spatial Auxiliary NetworkabstractIn most video platforms, such as Youtube, Kwai, and TikTok, the played videos usually have undergone multiple video encodings such as hardware encoding by recording devices, software encoding by video editing apps, and single/multiple video transcoding by video application servers. Previous works in compressed video restoration typically assume the compression artifacts are caused by one-time encoding. Thus, the derived solution usually does not work very well in practice. In this paper, we propose a new method, temporal spatial auxiliary network (TSAN), for transcoded video restoration. Our method considers the unique traits between video encoding and transcoding, and we consider the initial shallow encoded videos as the intermediate labels to assist the network to conduct self-supervised attention training. In addition, we employ adjacent multi-frame information and propose the temporal deformable alignment and pyramidal spatial fusion for transcoded video restoration. The experimental results demonstrate that the performance of the proposed method is superior to that of the previous techniques. The code is available at https://github.com/icecherylXuli/TSAN. Li Xu 0008, Gang He 0002, Jinjia Zhou, Jie Lei 0001, Weiying Xie, Yunsong Li 0001, Yu-Wing Tai |
AAAI | 4 |
| 2022 | Interlayer Restoration Deep Neural Network for Scalable High Efficiency Video CodingabstractThis paper applies an interlayer restoration deep neural network (IRDNN) for scalable high efficiency video coding (SHVC) to improve visual quality and coding efficiency. It is the first time to combine deep neural network (DNN) and SHVC. Considering the coding architecture of SHVC, we elaborate a multi-frame and multi-layer neural network to restore the interlayer of SHVC by utilizing both the adjacent reconstructed frames of the base layer (BL) and enhancement layer (EL). Moreover, we analyze the temporal motion relationship of frames in one layer and the compression degradation relationship of frames between different layers, and propose the synergistic mechanism of motion restoration and compression restoration in our IRDNN. The network can generate an interlayer with higher quality serving for the EL coding and thus enhance the coding efficiency. A large-scale and various-quality-degradation dataset is self-made for the task of interlayer restoration of SHVC. The experimental results show that with our implementation on SHVC, the EL Bj$\phi $ntegaard delta bit-rate (BD-BR) reduction is 9.291% and 6.007% in signal-to-noise ratio scalability and spatial scalability, respectively. The code is available athttps://github.com/icecherylXuli/IRDNN. Gang He 0002, Li Xu 0008, Jie Lei 0001, Weiying Xie, Yunsong Li 0001, Yibo Fan, Jinjia Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | E2E-LIADE: End-to-End Local Invariant Autoencoding Density Estimation Model for Anomaly Target Detection in Hyperspectral ImageabstractHyperspectral anomaly target detection (also known as hyperspectral anomaly detection (HAD)] is a technique aiming to identify samples with atypical spectra. Although some density estimation-based methods have been developed, they may suffer from two issues: 1) separated two-stage optimization with inconsistent objective functions makes the representation learning model fail to dig out characterization customized for HAD and 2) incapability of learning a low-dimensional representation that preserves the inherent information from the original high-dimensional spectral space. To address these problems, we propose a novel end-to-end local invariant autoencoding density estimation (E2E-LIADE) model. To satisfy the assumption on the manifold, the E2E-LIADE introduces a local invariant autoencoder (LIA) to capture the intrinsic low-dimensional manifold embedded in the original space. Augmented low-dimensional representation (ALDR) can be generated by concatenating the local invariant constrained by a graph regularizer and the reconstruction error. In particular, an end-to-end (E2E) multidistance measure, including mean-squared error (MSE) and orthogonal projection divergence (OPD), is imposed on the LIA with respect to hyperspectral data. More important, E2E-LIADE simultaneously optimizes the ALDR of the LIA and a density estimation network in an E2E manner to avoid the model being trapped in a local optimum, resulting in an energy map in which each pixel represents a negative log likelihood for the spectrum. Finally, a postprocessing procedure is conducted on the energy map to suppress the background. The experimental results demonstrate that compared to the state of the art, the proposed E2E-LIADE offers more satisfactory performance. Kai Jiang 0001, Weiying Xie, Jie Lei 0001, Zan Li 0001, Yunsong Li 0001, Tao Jiang 0031, Qian Du 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | Boundary Extraction Constrained Siamese Network for Remote Sensing Image Change DetectionabstractChange detection (CD) is crucial to the understanding of relationships and interactions among multitemporal high-resolution remote sensing (RS) images. However, various inherent attributes of images have different impacts on CD judgment. How to effectively use helpful information to improve the performance of CD is still a challenge. In this article, we present a boundary extraction constrained Siamese network (BESNet) to dig out the efficacy of boundary information. BESNet is a joint learning network in which a novel multiscale boundary extraction (MSBE) module is embedded. In this way, traditional and deep learning techniques are leveraged to learn together to maximize their respective strengths through cooperation. In particular, a new boundary extraction constrained (BEC) loss function combined with a contractive loss function is used to optimize the BESNet. Considering the interaction between various extracted features, a channel-shuffle fusion strategy is developed to exploit their complementary advantages between features. Our experiments show that the proposed BESNet can significantly improve the CD performance and generate more complete and clearer object boundaries. Experiments conducted on two real datasets over different scenes demonstrate its state-of-the-art performance. Jie Lei 0001, Yijie Gu, Weiying Xie, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Sparse Coding-Inspired GAN for Hyperspectral Anomaly Detection in Weakly Supervised LearningabstractAnomaly detection (AD) from hyperspectral images (HSIs) is of great importance in both space exploration and Earth observations. However, the challenges caused by insufficient datasets, no labels, and noise corruption substantially downgrade the accuracy of detection. To solve these problems, this article proposes a sparse coding (SC)-inspired generative adversarial network (GAN) for weakly supervised hyperspectral AD (HAD), named sparseHAD. It can learn a discriminative latent reconstruction with small errors for background pixels and large errors for anomalous ones. First, a background-category searching step is built to alleviate the difficulty of data annotation. Then, an SC-inspired regularized network is integrated into an end-to-end GAN to form a weakly supervised spectral mapping model consisting of two encoders, a decoder, and a discriminator. This model not only makes the network more robust and interpretable experimentally and theoretically but also develops a new SC-inspired path for HAD. Subsequently, the proposed sparseHAD detects anomalies in a latent space rather than the original space, which also contributes to its noise robustness. Quantitative assessments and experiments over real HSIs demonstrate the unique promise of the proposed sparseHAD. The code, data, and trained models are available athttps://github.com/JiangThea/HAD. Yunsong Li 0001, Tao Jiang 0031, Weiying Xie, Jie Lei 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Algorithm/Hardware Codesign for Real-Time On-Satellite CNN-Based Ship Detection in SAR ImageryabstractRecently, the convolutional neural network (CNN)-based approach for on-satellite ship detection in synthetic aperture radar (SAR) images has received increasing attention since it does not rely on predefined imagery features and distributions that are required in conventional detection methods. To achieve high detection accuracy, most of the existing CNN-based methods leverage complex off-the-shelf CNN models for optical imagery. Unfortunately, this usually leads to expensive computational cost, which is hard to process in real time using resource-constrained devices deployed in the harsh satellite environment. In this article, we propose OSCAR-RT, the first end-to-end algorithm/hardware codesign framework for real-time on-satellite CNN-based SAR ship detection, which can simultaneously produce an accurate and hardware-friendly CNN model and an ultraefficient field-programmable gate array (FPGA)-based hardware accelerator that can be deployed on satellites. With the real-time on-satellite processing speed in mind, we start from a state-of-the-art compact CNN model for optical imagery. To eliminate the sharp decrease in the detection accuracy for SAR imagery, we analyze the discrepancy between the SAR domain and optical domain and propose to adapt the model by adjusting the output feature size to better detect relatively smaller objects in SAR imagery. To improve the detection speed, we propose to develop a fully pipelined interlayer streaming accelerator architecture, where all the layers of the CNN model can be concurrently processed using on-chip FPGA resources. To achieve this architecture, we first propose a hardware-guided, progressive, and structural pruning strategy, which is guided by our modeled hardware metrics and applies state-of-the-art coarse-grained and fine-grained filter pruning as well as mixed-precision quantization techniques. Moreover, to improve the reusability and portability of the hardware accelerator design, we develop a library of highly optimized CNN components in high-level synthesis, together with their performance and resource models. Finally, we map the pruned CNN model onto these hardware library components in a fully pipelined interlayer streaming fashion, by adjusting their parallelism factors to balance the execution of each layer and fit into the resource constraint. Experimental results using the adapted MobileNetV1, MobileNetV2, and SqueezeNet models on the widely used SAR ship detection dataset (SSDD) demonstrate the effectiveness of OSCAR-RT; for the MobileNetV1 model, it achieves an average precision of 94%, a detection speed of 652 frames/s on the Xilinx VC709 FPGA evaluation board while consuming about 5.8-W power. Geng Yang 0001, Jie Lei 0001, Weiying Xie, Zhenman Fang, Yunsong Li 0001, Xin Zhang 0092 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Rank-Aware Generative Adversarial Network for Hyperspectral Band SelectionabstractTraditional clustering-based band selection (BS) methods treat each band as individuals, and selection is conducted by enlarging the difference between clusters, which leads to the loss of band interaction and information saliency evaluation. In this article, we propose a BS method named rank-aware generative adversarial network (R-GAN) to address these problems. First, centralized reference feature extraction (FE) with GAN aids R-GAN to combine interpretability and interband relevance. Then, the reference feature is refined with the saliency estimation provided by the rank-aware strategy. According to data characteristics, there are two versions of rank computation including tensor and matrix. Finally, the structural similarity index measurement (SSIM) maps the saliency to the original data space to obtain the final BS result. Extensive comparison experiments with popular existing BS approaches on five hyperspectral images (HSIs) datasets show that the proposed R-GAN can address spectral saliency effectively and select more informative band subsets, which outperforms other competitors for both detection and classification tasks. For example, on the SD-1 dataset, the ten bands selected by R-GAN achieve 0.982 ± 0.003 with an improvement of 13.7% in the area under the curve (AUC) value of anomaly detection performance. The peaked accuracy surpasses the baseline by 0.46% for the classification on the PaviaU dataset. Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001, Geng Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Weakly Supervised Discriminative Learning With Spectral Constrained Generative Adversarial Network for Hyperspectral Anomaly DetectionabstractAnomaly detection (AD) using hyperspectral images (HSIs) is of great interest for deep space exploration and Earth observations. This article proposes a weakly supervised discriminative learning with a spectral constrained generative adversarial network (GAN) for hyperspectral anomaly detection (HAD), called weaklyAD. It can enhance the discrimination between anomaly and background with background homogenization and anomaly saliency in cases where anomalous samples are limited and sensitive to the background. A novel probability-based category thresholding is first proposed to label coarse samples in preparation for weakly supervised learning. Subsequently, a discriminative reconstruction model is learned by the proposed network in a weakly supervised fashion. The proposed network has an end-to-end architecture, which not only includes an encoder, a decoder, a latent layer discriminator, and a spectral discriminator competitively but also contains a novel Kullback-Leibler (KL) divergence-based orthogonal projection divergence (OPD) spectral constraint. Finally, the well-learned network is used to reconstruct HSIs captured by the same sensor. Our work paves a new weakly supervised way for HAD, which intends to match the performance of supervised methods without the prerequisite of manually labeled data. Assessments and generalization experiments over real HSIs demonstrate the unique promise of such a proposed approach. Tao Jiang 0031, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | LREN: Low-Rank Embedded Network for Sample-Free Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) is a challenging task because it explores the intrinsic structure of complex high-dimensional signals without any samples at training time. Deep neural networks (DNNs) can dig out the underlying distribution of hyperspectral data but are limited by the labeling of large-scale hyperspectral datasets, especially the low spatial resolution of hyperspectral data, which makes labeling more difficult. To tackle this problem while ensuring the detection performance, we present an unsupervised low-rank embedded network (LREN) in this paper. LREN is a joint learning network in which the latent representation is specifically designed for HAD, rather than merely as a feature input for the detector. And it searches the lowest rank representation based on a representative and discriminative dictionary in the deep latent space to estimate the residual efficiently. Considering the physically mixing properties in hyperspectral imaging, we develop a trainable density estimation module based on Gaussian mixture model (GMM) in the deep latent space to construct a dictionary that can better characterize the complex hyperspectral images (HSIs). The closed-form solution of the proposed low-rank learner surpasses existing approaches on four real hyperspectral datasets with different anomalies. We argue that this unified framework paves a novel way to combine feature extraction and anomaly estimation-based methods for HAD, which intends to learn the underlying representation tailored for HAD without the prerequisite of manually labeled data. Code available at https://github.com/xdjiangkai/LREN. Kai Jiang 0001, Weiying Xie, Jie Lei 0001, Tao Jiang 0031, Yunsong Li 0001 |
AAAI | 3 |
| 2021 | PTGAN: A Proposal-Weighted Two-Stage GAN with Attention for Hyperspectral Target DetectionabstractIn this paper, a proposal-weighted two-stage generative adversarial network (GAN) with attention mechanism is proposed for hyperspectral target detection (HTD). PTGAN leverages GAN to estimate spectral background distribution and realize mapping from the latent space to the spectral space. Meanwhile, PTGAN conducts the reversed mapping through latent-spectral-latent and spectral-latent-spectral learning. On this basis, PTGAN implements accurate reconstruction of background spectrum via latent space. Therefore, targets of interest can be detected through larger pixel-level reconstruction error. In particular, the variance attention module is designed to make full use of global information among spectral bands to selectively emphasize channel-wise spectral features. Furthermore, a proposal-weighted strategy in a two-stage manner reduces the false alarm of detection by refining the previous detection proposal. Finally, exponential nonlinear fusion combines the discriminative feature from two stages to suppress the background. Extensive experiments on two real hyperspectral images (HSIs) verify the effectiveness of PTGAN. Weiying Xie, Yunsong Li 0001, Kai Jiang 0001, Jie Lei 0001, Qian Du 0001 |
IGARSS | 5 |
| 2021 | Spectral mapping with adversarial learning for unsupervised hyperspectral change detection
Jie Lei 0001, Meiqi Li, Weiying Xie, Yunsong Li 0001, Xiuping Jia |
Neurocomputing | 1 |
| 2021 | A Specially Optimized One-Stage Network for Object Detection in Remote Sensing ImagesabstractWith great significance in military and civilian applications, detecting indistinguishable small objects in wide-scale remote sensing images is still a challenging topic. In this letter, we propose a specially optimized one-stage network (SOON) focusing on extracting spatial information of high-resolution images by understanding and analyzing the combination of feature and semantic information of small objects. The SOON model consists of feature enhancement, multiscale detection, and feature fusion. The first part is implemented by constructing a receptive field enhancement (RFE) module and incorporating it into the network's specific parts where the information of small objects mainly exists. The second part is achieved by four detectors with different sensitivities, which access to the fused and enhanced features to enable the network to make full use of features in different scales. The third part consolidates the high-level and low-level features by adopting upsampling, concatenation, and convolution operations to build a feature pyramid structure, which explicitly yields strong feature representation and semantic information. In addition, we introduce the soft-nonmaximum suppression to preserve accurate bounding boxes in the postprocessing stage for densely arranged objects. Note that the split and merge strategy and the multiscale training strategy are employed. Extensive experiments and thorough analysis are performed on the NorthWestern Polytechnical University Very-High-Resolution (NWPU VHR)-10-v2 data set and the airplane, car and ship (ACS) data set as compared with several state-of-the-art methods. The satisfactory performance in experiments verifies the effectiveness of the design and optimization. Yunsong Li 0001, Jie Lei 0001, Weiying Xie |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Self-spectral learning with GAN based spectral-spatial target detection for hyperspectral image
Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Xiuping Jia |
Neural Networks | 3 |
| 2021 | Dual feature extraction network for hyperspectral image analysis
Weiying Xie, Jie Lei 0001, Shuo Fang, Yunsong Li 0001, Xiuping Jia, Mingsuo Li |
Pattern Recognit. | 2 |
| 2021 | Weakly Supervised Low-Rank Representation for Hyperspectral Anomaly DetectionabstractIn this article, we propose a weakly supervised low-rank representation (WSLRR) method for hyperspectral anomaly detection (HAD), which formulates deep learning-based HAD into a low-lank optimization problem not only characterizing the complex and diverse background in real HSIs but also obtaining relatively strong supervision information. Different from the existing unsupervised and supervised methods, we first model the background in a weakly supervised manner, which achieves better performance without prior information and is not restrained by richly correct annotation. Considering reconstruction biases introduced by the weakly supervised estimation, LRR is an effective method for further exploring the intricate background structures. Instead of directly applying the conventional LRR approaches, a dictionary-based LRR, including both observed training data and hidden learned data drawn by the background estimation model, is proposed. Finally, the derived low-rank part and sparse part and the result of the initial detection work together to achieve anomaly detection. Comparative analyses validate that the proposed WSLRR method presents superior detection performance compared with the state-of-the-art methods. Weiying Xie, Xin Zhang 0092, Yunsong Li 0001, Jie Lei 0001, Jiaojiao Li 0001, Qian Du 0001 |
IEEE Trans. Cybern. | 4 |
| 2021 | HPGAN: Hyperspectral Pansharpening Using 3-D Generative Adversarial NetworksabstractHyperspectral (HS) pansharpening, as a special case of the superresolution (SR) problem, is to obtain a high-resolution (HR) image from the fusion of an HR panchromatic (PAN) image and a low-resolution (LR) HS image. Though HS pansharpening based on deep learning has gained rapid development in recent years, it is still a challenging task because of the following requirements: 1) a unique model with the goal of fusing two images with different dimensions should enhance spatial resolution while preserving spectral information; 2) all the parameters should be adaptively trained without manual adjustment; and 3) a model with good generalization should overcome the sensitivity to different sensor data in reasonable computational complexity. To meet such requirements, we propose a unique HS pansharpening framework based on a 3-D generative adversarial network (HPGAN) in this article. The HPGAN induces the 3-D spectral-spatial generator network to reconstruct the HR HS image from the newly constructed 3-D PAN cube and the LR HS image. It searches for an optimal HR HS image by successive adversarial learning to fool the introduced PAN discriminator network. The loss function is specifically designed to comprehensively consider global constraint, spectral constraint, and spatial constraint. Besides, the proposed 3-D training in the high-frequency domain reduces the sensitivity to different sensor data and extends the generalization of HPGAN. Experimental results on data sets captured by different sensors illustrate that the proposed method can successfully enhance spatial resolution and preserve spectral information. Weiying Xie, Yuhang Cui, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001, Jiaojiao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Characterization of Background-Anomaly Separability With Generative Adversarial Network for Hyperspectral Anomaly DetectionabstractHyperspectral images (HSIs) have unique advantages in distinguishing subtle spectral differences of different materials. However, due to complex and diverse backgrounds, unknown prior knowledge, and imbalanced samples, it is challenging to separate background and anomaly. In this article, we present a novel characterization of background-anomaly separability with a generative adversarial network (BASGAN) for hyperspectral anomaly detection. The key contribution is the proposal to explicitly constrain the background and anomaly separability by characterizing background spectral samples while avoiding anomaly reconstruction. First, we use a class saliency map extraction algorithm to obtain pseudobackground and anomaly samples for adversarial training. To further mitigate the suffering of anomaly contamination in background distribution estimation, we introduce background-anomaly separability constrained loss function to enhance the reconstruction of the background while weakening the anomaly reconstruction in a semisupervised way. Additionally, a discriminator is induced into the latent space to make the encoded representation resemble Gaussian distribution during adversarial training. The other is adversarial training in the reconstruction space so that the background estimation can be improved. Experiments conducted on real data sets illustrate the superior background-anomaly separability of the proposed method. Jiaping Zhong, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Unsupervised spectral mapping and feature selection for hyperspectral anomaly detection
Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Jiaojiao Li 0001, Xiuping Jia |
Neural Networks | 3 |
| 2020 | Semisupervised Spectral Learning With Generative Adversarial Network for Hyperspectral Anomaly DetectionabstractLimited by the anomalous spectral vectors in unlabeled hyperspectral images (HSIs), anomaly detection methods based on background distribution estimation often suffer from the contamination of anomalies, which decreases the estimation accuracy and, thus, weakens the detection performance. To address this problem, we proposed a novel semisupervised spectral learning (SSL) for the hyperspectral anomaly detection framework based on the generative adversarial network (GAN). GAN is applied and developed to estimate the background distribution in a semisupervised manner and obtain an initial spectral feature because of its strong representational capability and adversarial training advantage. In the proposed framework, an initial spatial feature is generated via morphological attribute filtering. Finally, an exponential constrained nonlinear suppression fusion technique is adopted to suppress the background and combine the complementary information in different features to obtain a fused detection map. The performance of the proposed anomaly detection technique is evaluated on a series of HSIs. Experimental results demonstrate that our method can outperform state-of-the-art anomaly detection methods. Kai Jiang 0001, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Gang He 0002, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Discriminative Reconstruction for Hyperspectral Anomaly Detection With Spectral LearningabstractRecently, autoencoder (AE)-based anomaly detection has drawn considerable interest in hyperspectral image (HSI) analysis. In this article, we propose a novel discriminative reconstruction method for hyperspectral anomaly detection images with spectral learning (SLDR). The proposed algorithm has the following innovations. First, we use the spectral error map (SEM) to detect anomalies because the SEM can preferably reflect the spectral similarity of each pixel between the input and the reconstruction. Second, the loss function of the proposed SLDR model additionally introduces the spectral angle distance (SAD), which constrains the model to generate a reconstruction having greater spectral similarity to the input. Third, a constraint is imposed on the encoder, forcing it to generate latent variables that obey a unit Gaussian distribution, which helps the decoder to reconstruct a better background with respect to the input. Compared with the Reed-Xiaoli (RX), collaborative representation detection (CRD), attribute and edge-preserving filtering-based anomaly detection (AED) and adversarial autoencoder-based anomaly detection (AAE), through two real HSI data sets, the detection performance of the proposed SLDR method is found to be competitive. Jie Lei 0001, Shuo Fang, Weiying Xie, Yunsong Li 0001, Chein-I Chang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Spectral Adversarial Feature Learning for Anomaly Detection in Hyperspectral ImageryabstractTheoretically, hyperspectral images (HSIs) are capable of providing subtle spectral differences between different materials, but in fact, it is difficult to distinguish between background and anomalies because the samples of anomalous pixels in HSIs are limited and susceptible to background and noise. To explore the discriminant features, a spectral adversarial feature learning (SAFL) architecture is specially designed for hyperspectral anomaly detection in this article. In addition to reconstruction loss, SAFL also introduces spectral constraint loss and adversarial loss in the network with batch normalization to extract the intrinsic spectral features in deep latent space. To further reduce the false alarm rate, we present an iterative optimization approach by a weighted suppression function that depends on the contribution rate of each feature to the detection. In particular, the structure tensor matrix is adopted to adaptively calculate the contribution rate of each feature. Benefiting from these improvements, the proposed method is superior to the typical and state-of-the-art methods either in detection probability or false alarm rate. Weiying Xie, Baozhu Liu, Yunsong Li 0001, Jie Lei 0001, Chein-I Chang, Gang He 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Autoencoder and Adversarial-Learning-Based Semisupervised Background Estimation for Hyperspectral Anomaly DetectionabstractReliable detection of anomalies without any prior information is a critical yet challenging task in many applications, not least military and civilian fields. An intelligent anomaly detection system would use the material-specific spectral information in hyperspectral images (HSIs), thereby avoiding the loss of visually confusing objects. However, conventional hyperspectral anomaly detection methods are mainly achieved in an unsupervised way leading to limited performance due to lack of prior knowledge. In this article, we propose a novel autoencoder and adversarial-learning based semisupervised background estimation model (SBEM) that is trained only on the background spectral samples in order to accurately learn the background distribution. In particular, an unsupervised background searching method is firstly conducted on the original HSIs to search the background spectral samples. Our proposed SBEM consists of an encoder, a decoder, and a discriminator to thoroughly capture background distribution. Furthermore, jointly minimizing the reconstruction loss, spectral loss, and adversarial loss during training aids the model to learn the background distribution as required. Experiments on four real HSIs demonstrate that compared to the current state-of-the-art, the proposed framework yields higher detection capability and lower false alarm rate, which shows that it has a significant benefit in the tradeoff between detection accuracy and false alarm rate. Weiying Xie, Baozhu Liu, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Hyperspectral Band Selection for Spectral-Spatial Anomaly DetectionabstractOwing to significantly improved spectral resolution, a hyperspectral imaging sensor can now uncover many unknown subtle material substances. In many cases, anomalies are usually embedded in the background. To develop a means through which these anomalies may be detected and separated from the background, we propose a spectral-spatial anomaly detection method based on a selected band subset. To be specific, we constrain an unsupervised network by making full use of the underlying physical characteristics which are beneficial to hyperspectral anomaly detection. Based on that, a selection criterion is constructed to adaptively select a subset of bands that essentially contain discriminative and informative features between the anomaly and background in an unsupervised manner. Then, the selected bands are simultaneously inputted into the spatial detector and spectral detector. To overcome the deficiencies of detecting anomalies in only one aspect, an adaptive combination of spatial result and the spectral result is introduced. Finally, a simple and powerful iterative suppression is conducted on the initial detection map to further reduce false alarm rate while ensuring detection capability. Extensive empirical researches performed on eighteen publicly available hyperspectral images (HSIs) of different sizes over different scenes demonstrate that our proposed method can achieve an average detection capability of 0.99564, and the average false alarm rate is one order of magnitude lower than the second one. Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Chein-I Chang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Deep Latent Spectral Representation Learning-Based Hyperspectral Band Selection for Target DetectionabstractHyperspectral images (HSIs) can provide discriminative spectral signatures regarding the physical nature of different materials. It is this unique nature that makes HSIs to be of great interest in many fields. However, HSI application faces various challenges due to high dimensionality, redundant information, noisy bands, and insufficient samples. To address these problems, we propose an unsupervised band selection method based on deep latent spectral representation learning, called DLSRL, in this article. It imposes spectral consistency on deep latent space that resolves the issue of insufficient samples and spectral information lost in HSI interpretation. It pursues the low-dimensional optimal representation of the high-dimensional HSIs. In particular, an adaptive mapping relationship is constructed between the deep latent representation and the optimal subset to preserve physical significance optimally. Furthermore, a hierarchical optimization approach is introduced to achieve target detection with the selected subset. To verify the superiority of the proposed method, experiments have been conducted on four data sets captured by different sensors over different scenes. Comparative analyses validate that the proposed method presents superior performance in terms of high detection accuracy and low false alarm rate. Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | SRUN: Spectral Regularized Unsupervised Networks for Hyperspectral Target DetectionabstractThe high dimensionality of a hyperspectral image (HSI) provides the possibility of deeply capturing the underlying and intrinsic characteristics in spectra, such that targets embedded in the background can be detected. However, redundant information, deteriorated bands, and other interferences from background challenge the target detection problem. In this article, an effective feature extraction method based on unsupervised networks is proposed to mine intrinsic properties underlying HSIs. Our approach, called spectral regularized unsupervised networks (SRUN), imposes spectral regularization on autoencoder (AE) and variational AE (VAE) to emphasize spectral consistency, which is more suitable for characterizing spectral information of HSIs by hidden nodes than the original AE and VAE models. Then, we conduct a simple feature selection algorithm on the hidden nodes in the deepest code to select specific nodes that contain distinguishability between target and background, which is based on the spectral angular difference between a known target spectrum and spectra of other pixels in input. The selected nodes are further weighted adaptively to obtain a discriminative map depending on the observation that each selected node provides different contribution rates to target detection. Experimental results on several data sets illustrate that the proposed SRUN-based target detection algorithm is suitable for targets at the subpixel level and those with structural information. Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Qian Du 0001, Gang He 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Hyperspectral Pansharpening With Deep PriorsabstractHyperspectral (HS) image can describe subtle differences in the spectral signatures of materials, but it has low spatial resolution limited by the existing technical and budget constraints. In this paper, we propose a promising HS pansharpening method with deep priors (HPDP) to fuse a low-resolution (LR) HS image with a high-resolution (HR) panchromatic (PAN) image. Different from the existing methods, we redefine the spectral response function (SRF) based on the larger eigenvalue of structure tensor (ST) matrix for the first time that is more in line with the characteristics of HS imaging. Then, we introduce HFNet to capture deep residual mapping of high frequency across the upsampled HS image and the PAN image in a band-by-band manner. Specifically, the learned residual mapping of high frequency is injected into the structural transformed HS images, which are the extracted deep priors served as additional constraint in a Sylvester equation to estimate the final HR HS image. Comparative analyses validate that the proposed HPDP method presents the superior pansharpening performance by ensuring higher quality both in spatial and spectral domains for all types of data sets. In addition, the HFNet is trained in the high-frequency domain based on multispectral (MS) images, which overcomes the sensitivity of deep neural network (DNN) to data sets acquired by different sensors and the difficulty of insufficient training samples for HS pansharpening. Weiying Xie, Jie Lei 0001, Yuhang Cui, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | SOON: Specifically Optimized One-Stage Network for Object Detection in Remote Sensing ImageryabstractWith great significance in military and civilian applications, detecting indistinguishable small objects in wide-scale remote sensing images is still a challenging topic. In this work, we propose a specially optimized one-stage network (SOON) focusing on extracting spatial information of high-resolution images by understanding and analyzing the combination of feature and semantic information of small objects, which consists of feature enhancement, multi-scale detection, and feature fusion. The first part is implemented by constructing a receptive field enhancement (RFE) module and incorporating it into the specific parts of the network where the information of small objects mainly exists. The second part is achieved by four detectors with different sensitivities accessing to the fused and enhanced features, which enables the network to make full use of features in different scales. The third part consolidates the high-level and low-level features by adopting up-sampling, concatenation and convolution operations to build a feature pyramid structure, which explicitly yields strong feature representation and semantic information. In addition, we introduce the Soft-NMS to preserve accurate bounding boxes in the post-processing stage for densely arranged objects. Note that the split and merge strategy, as well as the multi-scale training strategy, are employed in this work. Extensive experiments and thorough analysis are performed on the NWPU VHR-10-v2 dataset and the ACS dataset as compared with several state-of-the-art methods, in which satisfactory performance verifies the effectiveness of the design and optimization. The code will be released for reproduction. Yunsong Li 0001, Jie Lei 0001, Weiying Xie |
ICTAI | 4 |
| 2019 | Discriminative Feature Learning With Distance Constrained Stacked Sparse Autoencoder for Hyperspectral Target DetectionabstractTarget detection (TD) is one of the major tasks in hyperspectral image (HSI) processing, and its performance is greatly affected by the background. Feature extraction (FE) has been an effective way to mine discriminative information, especially FE based on deep learning, which can learn the intrinsic properties of data to further improve the detection performance. Unlike supervised networks, unsupervised stacked sparse autoencoders (SSAEs) can learn deep and nonlinear features without any labeled data. However, SSAEs usually require a supervised fine-tuned model to obtain better discrimination, which is not feasible for TD, since the prior information is generally insufficient. In this letter, we introduce a distance constraint that is added to the SSAE to form a new distance constrained SSAE (DCSSAE) network. Specifically, the distance constraint maximizes the distinction between the target pixels and other background pixels in the feature space. Then, using the discriminative features learned from the DCSSAE, a simple detector using radial basis function kernel is derived for background suppression. Experiments on two HSIs demonstrate that the deep spectral features learned from the DCSSAE are more distinguishable, and our proposed detector, namely, the DCSSAE detector, outperforms several popular detectors, especially in background suppression. Yanzi Shi, Jie Lei 0001, Yaping Yin, Kailang Cao, Yunsong Li 0001, Chein-I Chang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2019 | Spectral constraint adversarial autoencoders approach to feature representation in hyperspectral anomaly detection
Weiying Xie, Jie Lei 0001, Baozhu Liu, Yunsong Li 0001, Xiuping Jia |
Neural Networks | 2 |
| 2019 | High-quality spectral-spatial reconstruction using saliency detection and deep feature enhancement
Weiying Xie, Yanzi Shi, Yunsong Li 0001, Xiuping Jia, Jie Lei 0001 |
Pattern Recognit. | 5 |
| 2019 | Spectral-Spatial Feature Extraction for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection faces various levels of difficulty due to the high dimensionality of hyperspectral images (HSIs), redundant information, noisy bands, and the limited capability of utilizing spectral-spatial information. In this paper, we address these problems and propose a novel approach, called spectral-spatial feature extraction (SSFE), which is based on two main aspects. In the spectral domain, we assume that the anomalous pixels are rarely present and all (or most) of the samples around the anomalies belong to background (BKG). Using this fact, we introduce a suppression function to construct a discriminative feature space and utilize a deep brief network to learn spectral representation and abstraction automatically that are used as inputs to the Mahalanobis distance (MD)-based detector. In the spatial domain, the anomalies appear as a small area grouped by pixels with high correlation among them compared to BKG. Therefore, the objects appearing as a small area are extracted based on attribute filtering, and a guided filter is further employed for local smoothness. More specifically, we extract spatial features of anomalies only from one single band obtained by fusing all bands in the visible wavelength range. Finally, we detect anomalies by jointly considering the spectral and spatial detection results. Several experiments are performed, which show that our proposed method outperforms the state-of-the-art methods. Jie Lei 0001, Weiying Xie, Yunsong Li 0001, Chein-I Chang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Structure Tensor and Guided Filtering-Based Algorithm for Hyperspectral Anomaly DetectionabstractAnomaly detection is one of the most important applications of hyperspectral imaging technology. It is a challenging task due to the high dimensionality of hyperspectral images (HSIs), redundant information, noisy bands, and the limited capability of utilizing spatial information. In this paper, we address these problems and propose a novel anomaly detection method in HSIs. Our approach, called structure tensor and guided filter (STGF)-based strategy for anomaly detection, is based on the characteristics of HSIs. First, a novel band selection algorithm is proposed to reduce dimension, remove noisy bands, and select bands with effective information. Second, the selected bands are decomposed into two parts according to the characteristics of anomalies that are usually in a small area. Followed by this step, the backgrounds are removed through a simple differential operation for each selected band. Considering that not all of the bands provide the same contributions to anomaly detection, we then fuse the differential maps by a novel adaptive weighting method to obtain an initial detection map. Finally, GF is conducted to rectify the previous map under the condition that the neighboring pixels usually have quite strong correlations with each other. Experiments have been conducted on real-scene remote sensing HSI. Comparative analyses validate that the proposed STGF method presents superior performance in terms of detection accuracy and computational time. Weiying Xie, Tao Jiang 0031, Yunsong Li 0001, Xiuping Jia, Jie Lei 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | Hyperspectral Image Super-Resolution Using Deep Feature Matrix FactorizationabstractHyperspectral images (HSIs) can describe the subtle differences in the spectral signatures of materials. However, they have low spatial resolution due to various hardware limitations. Improving it via postprocess without an auxiliary high-resolution (HR) image still remains a challenging problem. In this paper, we address this problem and propose a new HSI super-resolution (SR) method. Our approach, called deep feature matrix factorization (DFMF), blends feature matrix extracted by a deep neural network (DNN) with nonnegative matrix factorization strategy for super-resolving real-scene HSI. The estimation of the HR HSI is formulated as a combination of latent spatial feature matrix and spectral feature matrix. In the DFMF model, the input low-resolution (LR) HSI is first partitioned into several subsets according to the correlation matrix, and the key band is selected from each subset. Then, the key band group is super-resolved by a DNN model, and the HR key band group is then used as a guide to carry out deep spatial feature matrix. Specifically, the input LR HSI with prototype reflectance spectral vectors of the scene will be preserved when super-resolving in a spatial domain. Thus, the nonnegative spectral and spatial feature matrices are extracted simultaneously from alternately factorizing the pair of LR HSI and the HR key band group. Finally, the HR HSI is obtained by the integration of the spectral and spatial feature matrices. Experiments have been conducted on real-scene remote sensing HSI. Comparative analyses validate that the proposed DFMF method presents a superior super-resolving performance, as it preserves spectral information better. Weiying Xie, Xiuping Jia, Yunsong Li 0001, Jie Lei 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2016 | When Spark Meets FPGAs: A Case Study for Next-Generation DNA Sequencing AccelerationabstractFPGA-enabled datacenters have shown great potential for providing performance and energy efficiency improvement, and captured a great amount of attention from both academia and industry. In this paper we aim to answer one key question: how can we efficiently integrate FPGAs into state-of-the-art big-data computing frameworks? Although very important, this problem has not been well studied, especially for the integration of fine-grained FPGA accelerators that have short execution time but will be invoked many times. To provide a generalized methodology and insight for efficient integration, we conduct an in-depth analysis of challenges and corresponding solutions of integration at single-thread, single-node multi-thread, and multi-node levels. With a step-by-step case study for the next-generation DNA sequencing application, we demonstrate how a straightforward integration with 1000x slowdown can be tuned into an efficient integration with 2.6x overall system speedup and 2.4x energy efficiency improvement. Yuting Chen 0003, Jason Cong, Zhenman Fang, Jie Lei 0001, Peng Wei 0004 |
FCCM | 4 |
| 2016 | A High-throughput Architecture for Lossless Decompression on FPGA Designed Using HLS (Abstract Only)abstractIn the field of big data applications, lossless data compression and decompression can play an important role in improving the data center's efficiency in storage and distribution of data. To avoid becoming a performance bottleneck, they must be accelerated to have a capability of high speed data processing. As FPGAs begin to be deployed as compute accelerators in the data centers for its advantages of massive parallel customized processing capability, power efficiency and hardware reconfiguration. It is promising and interesting to use FPGAs for acceleration of data compression and decompression. The conventional development of FPGA accelerators using hardware description language costs much more design efforts than that of CPUs or GPUs. High level synthesis (HLS) can be used to greatly improve the design productivity. In this paper, we present a solution for accelerating lossless data decompression on FPGA by using HLS. With a pipelined data-flow structure, the proposed decompression accelerator can perform static Huffman decoding and LZ77 decompression at a very high throughput rate. According to the experimental results conducted on FPGA with the Calgary Corpus data benchmark, the average data throughput of the proposed decompression core achieves to 4.6 Gbps while running at 200 MHz. Jie Lei 0001, Yuting Chen 0003, Yunsong Li 0001, Jason Cong |
FPGA | 1 |
| 2015 | A Novel High-Throughput Acceleration Engine for Read AlignmentabstractThe Smith-Waterman (S-W) algorithm is widely adopted by the state-of-the-art DNA sequence aligners. Existing wave front-based methods ignored the fact that the S-W algorithm is fed with significantly varied-size inputs in modern aligners, in which the S-W algorithm is further optimized by exerting extensive pruning. In this paper, we propose an architecture, tailored for varied input sizes as well as harnessing software pruning strategies, to accelerate S-W. Our implementation demonstrates a 26.4x speedup over a 24-thread Intel Has well Xeon server, and outperforms wave front-based implementations by up to 6x with the same FPGA resource. Yuting Chen 0003, Jason Cong, Jie Lei 0001, Peng Wei 0004 |
FCCM | 3 |