EDBT 2026 Demo / reviewers in the wild / expert
Ran Tao 0003
dblp:99/955-3
· DBLP profile ↗
251ranked-venue papers
9as first author
153since 2021 · last 2027
0000-0002-5243-7189ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 128 · 4 first-author · 78 since 2021Graphics, computer vision, multimedia, augmented reality and games · 77 · 3 first-author · 43 since 2021Artificial intelligence and machine learning · 35 · 1 first-author · 26 since 2021Computer networks · 9 · 5 since 2021Security and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Learning tensor correlation filter with fused low-rank and smoothness priors for hyperspectral video object tracking
Wen-Shuai Hu, Jian-Li Wang, Ran Tao 0003, Qian Du 0001 |
Expert Syst. Appl. | 5 |
| 2026 | Single-UAV Synthetic-Aperture Passive Localization via Spatial Vector Backprojection
Hao Huan, Ran Tao 0003 |
WCNC | 3 |
| 2026 | ORSI Salient Object Detection via Progressive Interaction and Saliency-Guided EnhancementabstractAlthough great progress has been made in Salient Object Detection (SOD) in Optical Remote Sensing Images (ORSIs), it still faces critical challenges, particularly in dealing with irregular topological structures and complex contextual relationships. To address these issues, we propose a Progressive Interaction and Saliency-guided Enhancement Network (PISENet). Specifically, a Progressive Interaction Encoder (PIE) is proposed, which adopts a dual-path heterogeneous fusion architecture and a hierarchical progressive interaction mechanism to capture global irregular topological structures and local fine-grained image details. Meanwhile, it mitigates semantic gaps across multi-scale features and achieves effective cross-level feature fusion. Subsequently, a Global Context Enhancement Module (GCEM) is designed, which incorporates non-local blocks to enhance spatial correlations among features. A parallel multi-branch structure is further utilized to capture multi-level contextual information ranging from local details to long-range semantics, thereby strengthening the modeling of global context. Finally, a Multi-scale Progressive Attention Enhancement Decoder (MPAED) is devised, which adopts a saliency-guided attention mechanism to jointly model spatial and channel-wise dependencies, enhance responses in salient regions and boundaries, and progressively decode and aggregate deep semantic and shallow detailed features. Extensive experiments on three benchmark datasets demonstrate that our method achieves significant superiority over state-of-the-art approaches. Yunzuo Zhang, Liye Xue, Weiqi Lian, Ran Tao 0003 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2026 | HIMO: Cross-Arbitrary-Modality Image Invariant Feature Transform With Hierarchical Intrinsic Major OrientationabstractInvariant feature extraction is a critical challenge in intelligent image processing, particularly with the rapid advancement of multi-source/modal imaging. Cross-modal matching has attracted considerable attention, yet current studies primarily focus on targeted modalities rather than realizing a general approach. In this paper, cross-arbitrary-modal image invariant feature extraction and matching is studied. Inspired by human vision, a purely handcrafted invariant feature transform is proposed for universal cross-modal image matching, named Hierarchical Intrinsic Major Orientation (HIMO). Based on orientation information, a full-chain non-data-driven algorithm is designed that hinges on an Intrinsic Major Orientation (IMO) extraction. The HIMO incorporates a novel keypoint detector utilizing Difference-of-Feature Suppression (DoFS), a Polar-Pyramid descriptor (PolarP), and a Cascaded Dynamic Multi-scale Strategy (CDMS) to effectively address common challenges such as intensity distortion, rotation, scale differences, geometric deformation, and image noise. To validate the proposed method, two massive cross-modal datasets-General Cross-modal Zone (GCZ) and Wide-area Diverse Sources (WDS)-are introduced, alongside two practical evaluation metrics. Comprehensive experiments compared with 10 traditional and 15 deep-learning state-of-the-art algorithms on 5 datasets fully demonstrate that the proposed HIMO achieves superior performance in terms of robustness, stability, and generalization across diverse imaging conditions. Chenzhong Gao, Wei Li 0032, Desheng Weng, Ran Tao 0003, Xiang-Gen Xia 0001, Qian Du 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Alliance: All-in-One Spectral-Spatial-Frequency Awareness Foundation ModelabstractFrequency domain analysis reveals fundamental image patterns difficult to observe in raw pixel values, while avoiding redundant information in original image processing. Although recent remote sensing foundation models (FMs) have made progress in leveraging spatial and spectral information, they have limitations in fully utilizing frequency characteristics that capture hidden features. Existing FMs that incorporate frequency properties often struggle to maintain connections with the original image content, creating a semantic gap that affects downstream performance. To address these challenges, we propose the All-in-One Spectral-Spatial-Frequency Awareness Foundation Model (Alliance), a framework that effectively integrates information across all three domains. Alliance introduces several key innovations: (1) a progressive frequency decoding mechanism inspired by human visual cognition that minimizes multi-domain information gaps while preserving connections between general image information and frequency characteristics, progressively reconstructing from low to mid to high frequencies to extract patterns difficult to observe in raw pixel values; (2) a triple-domain fusion attention module that separately processes amplitude, phase, and spectral-spatial relationships for comprehensive feature integration; and (3) frequency embedding with frequency-aware Cls token initialization and frequency-specific mask token initialization that achieves fine-grained modeling of different frequency band information. Additionally, to evaluate FMs generalizability, we construct the Yellow River dataset, a large-scale multi-temporal collection that introduces challenging cross-domain tasks and establishes more rigorous standards for FMs assessment. Extensive experiments across six downstream tasks demonstrate Alliance's superior performance. Wei Li 0032, Yuxiang Zhang 0005, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Synchrosqueezing in rotated time-Frequency domain
Xiaoping Liu 0005, Gong Chen 0003, Jun Shi 0003, Ran Tao 0003 |
Signal Process. | 4 |
| 2026 | On minimization/maximization of the generalized multi-order complex quadratic form with constant-modulus constraints
Chunxuan Shi, Yongzhe Li, Ran Tao 0003 |
Signal Process. | 3 |
| 2026 | Dual-Domain Fractional Fourier Transformer for Underwater Image Degradation RemovalabstractIn this letter, we propose DEFriT, a dual-domain framework for underwater image restoration targeting degradation caused by scattering and absorption. Supervised training is enabled by simulating underwater degradation from land images using a physics-based formation model. To address the spectrally non-stationary nature of underwater attenuation, we introduce the fractional Fourier transform (FrFT) to bridge spatial and spectral representations. A Fractional Transform Block (FrTB) with real-imaginary decomposition and Hybrid Time-Frequency Self-Attention (HTFSA) supports joint spatial-spectral modeling. Empirical analysis shows that degradation mainly affect the amplitude spectrum, with learned fractional orders concentrated in a narrow range ($p \in [0.48, 0.52]$). To the best of our knowledge, this is the first work to introduce FrFT into underwater image restoration and to systematically establish a fractional order prior. DEFriT achieves state-of-the-art performance on five benchmark datasets. Chuangxi Chen, Xudong Zhao 0003, Yixiao Yang, Ran Tao 0003, Binghua Su |
IEEE Signal Process. Lett. | 6 |
| 2026 | Design of Multiple ISL-Aware Waveforms With SINR Guarantees for Integrated Sensing and Multi-User Communications
Chunxuan Shi, Yongzhe Li, Ran Tao 0003 |
IEEE Signal Process. Lett. | 3 |
| 2026 | Parametric Chunk Quantization Algorithm for Fast Passive Emitter Localization
Mingrui Wu, Hao Huan, Ran Tao 0003 |
IEEE Signal Process. Lett. | 3 |
| 2026 | Structure-Aware Dual Semantic Augmentation Alignment for Unsupervised Person Re-IdentificationabstractIn recent years, the inherent noise labels in unsupervised training can easily lead to structural distortion in the feature space and the mutual reinforcement of pseudo-label noise. Existing methods mostly focus on single-scale feature alignment, lacking collaborative perception and alignment of local topological structure and cross-camera distribution structure. To address this, we proposes a Structure-Aware Dual Semantic Augmentation Alignment network (SDAA). Firstly, we design an Adaptive Corrected Spatial Attention (ACSA) to achieve adaptive fusion of local detail structure and global semantic structure through multi-scale hybrid pooling and progressive spatial upsampling. Furthermore, we propose a Structural Semantic Decoupling module (SSD) to suppress microstructural noise by perceiving and reconstructing the local topological structure of the feature space. Finally, a Domain Semantic Alignment module (DSA) achieves explicit alignment of cross-camera distribution structure through camera-aware surrogate memory. Experiments on multiple public datasets, including Market-1501, MSMT17, and PersonX, demonstrate that our proposed method outperforms existing methods, validating the effectiveness and generality of the proposed module. Yunzuo Zhang, Weiqi Lian, Zhiwei Tu, Liye Xue, Ran Tao 0003 |
IEEE Signal Process. Lett. | 6 |
| 2025 | DFT-Spread-Based OTFS Waveform Design With Good Peak-to-Average Power Ratio for Joint Sensing and CommunicationsabstractWe study the problem of orthogonal time frequency space (OTFS) waveform design for joint sensing and communications. Our main objective is to achieve a low peak-to-average power ratio (PAPR) for the OTFS waveform with DFT spread to communication symbols, so that good parameter estimation and bit error rate performances in the presence of phase noise can be obtained for the sensing and communication sides, respectively. To this end, we first devise a symbol pattern scheme with pilot and reference components elaborated in the delay-Doppler domain for the OTFS, with the aim of improving phase noise and channel estimations. Based on this, we then perform power allocations on the designed pattern to further implement the reduction of PAPR for the OTFS waveform. The overall OTFS design is formulated as an optimization problem that incorporates the PAPR metric into the objective function for minimization. By replacing the infinity norm with an equivalent form, we convert the formulated optimization problem into a new tractable form, which allows us to apply majorization-minimization techniques for finding solutions. Our major contributions also lie in elaborating the majorant for the resulting problem and transforming into its dual problem. Simulation results verify the effectiveness of our design. Zhiying Chen, Yongzhe Li, Ran Tao 0003 |
ICASSP | 3 |
| 2025 | Low-Correlation OFDM Waveform Design With Optimally Coded Sub-Carriers for the Joint Sensing and CommunicationsabstractWe study the design of orthogonal frequency division multiplexing (OFDM) waveform for the joint sensing and communications (JSC), whose sub-carriers are to be optimally coded by a set of sequences. To obtain such waveform that simultaneously exploits frequency diversity and modulation flexibility for the JSC, we take the correlation level of waveform for sensing and the orthogonality of modulation sequences for communications into consideration. Specifically, we choose to minimize the integrated sidelobe level (ISL) of the OFDM waveform, and meanwhile, to ensure reasonable constraints enforced on sub-carriers with code modulations. In view of this, an ISL-minimization based design with respect to the modulation sequences is therefore formulated, which is generally non-convex. To tackle it, we introduce virtually auxiliary variables to help reformulate the original optimization problem, and then apply the framework of consensus alternating direction method of multipliers for finding solutions. Our major contributions also lie in elaborating an augmented Lagrangian for the newly obtained optimization problem via a first-order Taylor expansion on its objective function, based on which a closed-form solution is achieved for iterations. Simulation results verify the superiority of our proposed OFDM waveform design. Chunxuan Shi, Yongzhe Li, Ran Tao 0003 |
ICASSP | 3 |
| 2025 | Design of Multiple Binary Waveforms for the Joint MIMO Radar and CommunicationsabstractIn this paper, we focus on the multi-waveform design for the joint radar and communications, wherein the elements of waveforms are required to take binary values and to support both good integrated sidelobe level (ISL) and information embedding (IE) performances simultaneously. Since the binary waveform attribute enables limited degrees of freedom for the design, we partially choose to exploit both the phase and index modulations (PIM) of waveform elements to embed communication symbols in the fast-time domain, while we reserve a portion of them to improve the overall ISL of waveforms without PIM. To this end, we divide the fast-time waveforms to be designed into multiple blocks, each of which contains multiple uniform segments. We further develop a rule to instruct the elaboration of segments for IE via PIM within each block, and we leave the remaining segments among blocks unconstrained for the reduction on overall ISL of the binary waveforms. Based on the above, we formulate the design into a non-convex optimization problem that incorporates both the ISL minimization of waveforms and constraints on the fast-time IE. To tackle this problem, we reformulate it into a solvable integer optimization form, for which we employ the coordinate descent framework to find solutions. To make the work complete, we propose an associated method to determine the proper number of IE segments in the waveform design. Simulation results verify the effectiveness of our proposed design. Xiaohan Zhao, Yongzhe Li, Ran Tao 0003 |
ICASSP | 3 |
| 2025 | The OFDM Waveform Design With Optimally Coded Subcarriers for Joint Sensing and CommunicationsabstractWe study the design of the orthogonal frequency division multiplexing (OFDM) waveform that aims to simultaneously exploit frequency diversities and modulation flexibilities for joint sensing and communications (JSC). Such type of OFDM waveform adopts an optimal coding on subcarriers through a set of pre-designed sequences. To this end, we take into consideration the correlation level, peak-to-average-power ratio (PAPR), and orthogonality of modulation sequences of the OFDM waveform for JSC, based on which we devise two major corresponding designs. Strategically, we choose to minimize the integrated sidelobe level (ISL) of the OFDM waveform to obtain good correlation property in the first type of design, and meanwhile, to enforce reasonable constraints on OFDM subcarriers for achieving strict mutual orthogonality between sequences. For the second type of design, we further incorporate the PAPR reduction on the OFDM waveform into the development for easy practical implementations, wherein the PAPR is treated as either an objective for minimization or a constraint for bounding. Technically, we formulate both types of designs into different optimization problems. To tackle them, we introduce virtually auxiliary variables and reformulate the original problems into tractable forms, wherein the elaboration on constraints are also involved. Then, we apply the consensus alternating direction method of multipliers (ADMM) framework to find solutions, wherein proper augmented Lagrangians are elaborated to facilitate ADMM procedures. Our proposed methods enable closed-form solutions at iterations. For the development of corresponding algorithms, our major contributions also lie in converting the PAPR to a solvable form via an ℓp-norm based approximation and exploring fast implementations to conduct iterations. Simulation results show the superiority of our proposed designs in terms of different aspects. Chunxuan Shi, Yongzhe Li, Ran Tao 0003 |
IEEE Internet Things J. | 3 |
| 2025 | A Passive Synthetic Aperture Localization Method Based on Sparse Sampling ReconstructionabstractIn passive localization, the synthetic aperture positioning (SAP) method can achieve high precision and high-resolution positioning. However, existing research neglects the issue of target adaptability. For radar emitter targets, receivers can only periodically capture signals when the emitter’s beam scans toward the receiving antenna, resulting in spectral aliasing of the received signals. This leads to multiple false targets in localization images and reduced accuracy. This study employs the Fractional Fourier Transform (FrFT) integrated with compressed sensing for continuous signal reconstruction, aiming to eliminate spurious targets and enhance positioning accuracy. Initially, spectral aliasing is suppressed through FrFT, capitalizing on the approximately linear frequency-modulated (LFM) characteristics inherent in Doppler signals. Subsequently, continuous signal is reconstructed using compressed sensing with FrFT basis vectors forming the sensing matrix. Finally, the SAP method is implemented to achieve precise positioning. The effectiveness of the proposed method has been validated through simulations and Unmanned Aerial Vehicle (UAV) experiments, demonstrating that it significantly enhances the adaptability of SAP methods to radar emitter targets. Hao Huan, Ran Tao 0003, Yue Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | Polynomial Fitting Emitter Localization Method Based on Multisubaperture Phase StitchingabstractIn passive localization, the synthetic aperture positioning (SAP) method enables high-precision positioning under low signal-to-noise ratio (SNR) conditions. However, higher-order phase errors induced by platform self-localization errors degrade image focusing and reduce localization accuracy. In this paper, a polynomial fitting approach based on designing optimal pre-whitening filters using autoregressive models and employing iteratively reweighted least squares (IRLS) is applied to the unwrapped phase to eliminate higher-order error components. Additionally, a multiple sub-aperture phase stitching method is proposed to mitigate phase susceptibility to noise interference and error accumulation during phase unwrapping. The effectiveness of the proposed method is validated through both simulations and UAV experiments. Results demonstrate that meter-level localization accuracy can be achieved for the emitter target. Hao Huan, Ran Tao 0003, Yue Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | An interpretable convolutional neural network via generalized time-frequency scattering
Xiaoping Liu 0005, Gong Chen 0003, Jun Shi 0003, Ran Tao 0003 |
Signal Process. | 4 |
| 2025 | Robust Multitask Diffusion Bias Compensation M-Estimate Algorithms for Distributed Adaptive Learning With Noisy InputabstractThis letter studies the issue of robust multitask distributed estimation under the error-in-variable (EIV) model where input noise and output impulsive noise are considered. In such cases, existing distributed algorithms suffer from severe performance degradation. To tackle this problem, a robust multitask diffusion bias-compensated least mean M-estimate (R-MD-BCLMM) is proposed. We adopt a new real-time input noise variance estimation method which utilizes piecewise linearity of the modified Huber function to resist input noises. To further improve network information exchange capability and estimation performance, a robust spatial average combination based multitask adaptive clustering strategy is proposed. Finally, simulations demonstrate that the proposed R-MD-BCLMM algorithm outperforms some state-of-the-art distributed algorithms. Senran Peng, Lijuan Jia, Zi-Jiang Yang, Ran Tao 0003, Yue Wang 0001 |
IEEE Signal Process. Lett. | 4 |
| 2025 | Research on Distributed Synthetic Aperture Passive Positioning and Optimal Geometric Configuration
Junhua Yang, Hao Huan, Ran Tao 0003 |
IEEE Signal Process. Lett. | 4 |
| 2025 | Cross-Domain Few-Shot Learning Method Based on Fractional Domain Information for Hyperspectral Image Multi-Class Change DetectionabstractHyperspectral image multi-class change detection (HSI-MCD) based on deep learning (DL) rely significantly on the number of labeled data. Due to the high cost of manually labeling for hyperspectral images (HSIs), obtaining a large amount of labeled samples is difficult. Moreover, for multi-class change detection (MCD) tasks, there is the phenomenon of semantic cross-coupling of changes due to complex change scenarios. To solve the above problems, a cross-domain few-shot learning method based on fractional domain information for HSI-MCD (FrCFSL) is proposed. Firstly, a spectral-spatial-fractional information extraction module is proposed, which can extract spectral-spatial-fractional domain joint feature. Thus, the module can obtain more comprehensive and discriminative representations of land cover categories, alleviating the phenomenon of semantic cross-coupling between classes. Afterward, a cross-domain fewshot learning strategy is introduced, where it learns task-relevant category discrimination meta-knowledge from a pair of richly labeled very high-resolution optical images (VHRIs) dataset and transfers it to the bitemporal HSIs dataset. Thus, the model can achieve better MCD performance with a small number of labeled samples. Finally, to mitigate the domain distribution differences between VHRIs data and HSIs data, a topological structure alignment module is proposed to align the intrinsic topological relationships between land cover categories, thus narrowing the gap between the two domain distributions. Through experiments conducted on three HSI-MCD datasets and comparative analysis with six state-of-the-art methods, the validity and stability of the proposed method are indicated. Shou Feng, Jinghe Zhang, Yuanze Fan, Xinyao Liu, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Cross-Domain Hyperspectral Image Classification Based on Bi-Directional Domain AdaptationabstractUtilizing hyperspectral remote sensing technology enables the extraction of fine-grained land cover classes. Typically, satellite or airborne images used for training and testing are acquired from different regions or times, where the same class has significant spectral shifts in different scenes. In this paper, we propose a Bi-directional Domain Adaptation (BiDA) framework for cross-domain hyperspectral image (HSI) classification, which focuses on extracting both domain-invariant features and domain-specific information in the independent adaptive space, thereby enhancing the adaptability and separability to the target scene. In the proposed BiDA, a triple-branch transformer architecture (the source branch, target branch, and coupled branch) with semantic tokenizer is designed as the backbone. Specifically, the source branch and target branch independently learn the adaptive space of source and target domains, a Coupled Multi-head Cross-attention (CMCA) mechanism is developed in coupled branch for feature interaction and inter-domain correlation mining. Furthermore, a bi-directional distillation loss is designed to guide adaptive space learning using inter-domain correlation. Finally, we propose an Adaptive Reinforcement Strategy (ARS) to encourage the model to focus on specific generalized feature extraction within both source and target scenes in noise condition. Experimental results on cross-temporal/scene airborne and satellite datasets demonstrate that the proposed BiDA performs significantly better than some state-of-the-art domain adaptation approaches. In the cross-temporal tree species classification task, the proposed BiDA is more than 3%∼5% higher than the most advanced method. The codes will be available from the website: https://github.com/YuxiangZhang-BIT/IEEE TCSVT BiDA. Yuxiang Zhang 0005, Wei Li 0032, Wen Jia, Mengmeng Zhang 0005, Ran Tao 0003, Shunlin Liang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | A Prototype-Aware Learning and Dual-View Regularization Network for Weakly Supervised Change Detection in VHR Remote Sensing ImagesabstractChange detection (CD) is a critical task for monitoring the spatiotemporal evolution of the Earth’s surface. Recently, due to the advantages of reduced annotation cost and improved labeling efficiency, weakly supervised change detection (WSCD) has attracted increasing attention. However, existing WSCD methods encounter several critical challenges, including incomplete activation of class activation maps (CAMs), interference from noisy pseudo-labels during training, and instability in change recognition caused by illumination and environmental variations. To address these issues, we propose a prototype-aware learning and dual-view regularization network (PDRNet) for image-level WSCD. Specifically, to address the issue of incomplete activation caused by the tendency of CAM to focus excessively on locally discriminative regions, PDRNet devises a prototype-aware module (PAM), which captures stable category prototypes and refines CAM quality by reactivating hierarchical features. Furthermore, to mitigate the network’s sensitivity to noisy pseudo-labels, a dual-view regularization strategy (DRS) is designed to partition pseudo-labels into clean and noisy regions. Region-specific regularization is subsequently employed to improve the robustness of the model against noisy supervision. Finally, to enhance the capability of identifying changed regions, PDRNet constructs a wavelet-based change enhancement module (WCEM) to decompose bi-temporal features into multiple frequency bands. This facilitates the comprehensive utilization of low-frequency structural semantics and high-frequency texture details. Extensive experiments and analyses conducted on three publicly available CD datasets yield the superiority of PDRNet. Shou Feng, Chunhui Zhao 0003, Yingjie Tang, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Fractional Fourier-Enhanced Fusion Network Based on Pareto Optimization for Hyperspectral and LiDAR Data ClassificationabstractIn recent years, the utilization of hyperspectral image (HSI) and light detection and ranging (LiDAR) for collaborative classification has emerged as a significant research direction in earth observation tasks, with diverse joint classification algorithms showing promising performance using varying network architectures. However, these methodologies infrequently address the challenge of fusion arising from the substantially larger volume of HSI feature information compared to LiDAR features. Moreover, the effective learning of HSI and LiDAR features while mitigating modality conflicts remains an area that necessitates further investigation. As such, a Fractional Fourier Enhanced Fusion Network based on Pareto Optimization (FrFENet) is proposed for HSI and LiDAR Data classification. To address the disparity in information volume between modalities, a weighted fractional Fourier enhanced fusion module (WFrFEF) is introduced, which applies a weighted fractional Fourier transform to HSI features, enhancing their representations and facilitating balanced fusion with LiDAR features. Furthermore, a Pareto-based soft optimization strategy, HLPareto, is designed to balance learning rates across HSI and LiDAR features in a dual-branch network, effectively avoiding optimization conflicts. Additionally, a spatial-spectral integration module (SSIM) and an elevation information enhancement module (EIEM) are developed to improve feature extraction. The SSIM enables effective spatial-spectral fusion by facilitating token-level interactions, while the EIEM enhances elevation feature representation, preserving spatial geometric information in LiDAR data. Extensive experiments and comparative analyses conducted on three widely utilized HSI and LiDAR datasets have shown that the proposed FrFENet exhibits superior classification performance. Shou Feng, Hongtao Deng, Yabin Hu, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Transformer-Based Cross-Domain Few-Shot Learning for Hyperspectral Target DetectionabstractDeep learning-based methods have made significant progress in hyperspectral target detection (HTD). Unfortunately, limited target prior information and imbalance class resulting from the low occurrence probability of target leaves deep learning-based methods to confront bottlenecks. To ameliorate the abovementioned issues, a Transformer-based cross-domain few-shot learning (TCFSL) method is proposed for HTD. First, the TCFSL leverages cross-domain few-shot learning (FSL) to establish FSL tasks in both the source domain (SD) and the target domain (TD). This allows the TCFSL to learn transferable knowledge of the SD and distinguishable feature embedding model for the TD, to address the problems of target priori lacking and imbalance class. Second, feature-level and distribution-level domain adaptation (DA) is used to tackle the problem of domain shift in cross-domain FSL. The feature-level DA extracts intradomain information of the SD and TD to learn their common features to alleviate domain shift. The distribution-level DA based on cross-Transformer present interdomain distribution-level information aggregation and captures domain similarities of two data domains. By pursuing similarities between two data domains, the distribution-level DA block prompts specific FSL tasks in each domain, facilitating the target detection task. Finally, cross-domain FSL and DA blocks are trained in a unitary manner, which facilitates real-time information interaction and parameter adjustment between different blocks to achieve the optimal model. Experiments conducted on six HSI datasets indicate that the TCFSL outperforms 12 compared methods. Shou Feng, Fengchao Xiong, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Fractional-Domain Information-Enhanced Hyperspherical Prototype Learning Method for Hyperspectral Image Open-Set ClassificationabstractIn recent years, research in the field of hyperspectral image classification (HSIC) has increasingly focused on the open-set problem. Open-set classification demands not only accurately classifying the known categories but also identifying the unknown samples that are not labeled or included within the training data during testing stage. Existing open-set methods often suffer from misclassification between the known and unknown categories due to their inadequate utilization of metric space. Moreover, relying on a single threshold strategy performs poorly for identifying unknown categories in complex open environments. In this paper, a fractional domain information enhanced hyperspherical proto-type learning method (FrHSPL) is proposed for hyperspectral image open-set classification. FrHSPL develops a hyperspherical prototype learning (HSPL) strategy that ensures the features of known categories are uniformly distributed on the hypersphere. Therefore, HSPL can effectively enhance inter-class separability and optimize the exploitation of metric space. Subsequently, to enhance the discrimination capability of spectral features, a frequency-spatial-spectral information aggregation module is devised to deeply integrate fractional domain information with spatial and spectral information. Finally, an open-set recognition module is designed to identify unknown categories by using the prototypes of each known category along with the corresponding prototype radii. Extensive experiments on four common HSI datasets indicate that the proposed FrHSPL exhibits superior performance in comparison with both closed-set and open-set methods. Shou Feng, Cong'an Xu, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | A Nonlinear Weighted Graph Convolution Network Based on Manifold Geometric Regularization for Hyperspectral Image ClassificationabstractExtracting spatial-spectral joint features has become a critical approach for improving model classification performance in the field of hyperspectral image classification (HSIC). However, existing methods fail to fully exploit nonlinear spatial-spectral information. Unlike traditional convolutional neural networks (CNNs), graph convolutional neural networks (GCNs) can extract nonlinear spatial information. Nevertheless, both methods lack an accurate measurement of local neighborhood information, leading to blurred classification boundaries for ground objects. Additionally, the high-dimensional nature of hyperspectral data results in poor generalization and redundant information of trained models. To address these three issues, a nonlinear weighted graph convolution network based on manifold geometric regularization (MGR-NWGCN) method is devised for HSIC. Specifically, a nonlinear weighted graph convolution (NWGCN) module is designed, which utilizes a Graph-in-Graph structure based on cosine similarity-based normalized weighted graph convolution to extract nonlinear spatial-spectral information. Then, the manifold curvature regularization (C-MGR) module is implemented to improve the accuracy of similarity measurement and to enhance the generalization ability of the model, which constrains the model to form flatter feature manifold surfaces. Finally, the manifold intrinsic dimensionality regularization (ID-MGR) module is developed with the aim of eliminating redundant information, which embeds noise onto the surface of a low-dimensional manifold. The superior classification performance and robustness of the proposed MGR-NWGCN method are validated through extensive experiments on four datasets, with comparisons conducted against nine methods. Shou Feng, Cong'an Xu, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | DSNet: Dynamic Stitchable Neural Network for Hyperspectral Image ClassificationabstractHyperspectral image classification (HSIC) aims to identify land cover categories by leveraging the spectral and spatial information contained in hyperspectral images (HSI). Currently, many deep learning approaches utilize dual-branch networks to process spectral and spatial data separately, followed by the application of specialized modules to facilitate feature interaction or fusion. However, the design of these modules demands considerable time and effort from researchers and may not adequately capture the inherent relationships between independent spatial and spectral features in a dynamic manner. To address these issues, we propose the dynamic stitchable neural network (DSNet) for HSIC. While the DSNet maintains a dual-branch structure, it operates without traditional feature fusion or interaction. Instead, it employs a stitching network approach to integrate the two branches. Specifically, a spatial-spectral stitching module is presented to incorporates multiple stitching layers at various positions between the two network branches, creating new stitched networks that retain the strengths of both original networks. Additionally, a reinforcement learning-based strategy is designed for dynamically selecting stitching positions tailored to specific datasets, enabling the model to adaptively optimize the integration of spatial and spectral features. Recognizing the effectiveness of vision transformer (ViT) in learning spatial information and the capability of 1D convolutional neural network (1DCNN) in capturing spectral details, the DSNet directly stitches these two networks together. This fusion maximizes the utilization of both foundational networks, yielding a new hybrid network that delivers exceptional performance while also alleviating the burden on researchers to develop new architectures from scratch. Extensive experiments and analyses conducted on three public HSI datasets demonstrate the superiority of the proposed method, validating the effectiveness of our innovative modules. The codes of this work will be available from the website: https://github.com/ZZC/IEEE-TGRS-DSNet. Shou Feng, Zicheng Zhao, Bobo Xi, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Unsupervised Domain Adaptation With Hierarchical Masked Dual-Adversarial Network for End-to-End Classification of Multisource Remote Sensing DataabstractAlthough unsupervised domain adaptation (UDA) has been successfully applied for cross-scene classification of multisource remote sensing (MSRS) data, there are still some tough issues: 1) The vast majority of them are patch-based, requiring pixel by pixel processing at high complexity and ignoring the roles of unlabeled data between different domains. 2) Traditional masked autoencoder (MAE)-based methods lack effective multiscale analysis and require pre-training, ignoring the roles of low-level representations. As such, a hierarchical masked dual-adversarial DA network (HMDA-DANet) is proposed for cross-domain end-to-end classification of MSRS data. Firstly, a hierarchical asymmetric MAE (HAMAE) without pre-training is designed, containing a frequency dynamic large-scale convolutional (FDLConv) block to enhance important structural information in the frequency domain, and an intramodality enhancement and intermodality interaction (IAEIEI) block to embed some additional information beyond the domain distribution by expanding the cross-modal reconstruction space. Representative multimodal multiscale features can be extracted, while to some extent improving their generalization to the target domain. Then, a multimodal multiscale feature fusion (MMFF) block is built to model the spatial and scale dependencies for feature fusion and reduce the layer by layer transmission of redundancy or interference information. Finally, a dual-discriminator-based DA (DDA) block is designed for class-specific semantic feature and global structural alignments in both spatial and prediction spaces. It will enable HAMAE to model the cross-modal, cross-scale, and cross-domain associations, yielding more representative domain-invariant multimodal fusion features. Extensive experiments on five cross-domain MSRS datasets verify the superiority of the proposed HMDA-DANet over other state-of-the-art methods. Wen-Shuai Hu, Wei Li 0032, Heng-Chao Li 0001, Xudong Zhao 0003, Mengmeng Zhang 0005, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Collaborative Classification of Hyperspectral and LiDAR Date Based on Dynamic Multiple Fractional Fourier Domains FusionabstractCollaboratively utilizing the complementary information provided by hyperspectral imagery and light detection and ranging (LiDAR) data will extend the applications associated with land cover recognition and mapping. Existing joint classification algorithms mainly focus on learning complementary patterns in the pure spatial domain, while paying little attention to complementary cues in the spatial-frequency domain. The model’s expressive capability of these methods may be limited by an upper bound subject to the spatial domain. To fill this gap, a Dynamic Multiple Fractional Fourier Domains Fusion (DMFraF) is proposed for joint classification of hyperspectral and LiDAR data. Firstly, to comprehensively learn the complementary patterns between HSI and LiDAR data, we transform the features of two modalities into multiple fractional domains containing different spatial-frequency components for multimodal fusion. Secondly, to obtain the optimal representation from the multimodal features of multiple fractional domains, we propose a dynamic fusion scheme guided by the optimal transport (OT) technique, which can dynamically adjust the contributions from different fractional domains. Finally, to extract purer modality-specific features, we propose a channel aggregation Transformer encoder with central cross-attention (C2AT encoder), to aggregate channel-wise features of central pixels into the spatial branch and compress interference from noisy surroundings. Extensive experiments and analysis on three hyperspectral and LiDAR datasets suggest the superiority of the proposed method. Boao Qin, Shou Feng, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Language-Enhanced Dual-Level Contrastive Learning Network for Open-Set Hyperspectral Image ClassificationabstractIn recent years, language-supervised vision models have demonstrated impressive potential in learning open-world concepts. Some research has introduced this learning paradigm to the hyperspectral image (HSI) processing domain; however, there has been limited work integrating textual information into the hyperspectral open-set recognition task. To fill this gap, we leverage textual supervision information in open-set HSI classification (HSIC) and propose a language-enhanced dual-level contrastive learning network (LDCLNet). Specifically, we introduce a linguistic mode with prior knowledge as a supervised signal to enhance the metric distances between closed-set samples and provide supplementary semantic information for open-set samples. Second, a dual-level visual-language (V-L) contrastive learning (CL) approach, which can align visual and language embeddings separately at the instance level and manifold level, is proposed to establish a more accurate link between visual and language representations. Finally, a distance-refined open-set recognition method is proposed, which aims to effectively discover unknown class samples during testing by refining predictions of known and unknown classes. Extensive experiments and analysis on three public HSI datasets validate the effectiveness of LDCLNet. Boao Qin, Shou Feng, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | FCMMA: Fourier Conditional Mask-Based Mixed Attention Method for Hyperspectral Anomaly DetectionabstractIn recent years, reconstruction-based methods have achieved excellent detection results in the field of hyperspectral anomaly detection (HAD). These methods predominantly operate on two aspects regarding their working principles: 1) reconstructing background pixels and 2) suppressing anomalous pixels. However, most methods only tackle the HAD task from the spatial and spectral domains, making it challenging to effectively suppress anomalies. To eliminate these issues, this article proposes a Fourier conditional mask-based mixed attention (FCMMA) method. First, we propose the FCMMA method for HAD. FCMMA generates a conditional mask (CMASK) that suppresses anomalous high-frequency information and preserves background low-frequency information in the frequency domain, optimizing the anomaly detection process. In addition, to achieve fine-grained HAD, we propose the Fourier anomaly suppression filter (FASF). FASF uses Fourier techniques to manage background and anomalies, improving detection via precise frequency decoupling. Finally, a CMASK network is designed to effectively suppress anomalies. The CMASK network integrated the FASF module and the spatial-spectral multilayer perceptual (SSMLP) machine module together to enhance the transformation and representation capabilities of the generated masks, which can also help suppress anomalies. The results on five different datasets show that the proposed method is more effective and superior when compared to nine state-of-the-art methods. Shou Feng, Nan Su 0001, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | GeoFlowNet-SAR: Earthquake Displacement Estimation From Synthetic Aperture Radar ImagesabstractDisplacement estimation using remote sensing images is an effective approach for assessing surface displacement caused by natural disasters like earthquakes and landslides. By employing pixel correlation algorithms, high-precision displacement maps can be generated from images taken before and after surface movement. However, traditional methods often rely on spatial regularization or frequency masking to reduce high-frequency noise, which can smooth spatial details and result in biased displacement estimates, especially near sharp discontinuities typical of earthquake surface ruptures. Moreover, sub-pixel displacement estimation using Synthetic Aperture Radar (SAR) images remains a challenge compared to optical images, due to the strong impact of speckle noise. This paper presents GeoFlowNet-SAR, an innovative sub-pixel displacement estimation method leveraging SAR images. SAR offers advantages thanks to an all-weather observation and high penetration, making it suitable for conditions typically challenging for optical systems in the visible light spectrum. This study uses Sentinel-1 SAR Single Look Complex (SLC) images with dual-polarization (VV and VH modes) and Interferometric Wide (IW) swath mode to balance coverage and resolution. By training on simulated displacement datasets with realistic sharp discontinuities, GeoFlowNet-SAR directly predicts surface displacement fields, providing highly efficient, robust, and precise results, while overcoming some limitations of traditional methods. The effectiveness of the proposed methodological contribution is first quantitatively demonstrated using synthetic simulated earthquake datasets, including comparisons with state-of-the-art correlation methods. The method is further validated using two real remote sensing images from the 2019 Ridgecrest earthquake and from the 2023 Turkey-Syria earthquake. The observed results from these real datasets confirm the effectiveness of GeoFlowNet-SAR in practical applications. The codes are available at: https://gricad-gitlab.univ-grenoble-alpes.fr/giffards/geoflownet-sar. James Hollingsworth, Erwan Pathier, Tristan Montagnon, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003, Jocelyn Chanussot, Sophie Giffard-Roisin |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Tuple Perturbation-Based Contrastive Learning Framework for Multimodal Remote Sensing Image Semantic SegmentationabstractDeep learning models exhibit promising potential in multimodal remote sensing image semantic segmentation (MRSISS). However, the constrained access to labeled samples for training deep learning networks significantly influences the performance of these models. To address that, self-supervised learning (SSL) methods have garnered significant interest in the remote sensing community. Accordingly, this article proposes a novel multimodal contrastive learning framework based on tuple perturbation, which includes the pretraining and fine-tuning stages. First, a tuple perturbation-based multimodal contrastive learning network (TMCNet) is designed to better explore shared and different feature representations across modalities during the pretraining stage and the tuple perturbation module is introduced to improve the network’s ability to extract multimodal features by generating more complex negative samples. In the fine-tuning stage, we develop a simple and effective multimodal semantic segmentation network (MSSNet), which can reduce noise by using complementary information from various modalities to integrate multimodal features more effectively, resulting in better semantic segmentation performance. Extensive experiments have been carried out on two published multimodal image datasets including optical and synthetic aperture radar (SAR) pairs, and the results show that the proposed framework can obtain more superior performance of semantic segmentation than the current state-of-the-art methods in cases of limited labeled samples. The source code is available athttps://github.com/yeyuanxin110/TMCNet-MSSNet. Yuanxin Ye, Jinkun Dai, Keyi Duan, Ran Tao 0003, Wei Li 0032, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Chirplet Fourier Analysis Network for Cross-Scene Classification of Multisource Remote Sensing DataabstractThe joint application of multisource remote sensing (MSRS) data, such as hyperspectral image (HSI) and light detection and ranging (LiDAR), offers significant potential for accurate land cover classification. However, the existing applications often struggle with domain shifts across scenes caused by sensor, illumination, and phase variations. Focusing on this domain adaptation problem, a Chirplet Fourier analysis network (ChirpFAN) is proposed for cross-scene classification of MSRS data in this paper. Firstly, a fractional spatial-frequency-phase feature extraction module including the fractional Fourier transform and a learnable phase-aware weighting block is proposed to capture multi-domain features. Secondly, a Chirplet swin transformer (ChirpST) block integrates a Chirplet Fourier analysis (ChirpFA) layer within a Swin transformer is designed to analyze multi-scale textural and oscillatory patterns. Finally, a modality-shared network including ChirpST blocks is designed for inter-modal fusion and alignment. Extensive experiments demonstrate that the ChirpFAN framework achieves state-of-the-art performance with 3% average improvements on three challenging cross-scene MSRS datasets. Code will be released on GitHub. Xudong Zhao 0003, Qi Ming, Yixiao Yang, Wen-Shuai Hu, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | An Adaptive Weighted Metric Learning Network Based on Fractional Domain Decoupling for Hyperspectral Change DetectionabstractHyperspectral image change detection (HSI-CD) possesses strong capabilities in exploring subtle changes in land cover. Due to sensor noise and imaging conditions, different semantic land covers in the same spatial location may exhibit similar spectral characteristics, leading to pseudoinvariant phenomena (identification of changed areas as unchanged areas) and causing a higher rate of false negatives in the model. Existing methods primarily focus on obtaining auxiliary discriminative information from spatial correlations or temporal dependencies. However, the frequency domain, which possesses rich global gradient distribution information, is often overlooked. The fractional Fourier transform (FrFT) is an extension of the Fourier transform (FT), representing a temporal-frequency local transformation suitable for processing nonstationary signals. Furthermore, multiorder fractional Fourier domains provide more observable domains for change discrimination. In this work, the application of FrFT is extended to the field of HSI-CD, and an adaptive weighted metric learning network based on fractional domain decoupling (FrFTML) is proposed. Specifically, the fractional domain decoupling (FrDD) module transforms the original HSI into multiorder FrFT domains and extracts their rich spatial-frequency mixed information, effectively suppressing noise while enhancing the representation of subtle differences. In addition, an adaptive weighted metric learning (AWML) framework is designed to merge multiorder fractional Fourier domain information in an adaptively weighted fusion manner. It introduces deep metric learning to explore the distances between samples of different categories that have relatively high similarity, so as to guide the direction of adaptive weighted fusion. Finally, the differential mask attention (DMA) module is designed to explore global contextual differences between bitemporal HSIs, obtaining change features with well-represented differences. Some experiments conducted on three public datasets indicate that FrFTML outperforms other state-of-the-art methods. Furthermore, the proposed method exhibits superiority in dealing with land cover that may lead to pseudoinvariant phenomena (identification of changed areas as unchanged areas). Shou Feng, Tianyu Lan, Yuanze Fan, Mengmeng Zhang 0005, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Distribution-Independent Domain Generalization for Multisource Remote Sensing ClassificationabstractThe availability of multisource remote sensing data provides the possibility for comprehensive observation. Convolutional neural networks (CNNs) naturally integrate multisource feature extractors and classifiers into an end-to-end multilayer design. However, CNN assumes data are independent and identically distributed. In practice, it is not always possible to access the labels or even data of the testing scenes. Therefore, the CNN-based methods have exposed its limitation on generalization ability. To solve the issue, a feature-distribution-independent network (FDINet) is designed for multisource remote sensing cross-domain classification without feature alignment and decoupling operations. On one hand, an elegantly designed baseline is used for extracting multisource cross-domain features. The baseline extracts the common line and texture features through shallow weight-sharing networks. More importantly, the modality prediction probability is used to measure the similarity between the source domains and the target domains, thereby improving cross-domain collaboration capabilities. On the other hand, the sharpness-aware feature discriminating (SAFD) strategy is developed for model optimization. Specifically, the generalization ability is improved by minimizing the sharpness of local optima. To avoid the decrease in feature discrimination caused by the gradient conflict between sharpness and overall loss, the discrimination constraints are designed to balance feature discrimination and generalization ability. Comprehensive experiments are conducted on two datasets, which demonstrate that the proposed FDINet outperforms other competitors in terms of quantitative and qualitative analyses. Yunhao Gao, Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Global Clue-Guided Cross-Memory Quaternion Transformer Network for Multisource Remote Sensing Data ClassificationabstractMultisource remote sensing data classification is a challenging research topic, and how to address the inherent heterogeneity between multimodal data while exploring their complementarity is crucial. Existing deep learning models usually directly adopt feature-level fusion designs, most of which, however, fail to overcome the impact of heterogeneity, limiting their performance. As such, a multimodal joint classification framework, called global clue-guided cross-memory quaternion transformer network (GCCQTNet), is proposed for multisource data [i.e., hyperspectral image (HSI) and synthetic aperture radar (SAR)/light detection and ranging (LiDAR)] classification. First, a three-branch structure is built to extract the local and global features, where an independent squeeze-expansion-like fusion (ISEF) structure is designed to update the local and global representations by considering the global information as an agent, suppressing the negative impact of multimodal heterogeneity layer by layer. A cross-memory quaternion transformer (CMQT) structure is further constructed to model the complex inner relationships between the intramodality and intermodality features to capture more discriminative fusion features that fully characterize multimodal complementarity. Finally, a cross-modality comparative learning (CMCL) structure is developed to impose the consistency constraint on global information learning, which, in conjunction with a classification head, is used to guide the end-to-end training of GCCQTNet. Extensive experiments on three public multisource remote sensing datasets illustrate the superiority of our GCCQTNet with regards to other state-of-the-art methods. Wen-Shuai Hu, Wei Li 0032, Heng-Chao Li 0001, Ran Tao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Task-Wise Sampling Convolutions for Arbitrary-Oriented Object Detection in Aerial ImagesabstractArbitrary-oriented object detection (AOOD) has been widely applied to locate and classify objects with diverse orientations in remote sensing images. However, the inconsistent features for the localization and classification tasks in AOOD models may lead to ambiguity and low-quality object predictions, which constrains the detection performance. In this article, an AOOD method called task-wise sampling convolutions (TS-Conv) is proposed. TS-Conv adaptively samples task-wise features from respective sensitive regions and maps these features together in alignment to guide a dynamic label assignment for better predictions. Specifically, sampling positions of the localization convolution in TS-Conv are supervised by the oriented bounding box (OBB) prediction associated with spatial coordinates, while sampling positions and convolutional kernel of the classification convolution are designed to be adaptively adjusted according to different orientations for improving the orientation robustness of features. Furthermore, a dynamic task-consistent-aware label assignment (DTLA) strategy is developed to select optimal candidate positions and assign labels dynamically according to ranked task-aware scores obtained from TS-Conv. Extensive experiments on several public datasets covering multiple scenes, multimodal images, and multiple categories of objects demonstrate the effectiveness, scalability, and superior performance of the proposed TS-Conv. Zhanchao Huang, Wei Li 0032, Xiang-Gen Xia 0001, Hao Wang 0122, Ran Tao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | FDGNet: Frequency Disentanglement and Data Geometry for Domain Generalization in Cross-Scene Hyperspectral Image ClassificationabstractCross-scene hyperspectral image classification (HSIC) poses a significant challenge in recognizing hyperspectral images (HSIs) from different domains. The current mainstream approaches based on domain adaptation (DA) methods need to access target data when aligning distributions between domains, limiting the applicability of the model. In contrast, recent domain generalization (DG) methods aim to directly generalize to unseen domains, eliminating the requirements for target data during training. Nonetheless, most DG-based methods overly focus on randomizing sample styles, leading to semantically compromised samples. In addition, broadening the source distribution without ensuring reasonable support may result in undesired extended distributions. To address these issues, we propose a novel DG network with frequency disentanglement and data geometry (FDGNet) for cross-scene HSIC. Specifically, we first develop a spectral-spatial encoder based on frequency disentanglement (FDSS encoder), which facilitates synthesized domains to preserve their semantic consistency while simulating interdomain gaps with the source domain. Second, to avoid the generation of unrealistic samples, we incorporate data geometry into adversarial training. This helps diversify new domains while keeping the data geometry of extended domains in an explainable support. To improve the learning of domain-invariant representation, we propose an intermediate domain sampling strategy based on the class-wise perceptual manifold. This strategy synthesizes reliable intermediate domains by sampling from class-wise manifold flows estimated over the source and extended domains. Extensive experiments and analysis on three public HSI datasets yield the superiority of our proposed FDGNet. The codes will be available from the website: https://github.com/Qba-heu/FDGNet. Boao Qin, Shou Feng, Chunhui Zhao 0003, Bobo Xi, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | PRF-Net: A Progressive Remote Sensing Image Registration and Fusion NetworkabstractMost of the existing fusion algorithms are not robust to unregistered input images. Even after image registration, nonlinear nonregistration may persist in the local areas of the images, leading to poor quality in the fused image. So, as to tackle these challenges, a progressive remote sensing image registration and fusion network is proposed in this article, and named PRF-Net, which is particularly useful when two images are from different platforms. First, a registration network is designed to register the input image patches, which includes a global spatial transform network (GSTN) and a local spatial warp network (LSWN). The GSTN is primarily used for coarse registration, applying rigid transformation to globally align the input images. After coarse registration, the preliminarily registered moving image is input into the LSWN for local fine-tuning to maximize correlation between the input image patches. Subsequently, the fine registered images are degraded and input into the fusion network to generate the fused image. To maintain sufficient spectral and spatial information of the fused image, a multiscale feature extraction (MSFE) block with a highly interpretable spatial details attention (SDA) block is designed, which can enhance the ability of fusion network to extract and preserve spatial details and spectral information. Three groups of experiments conducted on four types of remote sensing images give evidence of that the proposed PRF-Net exhibits excellent performance in both reduced and full resolutions, showcasing its outstanding registration and fusion quality. Zhangxi Xiong, Wei Li 0032, Xiaobin Zhao, Baochang Zhang 0001, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Signal Reconstruction from Nonideal Samples in Fractional Fourier Transform DomainabstractIn this paper, we investigate a local average sampling and reconstruction problem using the fractional Fourier transform (FRFT). We present certain necessary and sufficient conditions under which there is an average sampling theorem for signals bandlimited in the FRFT domain. Xiaoping Liu 0005, Gong Chen 0003, Jun Shi 0003, Ran Tao 0003 |
ICASSP | 4 |
| 2024 | Fast Algorithm Design for the Constant-Envelope Precoding in Massive Mimo Communications with Interference ExploitationabstractWe study the problem of constant-envelope precoding in massive MIMO communications with interference exploitation, whose main challenge lies in the non-convexity introduced by its constant-modulus constraints. Different from conventional approaches that typically involve constant-modulus approximations, we devise a new method with direct phase manipulations for precoding. To this end, we first formulate the precoding design into an unconstrained phase optimization with a "min-max" scheme, based on which we then apply the majorization-minimization (MM) technique to transform it into a convex minimization problem. For the sake of efficient solutions to this convex optimization, we focus on dealing with its dual problem, whose objective is further transformed into a simple quadratic form by means of elaborating the surrogate for MM. Finally, we arrive to a fast gradient-based solution for iteration. Simulation results show the superiority of our proposed algorithm over the existing methods in terms of different aspects, especially in the context of large-scale optimization. Chunxuan Shi, Yongzhe Li, Ran Tao 0003 |
ICASSP | 3 |
| 2024 | OFDM Waveform Design with Good Correlation Level and Peak-to-Mean Envelope Power Ratio for the Joint MIMO Radar And CommunicationsabstractIn this paper, we focus on the orthogonal frequency division multiplexing (OFDM) waveform design for the joint multipleinput multiple-output radar and communications. An efficient method to simultaneously reduce the integrated sidelobe level (ISL) and peak-to-mean envelope power ratio (PMEPR) of OFDM waveforms is proposed from the radar side, which also guarantees high-quality information transmission for communications. Specifically, we exploit the spectral and phase randomness of waveforms to implement fast-time information embedding, based on which we formulate the design into a solvable nonconvex optimization problem. To solve it, we first rewrite its objective function into a quartic form by exploring the inherent algebraic structures and properties, which is then converted to a new quadratic form that is easy to be dealt with. Moreover, we obtain closed-form solutions at each iteration by means of a series of derivations involving majorization-minimization techniques. Simulation results verify the effectiveness of our method over existing works. Yongzhe Li, Ran Tao 0003, Tao Shan |
ICASSP | 3 |
| 2024 | Design of Spatial-Slow-Time Constant-Modulus Waveform Transmission and Receive Adaptive Filter for Dual-Function Radar Communications with Reconfigurable Intelligent SurfaceabstractWe study the problem of jointly designing spatial-slow-time unimodular waveforms and receive adaptive filter for dual-function radar communications with reconfigurable intelligent surface (RIS), which aims to mitigate interference for radar and meanwhile to transfer accurate symbols for communications. Hence, we maximize the signal-to-interference-plus-noise ratio at the radar receiver via devising a joint optimization of the waveform transmission, receive filter, and phase control of the RIS, wherein we enforce mean square error based constraints for communications. A non-convex optimization problem is therefore formulated. To solve it, we choose a cyclic manner that transforms the formulated design into sub optimization problems. In particular, we exploit majorization-minimization techniques followed by the consensus alternating direction method of multipliers for finding solutions to these sub problems. Therein lies a set of properly elaborated surrogate and Lagrangian functions, which finally lead to closed-form expressions for iterations. Simulation results verify the superiority of our proposed algorithm in terms of different aspects. Yuxuan Zhen, Chunxuan Shi, Yongzhe Li, Ran Tao 0003 |
ICASSP | 4 |
| 2024 | Motion Compensation for Synthetic Aperture Passive Localization Based on Weather Radar SignalsabstractIn emitter localization, the synthetic aperture positioning technique can achieve high-precision positioning even at a low signal-to-noise ratio (SNR). However, existing methods overlook the impact of receiver motion errors on the phase history of the received signal, leading to a reduction in localization accuracy. In this study, we propose a motion compensation (MoCo) technique for synthetic aperture passive localization using weather radar signals. Weather radar stations are chosen as reference stations due to their widespread coverage, high transmission power, and continuous signal transmission. The proposed method involves estimating and compensating for phase errors in the received signal, enabling phase coherency accumulation and achieving high-precision emitter localization. First, pulse compression is applied to the received weather radar signals to extract the phase information containing motion error details. Subsequently, we estimate motion errors using the extracted radar signal phase and apply phase compensation to the emitter target signal. Finally, the existing synthetic aperture positioning method is used to estimate the position of the target. Simulation results demonstrate that the method proposed in this paper offers superior accuracy in emitter localization compared to relying solely on real-time kinematic (RTK) for MoCo. The effectiveness of our proposed method is validated through actual unmanned aerial vehicle (UAV) experiments. Hao Huan, Ran Tao 0003, Yue Wang 0001, Xiaogang Tang |
WCNC | 3 |
| 2024 | Locality Robust Domain Adaptation for cross-scene hyperspectral image classification
Wei Li 0032, Yuxiang Zhang 0005, Ran Tao 0003 |
Expert Syst. Appl. | 5 |
| 2024 | Emitter Localization System Based on a Synthetic Aperture Map Drift Technique Aided by an InterferometerabstractIn emitter localization, the synthetic aperture positioning (SAP) technique can achieve high-precision positioning even at a low signal-to-noise ratio (SNR). However, existing methods require 2-D search implementation, which produces a huge amount of computation. In this study, the synthetic aperture map drift (MD) positioning method aided by an interferometer is proposed. This method reduces the amount of computation while achieving high-precision positioning. First, the exact azimuth position and approximate range position can be determined by a double-interference antenna via least-squares estimation. Doppler history is then used to correct range positioning error via an MD algorithm, which avoids the search of range and reduces the computational complexity. Simulation results show that the positioning accuracy of this method is close to the Cramer–Rao lower bound (CRLB). The effectiveness of the proposed method is verified through actual unmanned aerial vehicle (UAV) experiments. Hao Huan, Ran Tao 0003, Yue Wang 0001, Xiaogang Tang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Spectrum analysis for nonuniform sampling of bandlimited and multiband signals in the fractional Fourier domain
Yixiao Yang, Ran Tao 0003, Gang Li 0008, Chang Gao 0005 |
Signal Process. | 3 |
| 2024 | Communication-Efficient Secure Distributed Estimation With Noisy Measurement Against FDI AttackabstractWe study the problem of distributed estimation over wireless sensor networks (WSNs), where measurement noises and false data injection (FDI) attacks are considered. We propose the malicious node filtrating based secure diffusion bias compensated recursive least squares (MFS-dBCRLS) algorithm. To resist input noises, the BCRLS algorithm is introduced in local adaptation. To reduce the effect of network adversaries, a suitable time gating and a two-stage malicious node filtrating method are presented. Based on instantaneous state detection results in the first temporary filtrating stage, we permanently filtrate malicious neighbors in the second stage by proposing a new threshold test to decide final state. Besides, the selection range of the threshold is detailedly analyzed. Simulation results reveal communication-efficiency and robustness of our proposed method compared with some state-of-the-art algorithms under different FDI attacks. Zhanxi Zhang, Lijuan Jia, Senran Peng, Zi-Jiang Yang, Ran Tao 0003 |
IEEE Signal Process. Lett. | 5 |
| 2024 | VSS-Net: Visual Semantic Self-Mining Network for Video SummarizationabstractVideo summarization, with the target to detect valuable segments given untrimmed videos, is a meaningful yet understudied topic. Previous methods primarily consider inter-frame and inter-shot temporal dependencies, which might be insufficient to pinpoint important content due to limited valuable information that can be learned. To address this limitation, we elaborate on a Visual Semantic Self-mining Network (VSS-Net), a novel summarization framework motivated by the widespread success of cross-modality learning tasks. VSS-Net initially adopts a two-stream structure consisting of a Context Representation Graph (CRG) and a Video Semantics Encoder (VSE). They are jointly exploited to establish the groundwork for further boosting the capability of content awareness. Specifically, CRG is constructed using an edge-set strategy tailored to the hierarchical structure of videos, enriching visual features with local and non-local temporal cues from temporal order and visual relationship perspectives. Meanwhile, by learning visual similarity across features, VSE adaptively acquires an instructive video-level semantic representation of the input video from coarse to fine. Subsequently, the two streams converge in a Context-Semantics Interaction Layer (CSIL) to achieve sophisticated information exchange across frame-level temporal cues and video-level semantic representation, guaranteeing informative representations and boosting the sensitivity to important segments. Eventually, importance scores are predicted utilizing a prediction head, followed by key shot selection. We evaluate the proposed framework and demonstrate its effectiveness and superiority against state-of-the-art methods on the widely used benchmarks. Yunzuo Zhang, Yameng Liu, Weili Kang, Ran Tao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Fractional Fourier-Based Frequency-Spatial-Spectral Prototype Network for Agricultural Hyperspectral Image Open-Set ClassificationabstractAt present, hyperspectral image classification (HSIC) technology has been warmly concerned in all walks of life, especially in agriculture. However, existing classification methods operate under the closed-set assumption, which deviates from the real world with open properties. At the same time, there are more serious phenomena of different crops with similar spectrum and same crops with different spectrum in agricultural hyperspectral data, which is also a great challenge to existing methods. In this work, a fractional Fourier based frequency-spatial-spectral prototype network is proposed to address the challenges of open-set hyperspectral image classification in agricultural scenarios. Firstly, fractional Fourier transform is introduced into the network to combine the information in the frequency domain with the spatial-spectral information, so as to expand the difference between different classes on the premise of ensuring the similarity between classes. Then, the prototype learning strategy is introduced into the network to improve the feature recognition capability of the network through prototype loss. Finally, in order to break the stubbornly closed-set property of closed-set classification method, the open-set recognition module is proposed. The difference between the prototype vector and the feature vector is used to judge the unknown class. Experiments on three agricultural hyperspectral datasets show that this method can effectively identify unknown class without sacrificing the classification accuracy of closed-set, and has satisfactory classification performance. Maoyang Chen, Shou Feng, Chunhui Zhao 0003, Bo Qu, Nan Su 0001, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | High-Resolution Remote Sensing Image Change Detection Based on Fourier Feature Interaction and Multiscale PerceptionabstractAs a significant means of Earth observation, change detection in high-resolution remote sensing images has received extensive attention. Nevertheless, the variability in imaging conditions introduces style discrepancies and a range of pseudochange regions between bitemporal image pairs. Furthermore, changing objects possess diverse morphological representations, which makes accurately identifying change areas and delineating their boundaries within complex object distributions increasingly difficult. In response to the aforementioned challenges, we propose the Fourier feature interaction and multiscale perception (FIMP) model for effective change detection. To mitigate the impact of style discrepancies, FIMP employs the Fourier transform to adaptively filter bitemporal features in the frequency domain while mining the optimized bitemporal features relevant to the change detection task. To enhance the ability to recognize multiscale changing objects, FIMP aggregates and emphasizes the change areas with the introduced temporal change enhancement module (TCEM). By utilizing the U-fusion change perception module (UCPM) to perform multilevel bidirectional fusion of change features at different scales, FIMP can further enhance the ability to delineate complex semantic change boundaries. Experiments on three public datasets show that our approach outperforms seven state-of-the-art methods. Shou Feng, Chunhui Zhao 0003, Nan Su 0001, Wei Li 0032, Ran Tao 0003, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Cross-Domain Few-Shot Learning Based on Feature Disentanglement for Hyperspectral Image ClassificationabstractExisting hyperspectral cross-domain few-shot learning (FSL) methods focus mainly on elaborating on training strategies or domain alignment algorithms, while paying less attention to the biased meta-knowledge introduced by a large amount of source data and the implicit encouragement of learning target domain-specific attributes. In this paper, from the perspective of disentangled representation learning, a novel cross-domain FSL method based on feature disentanglement (FDFSL) is proposed for hyperspectral image classification (HSIC). Specifically, to suppress the representation biased towards the source data and enable the model to implicitly focus on the inherent knowledge of the target domain, an orthogonal low-rank feature disentanglement method is employed to acquire desired features of source and target pipelines. Furthermore, to preserve more shared and discriminative information from the heterogeneous data space (i.e., the spectral dimensions of the source and target scenes are typically different), a multi-order spectral interaction block based on central position encoding (MICD) is proposed to fully integrate the respective features into the spectral domain, which allows the model to emphasize informative spectral dimensions in a data-driven manner. Finally, to diversify the feature representation space while preventing the model overfitting domain alignment task, a self-distillation scheme is developed to facilitate the acquisition of task-relevant feature components. Extensive experiments and analysis on three public HSI datasets suggest the superiority of the proposed method. The code will be available on the website at https://github.com/Qba-heu/FDFSL. Boao Qin, Shou Feng, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003, Wei Xiang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Hyperspherical Structural-Aware Distillation Enhanced Spatial-Spectral Bidirectional Interaction Network for Hyperspectral Image ClassificationabstractThe existing methods for hyperspectral image classification (HSIC) mainly focus on the extraction of spectral and spatial features while paying less attention to the interaction of each other. Besides, most of them directly use a parameterized classifier as the final layer of the network. While this design is convenient for end-to-end optimization with the backbone, it overlooks the utilization of the metric space. In this article, a novel hyperspherical structural-aware distillation enhanced spatial–spectral bidirectional interaction network (HSDBIN) is proposed for HSIC. HSDBIN uses a dual-branch design combining the 1-D CNN and transformer to separately learn the detailed spectral correlations and global spatial relationships in parallel. Then, by interacting and aggregating the independent information between two parallel branches, a bidirectional interaction block across branches is designed to explore complementary clues between spectral and spatial pipelines. Finally, to enhance the utilization of metric space and keep compact intraclass relationship, we propose a hyperspherical structural-aware distillation (HSD) to transfer the geometric relationship of hyperspherical space into the metric space of output logits. Extensive experiments and analysis on three public HSI datasets suggest the superiority of the proposed method and verify the effectiveness of the proposed modules. Boao Qin, Shou Feng, Chunhui Zhao 0003, Bobo Xi, Wei Li 0032, Ran Tao 0003, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | An Object Fine-Grained Change Detection Method Based on Frequency Decoupling Interaction for High-Resolution Remote Sensing ImagesabstractChange detection is a prominent research direction in the field of remote sensing image processing. However, most current change detection methods focus solely on detecting changes without being able to differentiate the types of changes, such as “appear” or “disappear” of objects. Accurate detection of change types is of great significance in guiding decision-making processes. To address this issue, this article introduces the object fine-grained change detection (OFCD) task and proposes a method based on frequency decoupling interaction (FDINet). Specifically, in order to enhance the model’s ability to detect change types and improve its robustness to temporal information, a temporal exchange framework is designed. Additionally, to better capture spatial–temporal correlation in bi-temporal features, a wavelet interaction module (WIM) is proposed. This module utilizes wavelet transform for frequency decoupling, separating features into different components based on their frequency magnitudes. Then the module applies different interaction methods according to the characteristics of these frequency components. Finally, to aggregate complementary information from different-scale feature maps and enhance the representational capabilities of the extracted features, a feature aggregation and upsampling module (FAUM) is adopted. A series of experiments show the superiority of FDINet over most state-of-the-art methods, achieving good results on three different datasets. Yingjie Tang, Shou Feng, Chunhui Zhao 0003, Yuanze Fan, Qian Shi 0001, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Spatial-Temporal Weighted and Regularized Tensor Model for Infrared Dim and Small Target DetectionabstractDue to the confusion of target-like sparse structures and the interference of linear features in complex scenarios, many infrared small target detection methods struggle to effectively detect dim and small targets. In response to this challenge, we propose a new 3-D paradigm framework, which combines spatial-temporal weighting and regularization within a low-rank sparse tensor decomposition model. First, we design a novel spatial-temporal local prior structure tensor, named 3DST, which can significantly distinguish between targets and target-like sparse structures. Second, we introduce a three-directional log-based tensor nuclear norm (3DLogTNN) to provide a full characterization of the low-rankness of the background tensor. Third, we suggest a weighted three-directional total variation (3DTV) regularization to constrain smoothness features in background images. Finally, we develop an efficient alternating direction method of multipliers (ADMMs) to solve the proposed model. In particular, we devise a fast and accurate Sylvester tensor equation for accelerated subproblem solving. Extensive experimental results demonstrate that the proposed model has superior target detection and background suppression performance in complex scenarios compared with other detection methods. Jia-Jie Yin, Heng-Chao Li 0001, Yu-Bang Zheng, Gui Gao, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | SFSANet: Multiscale Object Detection in Remote Sensing Image Based on Semantic Fusion and Scale AdaptabilityabstractIn the field of computer vision, remote sensing image object detection plays an important role. Although the object detection algorithm has made significant progress, there are still problems in detecting objects with multi-scale in remote sensing image. Due to the insufficient utilization of object feature information, the detection accuracy of multi-scale objects is very low. To address the aforementioned issues, this paper proposes an effective object detection algorithm for remote sensing image based on semantic fusion and scale adaptability, known as SFSANet. Firstly, in view of the problem that the existing methods ignore the semantic differences between different scale feature maps, the semantic fusion (SF) module is proposed to enrich the semantic information and improve the ability to classify and locate objects. Next, to address the issue of the objects being easily interfered in complex background and the detection performance is poor, the spatial location attention (SLA) module is constructed to suppress background information and make key objects more prominent. Additionally, the scale adaptability module (SA) is designed to enrich the expression of feature information, realize the integration of global and local information, and ensure the integrity of image structure. Finally, we adopt the SIoU loss function as the localization loss to expedite model convergence. In order to verify the effectiveness of the proposed method, we conduct experiments on the mainstream datasets DIOR and NWPU VHR-10, which fully demonstrate the superiority of the proposed method. Yunzuo Zhang, Puze Yu, Shuangshuang Wang, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Relationship Learning From Multisource Images via Spatial-Spectral Perception NetworkabstractAdvances in multisource remote sensing have allowed for the development of more comprehensive observation. The adoption of deep convolutional neural networks (CNN) naturally includes spatial-spectral information, which has achieved promising performance in multisource data classification. However, challenges are still found with the extraction of spatial distribution and spectrum relationships, which eventually limit the classification performance. To solve the issue, a spatial-spectral perception network (S2PNet) is proposed to extract the advantages of different data sources and the cross information between data sources in a targeted manner. Specifically, the spatial perception network is developed to build the spatial distribution relationship from high-resolution images, while the spectral perception network extracts the spectrum relationship from spectral images. For perceiving cross information, a memory unit is utilized to store the features from different data sources in succession. In addition, the distance loss and reconstruction loss are introduced to keep the feature integrity, and the cross-entropy loss ensures that features can distinguish different classes. The comprehensive experiments are conducted on several datasets to validate the superiority of the proposed algorithm. The proposed S2PNet outperforms the considered classifiers with an average improvement of +0.77%, +5.62%, +1.58%, and +1.79% for overall accuracy values. Yunhao Gao, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003 |
IEEE Trans. Image Process. | 5 |
| 2024 | Multi-Scale Spatiotemporal Feature Fusion Network for Video Saliency PredictionabstractRecently, video saliency prediction has attracted increasing attention, yet the improvement of its accuracy is still subject to the insufficient use of multi-scale spatiotemporal features. To address this issue, we propose a 3D convolutional Multi-scale Spatiotemporal Feature Fusion Network (MSFFNet) to achieve the full utilization of spatiotemporal features. Specifically, we propose a Bi-directional Temporal-Spatial Feature Pyramid (BiTSFP), the first application of bi-directional fusion architectures in this field, which adds the flow of shallow location information on the basis of the previous flow of deep semantic information. Then, different from simple addition and concatenation, we design an Attention-Guided Fusion (AGF) mechanism that can adaptively learn the fusion weights of adjacent features to integrate them appropriately. Moreover, a Framewise Attention (FA) module is introduced to selectively emphasize the useful frames, augmenting the multi-scale temporal features to be fused. Our model is simple but effective, and it can run in real-time. Experimental results on the DHF1K, Hollywood-2, and UCF-sports datasets demonstrate that the proposed MSFF-Net outperforms existing state-of-the-art methods in accuracy. Yunzuo Zhang, Tian Zhang 0008, Cunyu Wu, Ran Tao 0003 |
IEEE Trans. Multim. | 4 |
| 2024 | FuBay: An Integrated Fusion Framework for Hyperspectral Super-Resolution Based on Bayesian Tensor RingabstractFusion with corresponding finer-resolution images has been a promising way to enhance hyperspectral images (HSIs) spatially. Recently, low-rank tensor-based methods have shown advantages compared with other kind of ones. However, these current methods either relent to blind manual selection of latent tensor rank, whereas the prior knowledge about tensor rank is surprisingly limited, or resort to regularization to make the role of low rankness without exploration on the underlying low-dimensional factors, both of which are leaving the computational burden of parameter tuning. To address that, a novel Bayesian sparse learning-based tensor ring (TR) fusion model is proposed, named as FuBay. Through specifying hierarchical sprasity-inducing prior distribution, the proposed method becomes the first fully Bayesian probabilistic tensor framework for hyperspectral fusion. With the relationship between component sparseness and the corresponding hyperprior parameter being well studied, a component pruning part is established to asymptotically approaching true latent rank. Furthermore, a variational inference (VI)-based algorithm is derived to learn the posterior of TR factors, circumventing nonconvex optimization that bothers the most tensor decomposition-based fusion methods. As a Bayesian learning methods, our model is characterized to be parameter tuning-free. Finally, extensive experiments demonstrate its superior performance when compared with state-of-the-art methods. Yinjian Wang, Wei Li 0032, Na Liu 0014, Yuanyuan Gui, Ran Tao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Representation-Enhanced Status Replay Network for Multisource Remote-Sensing Image ClassificationabstractDeep-learning-based methods are widely used in multisource remote-sensing image classification, and the improvement in their performance confirms the effectiveness of deep learning for classification tasks. However, the inherent underlying problems of deep-learning models still hinder the further improvement of classification accuracy. For example, after multiple rounds of optimization learning, representation bias and classifier bias are accumulated, which prevents the further optimization of network performance. In addition, the imbalance of fusion information among multisource images also leads to insufficient information interaction throughout the fusion process, thus making it difficult to fully utilize the complementary information of multisource data. To address these issues, a Representation-enhanced Status Replay Network (RSRNet) is proposed. First, a dual augmentation including modal augmentation and semantic augmentation is proposed to enhance the transferability and discreteness of feature representation, to reduce the impact of representation bias in the feature extractor. Then, to alleviate the classifier bias and maintain the stability of the decision boundary, a status replay strategy (SRS) is built to regulate the learning and optimization of the classifier. Finally, aiming to improve the interactivity of modal fusion, a novel cross-modal interactive fusion (CMIF) method is employed to jointly optimize the parameters of different branches by combining multisource information. Quantitative and qualitative results on three datasets demonstrate the superiority of RSRNet in multisource remote-sensing image classification, and its outperformance compared with other state-of-the-art methods. Wei Li 0032, Yinjian Wang, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | A Multistage Information Complementary Fusion Network Based on Flexible-Mixup for HSI-X Image ClassificationabstractMixup-based data augmentation has been proven to be beneficial to the regularization of models during training, especially in the remote-sensing field where the training data is scarce. However, in the process of data augmentation, the Mixup-based methods ignore the target proportion in different inputs and keep the linear insertion ratio consistent, which leads to the response of label space even if no effective objects are introduced in the mixed image due to the randomness of the augmentation process. Moreover, although some previous works have attempted to utilize different multimodal interaction strategies, they could not be well extended to various remote-sensing data combinations. To this end, a multistage information complementary fusion network based on flexible-mixup (Flex-MCFNet) is proposed for hyperspectral-X image classification. First, to bridge the gap between the mixed image and the label, a flexible-mixup (FlexMix) data augmentation strategy is designed, where the weight of the label increases with the ratio of the input image to prevent the negative impact on the label space because of the introduction of invalid information. More importantly, to summarize diverse remote-sensing data inputs including various modal supplements and uncertainties, a multistage information complementary fusion network (MCFNet) is developed. After extracting the features of hyperspectral and complementary modalities [X-modal, including multispectral, synthetic aperture radar (SAR), and light detection and ranging (LiDAR)] separately, the information between complementary modalities is fully interacted and enhanced through multiple stages of information complement and fusion, which is used for the final image classification. Extensive experimental results have demonstrated that Flex-MCFNet can not only effectively expand the training data, but also adequately regularize different data combinations to achieve state-of-the-art performance. Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Graph Information Aggregation Cross-Domain Few-Shot Learning for Hyperspectral Image ClassificationabstractMost domain adaptation (DA) methods in cross-scene hyperspectral image classification focus on cases where source data (SD) and target data (TD) with the same classes are obtained by the same sensor. However, the classification performance is significantly reduced when there are new classes in TD. In addition, domain alignment, as one of the main approaches in DA, is carried out based on local spatial information, rarely taking into account nonlocal spatial information (nonlocal relationships) with strong correspondence. A graph information aggregation cross-domain few-shot learning (Gia-CFSL) framework is proposed, intending to make up for the above-mentioned shortcomings by combining FSL with domain alignment based on graph information aggregation. SD with all label samples and TD with a few label samples are implemented for FSL episodic training. Meanwhile, intradomain distribution extraction block (IDE-block) and cross-domain similarity aware block (CSA-block) are designed. The IDE-block is used to characterize and aggregate the intradomain nonlocal relationships and the interdomain feature and distribution similarities are captured in the CSA-block. Furthermore, feature-level and distribution-level cross-domain graph alignments are used to mitigate the impact of domain shift on FSL. Experimental results on three public HSI datasets demonstrate the superiority of the proposed method. The codes will be available from the website: https://github.com/YuxiangZhang-BIT/IEEE_TNNLS_Gia-CFSL. Yuxiang Zhang 0005, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Cross-Scene Joint Classification of Multisource Data With Multilevel Domain Adaption NetworkabstractDomain adaption (DA) is a challenging task that integrates knowledge from source domain (SD) to perform data analysis for target domain. Most of the existing DA approaches only focus on single-source-single-target setting. In contrast, multisource (MS) data collaborative utilization has been extensively used in various applications, while how to integrate DA with MS collaboration still faces great challenges. In this article, we propose a multilevel DA network (MDA-NET) for promoting information collaboration and cross-scene (CS) classification based on hyperspectral image (HSI) and light detection and ranging (LiDAR) data. In this framework, modality-related adapters are built, and then a mutual-aid classifier is used to aggregate all the discriminative information captured from different modalities for boosting CS classification performance. Experimental results on two cross-domain datasets show that the proposed method consistently provides better performance than other state-of-the-art DA approaches. Mengmeng Zhang 0005, Xudong Zhao 0003, Wei Li 0032, Yuxiang Zhang 0005, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Fractional Fourier Image Transformer for Multimodal Remote Sensing Data ClassificationabstractWith the recent development of the joint classification of hyperspectral image (HSI) and light detection and ranging (LiDAR) data, deep learning methods have achieved promising performance owing to their locally sematic feature extracting ability. Nonetheless, the limited receptive field restricted the convolutional neural networks (CNNs) to represent global contextual and sequential attributes, while visual image transformers (VITs) lose local semantic information. Focusing on these issues, we propose a fractional Fourier image transformer (FrIT) as a backbone network to extract both global and local contexts effectively. In the proposed FrIT framework, HSI and LiDAR data are first fused at the pixel level, and both multisource feature and HSI feature extractors are utilized to capture local contexts. Then, a plug-and-play image transformer FrIT is explored for global contextual and sequential feature extraction. Unlike the attention-based representations in classic VIT, FrIT is capable of speeding up the transformer architectures massively and learning valuable contextual information effectively and efficiently. More significantly, to reduce redundancy and loss of information from shallow to deep layers, FrIT is devised to connect contextual features in multiple fractional domains. Five HSI and LiDAR scenes including one newly labeled benchmark are utilized for extensive experiments, showing improvement over both CNNs and VITs. Xudong Zhao 0003, Mengmeng Zhang 0005, Ran Tao 0003, Wei Li 0032, Wenzi Liao, Lianfang Tian, Wilfried Philips |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | New Perspectives on Observing Newton's RingsabstractNewton's rings experiment is a fundamental experiment. The rings counting method has been used to reveal the phenomenon of equal-thickness interference in physics for over 100 years. This paper proposes two perspectives on observing Newton's rings. From the perspective of signal processing, fractional Fourier transform are introduced into the Newton's rings experiment to reveal the mathematical nature of the fringe. From the perspective of data analysis, the deep neural network is trained so that it can intelligently analyze Newton's rings. High precision measurement of physical parameters such as the radius of curvature of a lens can be completed directly without counting rings. These two new perspectives provide good extensions to university physics experiments, which can be followed by additional theoretical and experimental sessions to further understand Newton's rings. The new session can be added to the current physics, signal processing and artificial intelligence courses. A brand-new course is designed to help students understand recent Newton's rings processing methods and know the pros and cons of them. which broaden their horizons and stimulate their creative thinking. Ming-Feng Lu, Jin-Min Wu, Wenming Yang, Feng Zhang 0011, Jihao Luo, Ran Tao 0003 |
FIE | 6 |
| 2023 | Multimodal Knowledge Distillation for Arbitrary-Oriented Object Detection in Aerial ImagesabstractRecently, many arbitrary-oriented object detection (AOOD) methods have been proposed and applied to remote sensing and other fields. For aerial platforms, lightweight structure and multimodal adaptations of convolutional neural network (CNN) models are urgently needed. Due to the limited model size, the performance of existing lightweight AOOD methods is low, especially in multimodal tasks. In this paper, a multimodal knowledge distillation (MKD) method is proposed for AOOD in aerial images. In MKD, a multimodal dynamic label assignment strategy is designed to select the optimal positive samples dynamically to adapt to different modalities and environments. Different multimodal localization and feature distillation modules are designed to make multimodal knowledge to be complementary and effectively learned by the lightweight model. Experiments on the public dataset demonstrated the effectiveness and advancement of MKD. Zhanchao Huang, Wei Li 0032, Ran Tao 0003 |
ICASSP | 3 |
| 2023 | Single-Shot Fractional Fourier Phase RetrievalabstractTraditional phase retrieval is generally concerned with re-covering a signal from its Fourier magnitude measurements whose inherent ambiguities make this problem especially difficult. In this work, we present an efficient phase retrieval technique from the single fractional Fourier transform (FrFT) magnitude measurement. Specifically, the FrFT measurement can be well-combined with signal priors via a generalized alternating projection framework, which can effectively alleviate the ambiguities of phase retrieval and the stagnation problem of numerical iterative processes. Through numerical simulations, we demonstrate that reconstructing an image from the single FrFT measurement leads to a significant performance improvement over that from the Fourier transform magnitude by using the proposed method. The source code is available at https://github.com/Yixiao-Yang/SFrFPR. Yixiao Yang, Ran Tao 0003 |
ICASSP | 2 |
| 2023 | Multi-Modal Domain Generalization for Cross-Scene Hyperspectral Image ClassificationabstractThe large-scale pre-training image-text foundation models have excelled in a number of downstream applications. The majority of domain generalization techniques, however, have never focused on mining linguistic modal knowledge to enhance model generalization performance. Additionally, text information has been ignored in hyperspectral image classification (HSI) tasks. To address the aforementioned shortcomings, a Multi-modal Domain Generalization Network (MDG) is proposed to learn cross-domain invariant representation from cross-domain shared semantic space. Only the source domain (SD) is used for training in the proposed method, after which the model is directly transferred to the target domain (TD). Visual and linguistic features are extracted using the dual-stream architecture, which consists of an image encoder and a text encoder. A generator is designed to obtain extended domain (ED) samples that are different from SD. Furthermore, linguistic features are used to construct a cross-domain shared semantic space, where visual-linguistic alignment is accomplished by supervised contrastive learning. Extensive experiments on two datasets show that the proposed method outperforms state-of-the-art approaches. Yuxiang Zhang 0005, Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003 |
ICASSP | 4 |
| 2023 | Efficent Large-Scale Multi-Unimodular Waveform Design with Good Correlation Properties via Direct Phase OptimizationsabstractIn this paper, we propose an efficient algorithm for designing large-scale multi-unimodular waveforms with low correlations. Different from existing approaches that commonly involve repetitive projections of complex values into their constant-modulus approximations, we conduct optimizations directly on the phase values of waveform elements. Specifically, we optimize the weighted integrated sidelobe level of waveforms, and formulate such design into an unconstrained optimization problem with respect to phase values of waveform elements. Then, we derive the gradient of the newly formulated objective function, through which we subsequently elaborate its majorant with the support of a properly designed Lipschitz-constant related quantity. Our major contributions also lie in obtaining a closed-form update of phase values that boils down to a gradient-descent regime, and calculating the update with fast implementations. Simulation results verify the superiority of our algorithm over existing state-of-the-art methods. Xiaohan Zhao, Yongzhe Li, Ran Tao 0003 |
ICASSP | 3 |
| 2023 | A survey on hyperspectral image restoration: from the view of low-rank tensor approximation
Na Liu 0014, Wei Li 0032, Yinjian Wang, Ran Tao 0003, Qian Du 0001, Jocelyn Chanussot |
Sci. China Inf. Sci. | 4 |
| 2023 | Remote Sensing Image Fusion With Task-Inspired Multiscale Nonlocal-Attention NetworkabstractRecently, convolutional neural networks (CNNs) have been developed for remote sensing image fusion (RSIF). To obtain competitive fusion performance, network design becomes more complicated by stacking convolutional layers deeper and wider. However, problems still remain when applying existing networks in practical applications. On the one hand, researchers focus on improving spatial resolution but ignore that the fused images will be used in subsequent interpretation applications, e.g., objection detection. On the other hand, RSIF involves different tasks with different image sources e.g., pansharpening of the panchromatic and multispectral image, hypersharpening of the panchromatic and hyperspectral image, etc. However, existing networks only solve one of them, failing to be compatible with other tasks. To address the above problems, a convenient task-inspired multiscale nonlocal-attention network (MNAN) is proposed for RSIF. The proposed MNAN focuses more on enhancing the multi-scale targets in the scene when improving the resolution of the fused image. In addition, the proposed network can be applied to both pansharpening and hypersharpening tasks without any modification. Na Liu 0014, Wei Li 0032, Xian Sun 0001, Ran Tao 0003, Jocelyn Chanussot |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | When Ramanujan sums meet affine Fourier transform
Hongxia Miao, Feng Zhang 0011, Ran Tao 0003, Mugen Peng |
Signal Process. | 3 |
| 2023 | Nonuniform MIMO Sampling and Reconstruction of Multiband Signals in the Fractional Fourier DomainabstractThis paper explores nonuniform multiple-input multiple-output (MIMO) sampling and reconstruction of signals with multiple bands in the fractional Fourier domain. We investigate discrete-time fractional Fourier transforms of nonuniformly sampled output signals of MIMO channel and study the resulting fractional spectral aliasing. In order to tackle the problem that the spectral aliasing differs with respect to multiple-output signals, we define combined aliasing boundaries and perform spectrum analysis within the fractional frequency sub-intervals separated by these elaborated boundaries. Moreover, we derive the conditions for combined reconstructing the fractional spectra of the input/output signals of MIMO channel and devise relevant reconstruction methods. Simulation results verify the effectiveness of our proposed methods. Gang Li 0008, Ran Tao 0003, Yongzhe Li |
IEEE Signal Process. Lett. | 3 |
| 2023 | Hyperspectral and LiDAR Data Classification Based on Structural Optimization TransmissionabstractWith the development of the sensor technology, complementary data of different sources can be easily obtained for various applications. Despite the availability of adequate multisource observation data, for example, hyperspectral image (HSI) and light detection and ranging (LiDAR) data, existing methods may lack effective processing on structural information transmission and physical properties alignment, weakening the complementary ability of multiple sources in the collaborative classification task. The complementary information collaboration manner and the redundancy exclusion operator need to be redesigned for strengthening the semantic relatedness of multisources. As a remedy, we propose a structural optimization transmission framework, namely, structural optimization transmission network (SOT-Net), for collaborative land-cover classification of HSI and LiDAR data. Specifically, the SOT-Net is developed with three key modules: 1) cross-attention module; 2) dual-modes propagation module; and 3) dynamic structure optimization module. Based on above designs, SOT-Net can take full advantage of the reflectance-specific information of HSI and the detailed edge (structure) representations of multisource data. The inferred transmission plan, which integrates a self-alignment regularizer into the classification task, enhances the robustness of the feature extraction and classification process. Experiments show consistent outperformance of SOT-Net over baselines across three benchmark remote sensing datasets, and the results also demonstrate that the proposed framework can yield satisfying classification result even with small-size training samples. Mengmeng Zhang 0005, Wei Li 0032, Yuxiang Zhang 0005, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Cybern. | 4 |
| 2023 | SiamBAG: Band Attention Grouping-Based Siamese Object Tracking Network for Hyperspectral VideosabstractA hyperspectral video contains frames with numerous spectral bands, providing fine reflectance information for object identification and tracking. Enriched features can be learned from spectral-spatial data using deep learning models. However, due to the difficulty in hyperspectral video collection, deep model training is often insufficient, causing reduced performance during the testing stage. To address this issue, we present a novel Band Attention Grouping-based Siamese framework (SiamBAG) for hyperspectral object tracking. SiamBAG employs massive color object tracking data to train a deep neural network. Band weights obtained by band attention module are used to group a hyperspectral image into multiple three-channel false-color images with approximate total group weights. Then multiple enhanced images obtained by histogram equalization are fed to the proposed SiamBAG network to generate a classification branch, a regression branch and a scale tuning branch. In the classification branch, the response maps of multiple groups are fused by regularized group weights to estimate the position of objects. Then the regression branch is used to obtain the initial object position of objects. The position offsets are fed back to the scale tune branch to relocate and fine-tune the object position by exploiting the similarity between template features and detection features. Experimental results demonstrate that the proposed tracker achieves superior tracking performance than other methods. The source codes of this paper will be released at https://github.com/zephyrhours/Hyperspectral-Object-Tracking-SiamBAG. Wei Li 0032, Zengfu Hou, Jun Zhou 0001, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | A Coarse-to-Fine Hyperspectral Target Detection Method Based on Low-Rank Tensor DecompositionabstractTo solve the problem of low target detection accuracy caused by the related quantities such as background, target and noise contained in hyperspectral images (HSIs), considering the use of the spatial spectrum and spectral characteristics while increasing the degree of discrimination between target and background, a coarse-to-fine hyperspectral image target detection algorithm based on low-rank tensor decomposition (HTDLTD) is proposed. The HTD based on low rank sparse decomposition mainly decomposes hyperspectral images in spectral dimension, which does not make full use of the spatial information of HSIs, resulting in low detection accuracy. In order to solve this problem, in view of the fact that the hyperspectral third-order tensor can describe the spatial information and spectral information of HSIs equally, the HTD method based on low-rank tensor decomposition (LRTD) is proposed to extract pure background information. Then, in order to solve the problem of low detection accuracy in the case of low target and background discrimination, the rough target detection method based on max over (SMF-MAX) target detection method is proposed to perform rough detection on the original HSI to obtain rough detection results. Finally, in order to further improve the performance of target detection, the fine target detection method based on spectral distance is proposed. By calculating the spectral distance between the original HSI and the synthesized HSI, the final reconstructed target detection result is obtained. Experimental results on three data sets show that the proposed HTDLTD exceeds eight state-of-the-art target detection methods used for comparison. Shou Feng, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | A Joint Optimization Based Pansharpening via Subpixel-Shift DecompositionabstractPatch-based spatial dictionary has been widely used to fuse a high-resolution panchromatic (PAN) image with a low-resolution multispectral (LMS) image under the framework of sparse representation. However, patch-based dictionary in the spatial domain is not sufficient to preserve spectral information, which may lead to large spectral distortion. To solve this problem, a new spectral dictionary based pansharpening method using subpixel-shift decomposition and joint optimization (termed as PANDA) is proposed. In this method, the model of pansharpening is formulated in a decomposed spectral domain under the sparse and low-rank constraint, as a joint optimization procedure of spectral dictionary and its coefficients. Specifically, a subpixel-shift decomposition is firstly constructed, to decompose the PAN image into a series of subimages with the same spatial resolution of the LMS image. Then, a new imaging model for the pansharpening problem of the LMS image and the decomposed PAN subimages is formulated, with sparse and low-rank constraints. And finally, a joint optimization procedure for the spectral dictionary and its coefficients are theoretically derived, using the spectral information provided by the LMS image and the spatial information provided by the entire PAN subimages, respectively. Experimental results on different datasets show that, the pansharpening performance of the proposed PANDA method outperforms the state-of-the-art methods in both spatial and spectral domains. Xiaolin Han 0001, Wei Leng, Qizhi Xu, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Multiarea Target Attention for Hyperspectral Image ClassificationabstractIn hyperspectral image (HSI) classification, objects corresponding to pixels of different classes exhibit varying size characteristics, which causes a challenge for effective pixelwise feature extraction and classification. In this article, we propose a novel multiscale model, called multiarea target attention (MATA). The proposed MATA uses an architecture that includes a shared feature extractor (FE) and classifier to capture multiscale spectral–spatial information effectively and efficiently. The FE uses a multiscale target attention module (MSTAM) to extract spectral–spatial information from target pixels and their similar pixels across multiscale areas, while$L_{2}$-normalization is used to address discrepancies between features of different scales. The classifier adopts a classwise decision weighting strategy to account for the varying sizes of different classes and the different contributions of semantic features at each scale to each class. Experimental results on five public HSI datasets demonstrate that the proposed MATA outperforms existing state-of-the-art single- and multiscale models, confirming its effectiveness and efficiency in HSI classification. Code is available athttps://github.com/huanliu233/MATA. Huan Liu 0015, Wei Li 0032, Xiang-Gen Xia 0001, Mengmeng Zhang 0005, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Tensor Spectral k-Support Norm Minimization for Detecting Infrared Dim and Small Target Against Urban BackgroundsabstractIn the low-altitude urban background with heavy interference, especially in the face of corner interference with higher intensity than the target, infrared (IR) dim and small target is extremely lack of prior information (i.e., size, shape and contrast information). In such case, the existing detection methods usually suffer from high false alarm or even failure. To deal with this situation, we develop a novel spatial-temporal tensor model with tensor spectralk-support norm minimization (STTM-TSNM) for detecting IR dim and small target. Firstly, the spatial-temporal information of the original image sequence can be preserved completely by constructing the holistic STTM. Then, according to the spatial-temporal related prior knowledge of the target and background, the target detection task is customized as an optimization problem of low rank and sparse tensor recovery. To better preserve the internal structure and capture more global information, the tensor spectralk-support norm minimization is introduced as the regularization term of the constraint background. Finally, draw support from the framework of alternating direction method of multipliers (ADMM) algorithm, the precise separation of target and background is achieved. In addition, to promote the prosperity of sequential detection methods, we released to the scientific community a small IR target dataset containing six image sequences with urban background. The experimental results on six real IR sequences demonstrate that our method outputs the most outstanding detection performance compared with the latest sequential detection methods. Dongdong Pang, Pengge Ma, Tao Shan, Ran Tao 0003, Qiuchun Jin |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | A Pansharpening Method Based on Hybrid-Scale Estimation of Injection GainsabstractThe injection scheme provides an efficient way for CS- and MRA-based pansharpening approaches. Within this paradigm, the estimation of injection gains is one of the keys to pansharpening outcomes, which has attracted much attention in the community. Most of the existing models are derived from the regression methodology. Hence, the reference is indispensable for the estimation. However, the reference is unavailable in practice, and therefore, the estimation is usually performed at a degraded scale. This article is devoted to the estimation of injection gains without reference. A hybrid-scale (HS) estimation, which involves both the high-resolution and low-resolution data, is proposed, along with three HS models. The proposed method features a context-based and fast implementation with fewer tunable parameters. Experimental results show that the HS models yield more accurate and robust results compared with the typical regression-based models, and they are also competitive with the state-of-the-art approaches. Yan Shi 0012, Aiyong Tan, Na Liu 0014, Wei Li 0032, Ran Tao 0003, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Remote-Sensing Scene Classification via Multistage Self-Guided Separation NetworkabstractIn recent years, remote sensing scene classification is one of research hotspots and has played an important role in the field of intelligent interpretation of remote sensing data. However, various complex objects and backgrounds form a variety of remote sensing scenes through spatial combination and correlation, which brings great challenges to accurately classify different scenes. Among them, the insufficient feature difference brought about the unbalanced change of background and target between inter-class sample and the feature representation inconsistency caused by the difference of representation among the intra-class samples have become obstacles to effectively distinguish different scene images. To address these issues, a Multi-stage Self-Guided Separation Network (MGSNet) is proposed for remote sensing scene classification. First of all, different from the previous work, it attempts to utilize the background information outside the effective target in the image as a decision aid through a target-background separation strategy to improve the distinguish ability between target similarity-background difference samples. In addition, the diversity of feature concerns among different network branches is expanded through contrastive regularization to improve the separation of target-background information. Additionally, a self-guided network is proposed to find common features between intra-class samples and improve the consistency of feature representation. It combines the texture and morphological features of images to guide feature learning, effectively reducing the impact of intra-class differences. Extensive experimental results on three benchmark demonstrate that MGSNet can achieve better classification performance compared to the state-of-the-art approaches. Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | CSTSUNet: A Cross Swin Transformer-Based Siamese U-Shape Network for Change Detection in Remote Sensing ImagesabstractChange detection (CD) in remote sensing images is a critical task that has achieved significant success by deep learning. Current networks often employ pixel-based differencing, proportion, classification-based, or feature concatenation methods to represent changes of interest. However, these methods fail to effectively detect the desired changes, as they are highly sensitive to factors such as atmospheric conditions, lighting variations, and phenological variations, resulting in detection errors. Inspired by the Transformer structure, we adopt a cross-attention mechanism to more robustly extract feature differences between bitemporal images. The motivation of the method is based on the assumption that if there is no change between image pairs, the semantic features from one temporal image can well be represented by the semantic features from another temporal image. Conversely if there is a change, there are significant reconstruction errors. Therefore, a Cross Swin Transformer based Siamese U-shaped network namely CSTSUNet is proposed for remote sensing change detection. CSTSUnet consists of encoder, difference feature extraction, and decoder. The encoder is based on a hierarchical Resnet with the Siamese U-net structure, allowing parallel processing of bitemporal images and extraction of multi-scale features. The difference feature extraction consists of four difference feature extraction modules that compute difference feature at multiple scales. In this module, Cross Swin Transformer is employed in each difference feature extraction module to communicate the information of bitemporal images. The decoder takes in the multi-scale difference features as input, injects details and boundaries iteratively level by level, and makes the change map more and more accurate. We conduct experiments on three public datasets, and the experimental results demonstrate that the proposed CSTSUNet outperforms other state-of-the-art methods in terms of both qualitative and quantitative analyses. Our code is available at https://github.com/l7170/CSTSUNet.git. Yaping Wu, Lu Li 0005, Nan Wang 0038, Wei Li 0032, Junfang Fan, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | FANet: An Arbitrary Direction Remote Sensing Object Detection Network Based on Feature Fusion and Angle ClassificationabstractHigh-precision remote sensing image object detection has broad application prospects in military defense, disaster emergency, urban planning, and other fields. However, the arbitrary orientation, dense arrangement, and small size of objects in remote sensing images lead to poor detection accuracy of existing methods. To achieve accurate detection, this paper proposes an arbitrary directional remote sensing object detection method, called FANet, based on feature fusion and angle classification. Initially, the angle prediction branch is introduced, and the circular smooth label method is used to transform the angle regression problem into a classification problem, which solves the difficult problem of abrupt changes in the boundaries of the rotating frame while realizing the object frame rotation. Subsequently, to extract robust remote sensing objects, innovative introduce pure convolutional model as a backbone network, while Conv is replaced by GSConv to reduce the number of parameters in the model along with ensuring detection accuracy. Finally, the strengthen connection feature pyramid network (SC-FPN) is proposed to redesign the lateral connection part for deep and shallow layer feature fusion, and add jump connections between the input and output of the same level feature map to enrich the feature semantic information. In addition, add a variable parameter to the original localization loss function to satisfy the bounding box regression accuracy under different IoU thresholds, and thus obtain more accurate object detection. The comprehensive experimental results on two public datasets for rotated object detection DOTA and HRSC2016 demonstrate the effectiveness of our method. Yunzuo Zhang, Cunyu Wu, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Morphological Transformation and Spatial-Logical Aggregation for Tree Species Classification Using Hyperspectral ImageryabstractHyperspectral image (HSI) consists of abundant spectral and spatial characteristics, which contribute to a more accurate identification of materials and land covers. However, most existing methods of hyperspectral image analysis primarily focus on spectral knowledge or coarse-grained spatial information while neglecting the fine-grained morphological structures. In the classification task of complex objects, spatial morphological differences can help to search for the boundary of fine-grained classes, e.g., forestry tree species. Focusing on subtle traits extraction, a spatial-logical aggregation network (SLA-NET) is proposed with morphological transformation for tree species classification. The morphological operators are effectively embedded with the trainable structuring elements, which contributes to distinctive morphological representations. We evaluate the classification performance of the proposed method on two tree species datasets, and the results demonstrate that the proposed SLA-NET significantly outperforms the other state-of-the-art classifiers. Mengmeng Zhang 0005, Wei Li 0032, Xudong Zhao 0003, Huan Liu 0015, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Language-Aware Domain Generalization Network for Cross-Scene Hyperspectral Image ClassificationabstractText information including extensive prior knowledge about land cover classes has been ignored in hyperspectral image (HSI) classification tasks. It is necessary to explore the effectiveness of linguistic mode in assisting HSI classification. In addition, the large-scale pretraining image–text foundation models have demonstrated great performance in a variety of downstream applications, including zero-shot transfer. However, most domain generalization methods have never addressed mining linguistic modal knowledge to improve the generalization performance of model. To compensate for the inadequacies listed above, a language-aware domain generalization network (LDGnet) is proposed to learn cross-domain-invariant representation from cross-domain shared prior knowledge. The proposed method only trains on the source domain (SD) and then transfers the model to the target domain (TD). The dual-stream architecture including the image encoder and text encoder is used to extract visual and linguistic features, in which coarse-grained and fine-grained text representations are designed to extract two levels of linguistic features. Furthermore, linguistic features are used as cross-domain shared semantic space, and visual–linguistic alignment is completed by supervised contrastive learning in semantic space. Extensive experiments on three datasets demonstrate the superiority of the proposed method when compared with the state-of-the-art techniques. The codes will be available from the website:https://github.com/YuxiangZhang-BIT/IEEE_TGRS_LDGnet. Yuxiang Zhang 0005, Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Infrared Small UAV Target Detection via Isolation ForestabstractThe illegal misuse of non-cooperative UAVs poses huge threats to society and life safety. Infrared imaging is reliable to monitor unmanned aerial vehicles (UAVs) and the anti-UAVs technology via infrared images has attracted more and more attention. In order to provide sufficient time for follow-up, UAVs are acquired at long distances, usually exhibiting the features of weak and small. Furthermore, infrared images are usually with low signal-to-clutter ratio (SCR). These factors make the correct detection of UAVs a challenge. Existing methods do not fully exploit the phenomenon that the UAVs are easily isolated, resulting in unsatisfactory detection results. For alleviating the issue, a novel detection method via isolation Forest (iForest) is proposed. In the proposed method, the multi-direction couple-order derivative properties are firstly analyzed, which enlarges the feature difference between UAVs and background. Then, a global iForest is constructed, which takes full advantage of the phenomenon that UAVs are susceptible to being isolated. As far as we know, this is the first time that iForest is constructed in infrared small targets detection field. Furthermore, a local iForest is created, which further eliminates the residual false alarms of the result of global iForest. Experiments on nine sequences demonstrate the performance of the proposed method, which is capable of detecting various UAVs under diverse background. Mingjing Zhao, Wei Li 0032, Lu Li 0005, Jin Hu 0004, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | High-Resolution Remote Sensing Bitemporal Image Change Detection Based on Feature Interaction and Multitask LearningabstractWith the development of remote sensing technology, high-resolution (HR) remote sensing optical images have gradually become the main source of change detection data. Albeit, the change detection for HR remote sensing images still faces challenges: 1) in complex scenes, a region contains a large amount of semantic information, which makes it difficult to accurately locate the boundaries between different semantics in the feature maps and 2) due to the inability to maintain consistent conditions such as light, weather, and other factors when acquiring bitemporal images, confounding factors such as the style of bitemporal data that are not related to change detection can cause detection difficulties. Therefore, a change detection method based on feature interaction and multitask learning (FMCD) is proposed in this article. To improve the ability to detect changes in complex scenes, FMCD models the context information of features through a multilevel feature interaction module, so as to obtain representative features, and to improve the sensitivity of the model to changes, the interaction between two temporal features is realized through the mix attention block (MAB). In addition, to eliminate the influence of weather and other factors, FMCD adopts a multitask learning strategy, takes domain adaptation as an auxiliary task, and maps the features of bitemporal images to the same space through the feature relationship adaptation module (FRAM) and feature distribution adaptation module (FDAM). Experiments on three datasets show that the proposed method is superior to other state-of-the-art methods. Chunhui Zhao 0003, Yingjie Tang, Shou Feng, Yuanze Fan, Wei Li 0032, Ran Tao 0003, Lifu Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Single-Source Domain Expansion Network for Cross-Scene Hyperspectral Image ClassificationabstractCurrently, cross-scene hyperspectral image (HSI) classification has drawn increasing attention. It is necessary to train a model only on source domain (SD) and directly transferring the model to target domain (TD), when TD needs to be processed in real time and cannot be reused for training. Based on the idea of domain generalization, a Single-source Domain Expansion Network (SDEnet) is developed to ensure the reliability and effectiveness of domain extension. The method uses generative adversarial learning to train in SD and test in TD. A generator including semantic encoder and morph encoder is designed to generate the extended domain (ED) based on encoder-randomization-decoder architecture, where spatial randomization and spectral randomization are specifically used to generate variable spatial and spectral information, and the morphological knowledge is implicitly applied as domain invariant information during domain expansion. Furthermore, the supervised contrastive learning is employed in the discriminator to learn class-wise domain invariant representation, which drives intra-class samples of SD and ED. Meanwhile, adversarial training is designed to optimize the generator to drive intra-class samples of SD and ED to be separated. Extensive experiments on two public HSI datasets and one additional multispectral image (MSI) dataset demonstrate the superiority of the proposed method when compared with state-of-the-art techniques. The codes will be available from the website:https://github.com/YuxiangZhang-BIT/IEEE_TIP_SDEnet. Yuxiang Zhang 0005, Wei Li 0032, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | Microscopic Hyperspectral Image Classification Based on Fusion Transformer With Parallel CNNabstractMicroscopic hyperspectral image (MHSI) has received considerable attention in the medical field. The wealthy spectral information provides potentially powerful identification ability when combining with advanced convolutional neural network (CNN). However, for high-dimensional MHSI, the local connection of CNN makes it difficult to extract the long-range dependencies of spectral bands. Transformer overcomes this problem well because of its self-attention mechanism. Nevertheless, transformer is inferior to CNN in extracting spatial detailed features. Therefore, a classification framework integrating transformer and CNN in parallel, named as Fusion Transformer (FUST), is proposed for MHSI classification tasks. Specifically, the transformer branch is employed to extract the overall semantics and capture the long-range dependencies of spectral bands to highlight the key spectral information. The parallel CNN branch is designed to extract significant multiscale spatial features. Furthermore, the feature fusion module is developed to effectively fuse and process the features extracted by the two branches. Experimental results on three MHSI datasets demonstrate that the proposed FUST achieves superior performance when compared with state-of-the-art methods. Weijia Zeng, Wei Li 0032, Mengmeng Zhang 0005, Hao Wang 0122, Yue Yang 0041, Ran Tao 0003 |
IEEE J. Biomed. Health Informatics | 7 |
| 2023 | Asymmetric Feature Fusion Network for Hyperspectral and SAR Image ClassificationabstractJoint classification using multisource remote sensing data for Earth observation is promising but challenging. Due to the gap of imaging mechanism and imbalanced information between multisource data, integrating the complementary merits for interpretation is still full of difficulties. In this article, a classification method based on asymmetric feature fusion, named asymmetric feature fusion network (AsyFFNet), is proposed. First, the weight-share residual blocks are utilized for feature extraction while keeping separate batch normalization (BN) layers. In the training phase, redundancy of the current channel is self-determined by the scaling factors in BN, which is replaced by another channel when the scaling factor is less than a threshold. To eliminate unnecessary channels and improve the generalization, a sparse constraint is imposed on partial scaling factors. Besides, a feature calibration module is designed to exploit the spatial dependence of multisource features, so that the discrimination capability is enhanced. Experimental results on the three datasets demonstrate that the proposed AsyFFNet significantly outperforms other competitive approaches. Wei Li 0032, Yunhao Gao, Mengmeng Zhang 0005, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Central Attention Network for Hyperspectral Imagery ClassificationabstractIn this article, the intrinsic properties of hyperspectral imagery (HSI) are analyzed, and two principles for spectral-spatial feature extraction of HSI are built, including the foundation of pixel-level HSI classification and the definition of spatial information. Based on the two principles, scaled dot-product central attention (SDPCA) tailored for HSI is designed to extract spectral-spatial information from a central pixel (i.e., a query pixel to be classified) and pixels that are similar to the central pixel on an HSI patch. Then, employed with the HSI-tailored SDPCA module, a central attention network (CAN) is proposed by combining HSI-tailored dense connections of the features of the hidden layers and the spectral information of the query pixel. MiniCAN as a simplified version of CAN is also investigated. Superior classification performance of CAN and miniCAN on three datasets of different scenarios demonstrates their effectiveness and benefits compared with state-of-the-art methods. Huan Liu 0015, Wei Li 0032, Xiang-Gen Xia 0001, Mengmeng Zhang 0005, Chenzhong Gao, Ran Tao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Hyperspectral and SAR Image Classification via Multiscale Interactive Fusion NetworkabstractDue to the limitations of single-source data, joint classification using multisource remote sensing data has received increasing attention. However, existing methods still have certain shortcomings when faced with feature extraction from single-source data and feature fusion between multisource data. In this article, a method based on multiscale interactive information extraction (MIFNet) for hyperspectral and synthetic aperture radar (SAR) image classification is proposed. First, a multiscale interactive information extraction (MIIE) block is designed to extract meaningful multiscale information. Compared with traditional multiscale models, it can not only obtain richer scale information but also reduce the model parameters and lower the network complexity. Furthermore, a global dependence fusion module (GDFM) is developed to fuse features from multisource data, which implements cross attention between multisource data from a global perspective and captures long-range dependence. Extensive experiments on the three datasets demonstrate the superiority of the proposed method and the necessity of each module for accuracy improvement. Wei Li 0032, Yunhao Gao, Mengmeng Zhang 0005, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Topological Structure and Semantic Information Transfer Network for Cross-Scene Hyperspectral Image ClassificationabstractDomain adaptation techniques have been widely applied to the problem of cross-scene hyperspectral image (HSI) classification. Most existing methods use convolutional neural networks (CNNs) to extract statistical features from data and often neglect the potential topological structure information between different land cover classes. CNN-based approaches generally only model the local spatial relationships of the samples, which largely limits their ability to capture the nonlocal topological relationship that would better represent the underlying data structure of HSI. In order to make up for the above shortcomings, a Topological structure and Semantic information Transfer network (TSTnet) is developed. The method employs the graph structure to characterize topological relationships and the graph convolutional network (GCN) that is good at processing for cross-scene HSI classification. In the proposed TSTnet, graph optimal transmission (GOT) is used to align topological relationships to assist distribution alignment between the source domain and the target domain based on the maximum mean difference (MMD). Furthermore, subgraphs from the source domain and the target domain are dynamically constructed based on CNN features to take advantage of the discriminative capacity of CNN models that, in turn, improve the robustness of classification. In addition, to better characterize the correlation between distribution alignment and topological relationship alignment, a consistency constraint is enforced to integrate the output of CNN and GCN. Experimental results on three cross-scene HSI datasets demonstrate that the proposed TSTnet performs significantly better than some state-of-the-art domain-adaptive approaches. The codes will be available from the website: https://github.com/YuxiangZhang-BIT/IEEE_TNNLS_TSTnet. Yuxiang Zhang 0005, Wei Li 0032, Mengmeng Zhang 0005, Ying Qu 0001, Ran Tao 0003, Hairong Qi 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Effective Velocity Calculation Method in Passive Synthetic Aperture for Emitter LocalizationabstractEmitter localization has been an active research subject in electronic reconnaissance, target tracking, emergency response, and satellite interference source localization. Recently, researchers applied passive synthetic aperture in the emitter localization to improve the positioning accuracy. Passive synthetic aperture uses Doppler rate and zero Doppler point to target range and azimuth locations, respectively. In spaceborne model, effective satellite velocity is used in the range equation to fit the passive synthetic aperture model. Therefore, this study proposed a convex optimization approach to calculate the effective satellite velocity. We established an objective function by calculating the relationship among beam footprint velocity, satellite velocity, and effective satellite velocity. The optimal effective satellite velocity and range distance are obtained through iterating different range distances. Finally, the experimental results show that the positioning accuracy of proposed method is one order of magnitude higher than that of traditional FOA and FDOA method. Hao Huan, Ran Tao 0003, Yue Wang 0001, Xiaogang Tang |
GLOBECOM | 3 |
| 2022 | Extracting and Distilling Direction-Adaptive Knowledge for Lightweight Object Detection in Remote Sensing ImagesabstractRecently, some lightweight convolutional neural network (CNN) models have been proposed for airborne or spaceborne remote sensing object detection (RSOD) tasks. However, these lightweight detectors suffer from performance degradation due to the compromise of limited computing resources on embedded devices. In order to narrow this performance gap, a direction-adaptive knowledge extraction and distillation (DKED) method is proposed. Specifically, a dynamic directional convolution (DDC) is developed to extract the typical arbitrary-oriented features, and a direction-adaptive knowledge distillation (DKD) strategy is designed for guiding the lightweight model to learn the intrinsic knowledge of the RSOD task from the high-performance model. Experiments on public datasets demonstrate that the proposed method can effectively improve the performance of the lightweight RSOD model without additional inference costs. Zhanchao Huang, Wei Li 0032, Ran Tao 0003 |
ICASSP | 3 |
| 2022 | Unimodular Waveform Design with Low Correlation Levels: A Fast Algorithm Development to Support Large-Scale Code LengthsabstractWe deal with the problem of unimodular waveform(s) design with low correlation levels for the case of large-scale code lengths that can reach tens of thousands. Our primary goals are to reduce the resulting complexity with high efficiency, and meanwhile, to ensure an integrated sidelobe level (ISL) or weighted ISL (WISL) of waveforms as low as possible. To this end, we study a generic model for the minimization of ISL/WISL, wherein the objective function is formulated to embed a Hadmard product into a high-order matrix norm. Our major contributions lie in the transformation of the objective into a proper form via multiple shift matrices and the reformulation of problem in order to use alternating direction method of multipliers (ADMM) technique. In particular, we introduce a virtual matrix to form an additional equality constraint for ADMM, whose augmented Lagrangian is elaborated to help derive a fast algorithm that iterates with closed-form solutions. Simulation results verify the superiority of our algorithm over existing state-of-the-art algorithms in different aspects. Yongzhe Li, Chunxuan Shi, Ran Tao 0003 |
ICASSP | 3 |
| 2022 | Geometric Low-Rank Tensor Approximation for Remotely Sensed Hyperspectral And Multispectral Imagery FusionabstractImproving the spatial resolution of a hyperspectral image (HSI) is of great significance in the remotely sensed field. By fusing a high-spatial-resolution multispectral image (MSI) with an HSI collected from the same scene, hyperspectral and multispectral (HS–MS) fusion has been an emerging technique to address the issue. Extracting complex spatial information from MSIs while maintaining abundant spectral information of HSIs is essential to generate the fused high-spatial-resolution HSI (HS2I). A common way is to learn low-rank/sparse representations from HSI and MSI, then reconstruct the fused HS2I based on tensor/matrix decomposition or unmixing paradigms, which ignore the intrinsic geometry proximity inherited by the low-rank property of the fused HS2I. This study proposes to estimate the high-resolution HS2I via low-rank tensor approximation with geometry proximity as side information learned from MSI and HSI by defined graph signals, which we name GLRTA. Row graph ${\mathcal{G}_r}$ and column graph ${\mathcal{G}_c}$ are defined on the horizontal slice and lateral slice of MSI tensor $\mathcal{M}$ respectively, while spectral band graph ${\mathcal{G}_b}$ is defined on a frontal slice of HSI tensor $\mathcal{H}$. Experimental results demonstrate that the proposed GLRTA can effectively improve the reconstruction results compared to other competitive works. Na Liu 0014, Wei Li 0032, Ran Tao 0003 |
ICASSP | 3 |
| 2022 | Dual Graph Cross-Domain Few-Shot Learning for Hyperspectral Image ClassificationabstractMost domain adaptation (DA) methods focus on the case where the source data (SD) and target data (TD) with the same classes are obtained by the same sensor in cross-scene hyperspectral image (HSI) classification tasks. However, the classification performance is significantly reduced when there are new classes in TD. In addition, domain alignment is carried out based on local spatial information in most methods, rarely taking into account the non-local spatial information (non-local relationships) with strong correspondence. A Dual Graph Cross-domain Few-shot Learning (DG-CFSL) framework is proposed, trying to make up for the above shortcomings by combining Few-shot Learning (FSL) with domain alignment. Both SD with all label samples and TD with a few label samples are implemented for FSL episodic training. Meanwhile, Intra-domain Distribution Extraction block (IDE-block) is designed to characterize and aggregate the intra-domain non-local relationships. Furthermore, feature- and distribution-level cross-domain graph alignments are used to mitigate the impact of domain shift on FSL. Experimental results on two public HSI data sets demonstrate the effectiveness of the proposed method. Yuxiang Zhang 0005, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003 |
ICASSP | 4 |
| 2022 | Multisource Cross-Scene Classification Using Fractional Fusion and Spatial-Spectral Domain AdaptationabstractTo solve the limitation of labeled samples in hyperspectral image (HSI) classification, cross-scene learning methods are developed recently. However, the disparity caused by environmental variation between HSI scenes is still a challenge. As a supplement, light detection and ranging (LiDAR) data provides elevation and spatial information regardless the variations. In this paper, we propose a multisource cross-scene classification method using fractional fusion and spatial-spectral domain adaptation to reduce disparity between scenes. The spatial information of HSI is preserved by fractional differential masks (FrDM) firstly. Then the LiDAR data is utilized for spectral alignment of HSI. The utilization of LiDAR data reduces the pixel-level disparity between scenes. At last, a spatial-spectral domain adaptation network is proposed for feature extraction and classification. Experimental results on HSI and LiDAR scenes show 5% improvements in overall accuracy compared with state-of-the-art methods. Xudong Zhao 0003, Mengmeng Zhang 0005, Ran Tao 0003, Wei Li 0032, Wenzi Liao, Wilfried Philips |
IGARSS | 3 |
| 2022 | Multisource Remote Sensing Data Classification Using Fractional Fourier TransformerabstractFocusing on joint classification of Hyperspectral image (HSI) and Light detection and ranging (LiDAR) data, a fractional Fourier image transformer (FrIT) is proposed as a backbone network in this paper. In the proposed FrIT, HSI and LiDAR data are firstly fused at pixel-level. Both multi-source and HSI feature extractors are utilized to capture local contexts. Then, a plug-and-play image transformer FrIT is explored for global contexts and sequential feature extraction. Unlike the attention-based representations in classic visual image transformer (VIT), FrIT is capable of speeding up the transformer architectures massively. To reduce the information loss from shallow to deep layers, FrIT is devised to connect contextual features in multiple fractional domains. At last, to evaluate the performance of FrIT, a new HSI and LiDAR benchmark is provided for extensive experiments, on which the proposed FrIT gains an improvement of 3% over state-of-the-art methods. Xudong Zhao 0003, Mengmeng Zhang 0005, Ran Tao 0003, Wei Li 0032, Wenzi Liao, Wilfried Philips |
IGARSS | 3 |
| 2022 | Retinex-Based Low-Light Hyperspectral Restoration Using Camera Response ModelabstractSpectral quality is one of the most critical issues that has to be considered in real hyperspectral image (HSI) application. Denoising, destriping, inpainting, deblurring and super-resolution are common techniques to improve the quality of HSIs from different aspects. These techniques have attracted much attention that a diversity of methods, algorithms, tools have been well developed to facilitate the development of HSI restoration. Although effectively improving the quality of HSIs, these technologies mainly focus on recovering an HSI captured in the normal sunlight. It is acknowledged that HSIs are captured via passive imaging mechanisms covering the spectral bands from visible& near-infrared to shortwave infrared spectral range (i.e., around 400nm to 2500nm). The imaging condition limits HSI spectrometers to capture HSIs without sunlight (e.g., in dark environments or night time). In this work, a low-light HSI restoration method is proposed, where we borrow idea of intrinsic decomposition based on Retinex theory in natural image low-light enhancement. Additionally, camera response function that describe the spectral degradation of RGB image and relationship between irradiance and pixel values are employed, respectively. The experimental results validate the effectiveness of the proposed method. Na Liu 0014, Yinjian Wang, Yixiao Yang, Wei Li 0032, Ran Tao 0003 |
IGARSS | 5 |
| 2022 | Mpanet: Multi-Patch Attention for Infrared Small Target Object DetectionabstractInfrared small target detection (ISTD) has attracted widespread attention and been applied in various fields. Due to the small size of infrared targets and the noise interference from complex backgrounds, the performance of ISTD using convolutional neural networks (CNNs) is restricted. Moreover, the constriant that long-distance dependent features can not be encoded by the vanilla CNNs also impairs the robustness of capturing targets' shapes and locations in complex scenarios. To this end, a multi-patch attention network (MPANet) based on the axial-attention encoder and the multi-scale patch branch (MSPB) structure is proposed. Specially, an axial-attention-improved encoder architecture is designed to highlight the effective features of small targets and suppress background noises. Furthermore, the developed MSPB structure fuses the coarse-grained and fine-grained features from different semantic scales. Extensive experiments on the SIRST dataset show the superiority performance and effectiveness of the proposed MPANet compared to the state-of-the-art methods. Wei Li 0032, Xin Wu 0001, Zhanchao Huang, Ran Tao 0003 |
IGARSS | 5 |
| 2022 | Strategy Designs for the Information Embedding of Joint MIMO Radar and Communications With SubarraysabstractIn this paper, we focus on designing strategies for information embedding (IE) of the joint MIMO radar and communications (JMRC) with subarrays. To fulfill the goal of IE as a secondary function for a MIMO radar with array division, we propose three designs which exploit the phase information conveyed by the construction of subarrays and/or the diversities of beampatterns synthesized by the array. We use the phase differences between subarrays to implement an index modulation for the IE in the first design, while we utilize the invariance property of beampattern rotations in the second design. The third design is a hybrid strategy, which combines the first and second designs together. For each design, the factors that affect the achievable communication data rate are analyzed with detailed derivations. Simulation results verify the effectiveness of our proposed IE designs for the JMRC. Yongzhe Li, Ran Tao 0003 |
WCNC | 3 |
| 2022 | Collaborative representation with background purification and saliency weight for hyperspectral anomaly detection
Zengfu Hou, Wei Li 0032, Ran Tao 0003, Pengge Ma, Weihua Shi |
Sci. China Inf. Sci. | 3 |
| 2022 | Blind adaptive identification and equalization using bias-compensated NLMS methods
Lijuan Jia, Ran Tao 0003, Yue Wang 0001 |
Sci. China Inf. Sci. | 3 |
| 2022 | Dynamic proximal unrolling network for compressive imaging
Yixiao Yang, Ran Tao 0003, Kaixuan Wei, Ying Fu 0001 |
Neurocomputing | 2 |
| 2022 | Passive Synthetic Aperture High-Precision Radiation Source Location by Single SatelliteabstractPassive localization is important because of its strong concealment and long detection distance. Space-borne passive radiation source location technologies have contradicting characteristics of wide coverage and high-precision position. Hence, accurate target identification is difficult to realize. A novel measurement that utilizes passive synthetic aperture is presented in this letter to measure the time of the satellite’s passing zenith precisely. In the beam coverage duration, the radiation carrier Doppler component is extracted and accumulated to synthesize an equivalent large virtual satellite antenna azimuth aperture. Taking advantage of the linear Doppler rate, a local matched filter is designed to implement a 2-D search. Doppler rate and zero Doppler point correspond to the target range and azimuth locations, respectively. The corresponding subastral point position radiation source can then be determined. Practical experiment results from an unmanned aerial vehicle platform demonstrated the effectiveness, wide coverage, and high-precision meter scale of the proposed method. Satellite data were also utilized to verify the feasibility of the system. Ang Li 0017, Hao Huan, Ran Tao 0003 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | A Novel Spatiotemporal Saliency Method for Low-Altitude Slow Small Infrared Target DetectionabstractThe effective monitoring of low-altitude slow small (LSS) targets represented by unmanned aerial vehicle (UAV) is a great challenge in the field of security in recent years. Most of the existing infrared (IR) small target algorithms focus on high-altitude target detection. However, the low-altitude background is complex and changeable, and high-intensity suspected targets exist widely. Existing methods usually cause high false alarm or failure detection for LSS targets. In this letter, we propose a novel spatiotemporal saliency method for LSS IR targets in image sequences. First, spatial variance saliency mapping and temporal gray saliency mapping are calculated in spatial domain and temporal domain, respectively. Then, the fusion saliency map is obtained by fusing the spatial saliency map and temporal saliency map. Finally, the target is extracted by a simple adaptive threshold segmentation. The proposed method is verified in five low-altitude IR image sequences. Experimental results demonstrate that the proposed method can achieve better detection performance than the existing state-of-the-art methods for LSS targets. Dongdong Pang, Tao Shan, Pengge Ma, Wei Li 0032, Shengheng Liu, Ran Tao 0003 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Low-Slow-Small Target Tracking Using Relocalization ModuleabstractWith the gradual opening of airspace, tracking of noncooperative low-altitude slow-speed small size (LSS) targets is important for the maintenance of security. It is still a challenging problem, especially for complex scenarios and real-time constraints. In this letter, an efficient tracking by relocalization (TRL) framework is proposed for small flying object tracking, aiming to alleviate the issue of losing moving targets in a complex background. Our designed relocalization module consists of a feature-aggregated module and a global search module. On the one hand, a feature-aggregated module is integrated into the designed framework to increase the ability to locate small targets. On the other hand, a global search module is developed to update the tracking performance, which attempts to address missed targets in long-term small object tracking tasks. What needs to be declared is that the basic tracking module cooperates with the relocalization module we designed to achieve the tracking of small targets. Performance evaluation of two small-flying target data sets and comparison with several state-of-the-art approaches demonstrate the effectiveness of the proposed framework. Wei Li 0032, Zhanchao Huang, Ran Tao 0003, Pengge Ma |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Visible-Assisted Infrared Image Super-Resolution Based on Spatial Attention Residual NetworkabstractInfrared images have a wide range of applications in military and civilian fields, including night vision, surveillance, and robotics. However, the most commonly used infrared images are low-resolution (LR), which lack texture details, and existing infrared image super-resolution (SR) algorithms are limited by the lack of spatial information utilization. To solve the above problems, a spatial attention residual network (SAResNet) is proposed. Specifically, the network consists of spatial attention residual block (SARB) with several short skip connections (SSCs). The SARB contains 20 spatial attention blocks (SAB), which adaptively adjusts weights of different spatial regions by considering interdependence between spatial features. Meanwhile, the visible images are considered as complementary sources; thus, a visible-assisted training strategy is designed for the infrared SR process, promoting details preservation. Furthermore, the spatial attention (SA) mechanism is utilized, which focuses more on spatial characteristics of the image and refines the main objects and target boundaries. Experimentally, the proposed method, SAResNet, is compared with existing SR methods, and the effectiveness of the proposed method is demonstrated based on both quantity and quality analyses. Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Phase retrieval from multiple FRFT measurements based on nonconvex low-rank minimization
Xinhua Su, Ran Tao 0003, Yongzhe Li |
Signal Process. | 2 |
| 2022 | Sliding Short-Time Fractional Fourier TransformabstractThe short-time fractional Fourier transform (STFRFT) has been shown to be a powerful tool for processing signals whose fractional frequencies vary with time. However, for real-time applications that require recalculating the STFRFT at each or several samples, the existing discrete algorithms are not suitable. To solve this problem, a new sliding algorithm is proposed, termed as the sliding STFRFT. First, the sliding STFRFT algorithm with the sliding step 1 is proposed. Then, it is derived to the circumstance when the sliding step turns to$\bm {p\;(p > 1)}$. The proposed sliding STFRFT algorithm directly computes the STFRFT at the time$\bm {m+1}$or$\bm {m+p}$using the STFRFT output result at the time$\bm {m}$, which greatly reduces the computation complexity. The theoretical analysis demonstrates that the proposed algorithm has the lowest computational cost among existing STFRFT algorithms. Gaowa Huang, Feng Zhang 0011, Ran Tao 0003 |
IEEE Signal Process. Lett. | 3 |
| 2022 | Adaptive Spatio-Temporal Tube for Fast Motion Segments Extraction of VideosabstractExisting motion segments extraction methods suffer from the problem of high computation complexity. To address this issue, we propose a method called adaptive spatio-temporal tube for fast motion segments extraction of videos. Firstly, initial spatio-temporal flow of sub-videos divided from input video is computed by adopting a novel Area-adjusted Spatio-Temporal Tunnel (A-STT) to screen preliminarily moiton segments. Secondly, the Sampling-line Adjustment Mechanism (SAM) is presented to avoid processing the entire amount of video spatial data and reduce computational complexity. The SAM is created by analyzing object consistency to produce a Sampling-line Adjustment Factor (SAF) which is used to dynamically obtain the sampling-line of various sub-videos. Finally, the adaptive spatio-temporal tubes are generated by integrating the initial spatio-temporal flow and SAF, which ensures the robustness of the proposed method. The proposed method is experimented on the public datasets VISOR, CAVIR and self-collected dataset. The experimental results demonstrate that the proposed method outperforms the state-of-the-art methods in terms of both computing speed and accuracy. Yunzuo Zhang, Kaina Guo, Ran Tao 0003 |
IEEE Signal Process. Lett. | 3 |
| 2022 | Spatial-Temporal Minimum Error Random Interaction Networks for Distributed EstimationabstractIn this letter, we study the problem of adaptive parameter estimation for distributed wireless sensor networks (WSNs), where the non-Gaussian impulsive noises in the communication links are considered. In such cases, if the traditional distributed collaborative strategy is still adopted, the estimation performance of the network will decline significantly. Aiming at this problem, we propose a new spatial-temporal minimum error random interaction strategy for the distributed estimation in the presence of impulsive link noises. Furthermore, the maximum correlation entropy criterion and the stochastic gradient descent method are used to update the combination factor so that the network has dynamic and real-time adaptability to the link noises. The proposed algorithm is also compared with the classic DLMS algorithm and a state-of-the-art algorithm designed for the impulsive link noises. Simulation results show that the proposed algorithm can not only reduce the network traffic effectively, but also be robust to the non-Gaussian link impulsive noises while maintaining its advantage over the non-cooperative algorithm and achieving the optimal estimation performance. Lijuan Jia, Zi-Jiang Yang, Ran Tao 0003 |
IEEE Signal Process. Lett. | 4 |
| 2022 | MS-HLMO: Multiscale Histogram of Local Main Orientation for Remote Sensing Image RegistrationabstractMulti-source image registration is challenging due to intensity, rotation, and scale differences among the images. Considering the characteristics and differences of multi-source remote sensing images, a feature-based registration algorithm named Multi-scale Histogram of Local Main Orientation (MS-HLMO) is proposed. Harris corner detection is first adopted to generate feature points. The HLMO feature of each Harris feature point is extracted on a Partial Main Orientation Map (PMOM) with a Generalized Gradient Location and Orientation Histogram-like (GGLOH) feature descriptor, which provides high intensity, rotation, and scale invariance. The feature points are matched through a multi-scale matching strategy. Comprehensive experiments on 17 multi-source remote sensing scenes demonstrate that the proposed MS-HLMO and its simplified version MS-HLMO+outperform other competitive registration algorithms in terms of effectiveness and generalization. Chenzhong Gao, Wei Li 0032, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Hyperspectral and Multispectral Classification for Coastal Wetland Using Depthwise Feature Interaction NetworkabstractThe monitoring of coastal wetlands is of great importance to the protection of marine and terrestrial ecosystems. However, due to the complex environment, severe vegetation mixture, and difficulty of access, it is impossible to accurately classify coastal wetlands and identify their species with traditional classifiers. Despite the integration of multisource remote sensing data for performance enhancement, there are still challenges with acquiring and exploiting the complementary merits from multisource data. In this article, the depthwise feature interaction network (DFINet) is proposed for wetland classification. A depthwise cross attention module is designed to extract self-correlation and cross correlation from multisource feature pairs. In this way, meaningful complementary information is emphasized for classification. DFINet is optimized by coordinating consistency loss, discrimination loss, and classification loss. Accordingly, DFINet reaches the standard solution-space under the regularity of loss functions, while the spatial consistency and feature discrimination are preserved. Comprehensive experimental results on two hyperspectral and multispectral wetland datasets demonstrate that the proposed DFINet outperforms other competitive methods in terms of overall accuracy. Yunhao Gao, Wei Li 0032, Mengmeng Zhang 0005, Jianbu Wang, Weiwei Sun 0005, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Infrared Attention Network for Woodland Segmentation Using Multispectral Satellite ImagesabstractSemantic segmentation of the remote sensing images (RSIs) has attracted increasing interest in recent years. However, large-area segmentation of the woodland presents challenges. The wide distribution and diverse tree species of the woodland make feature extraction difficult. For this reason, an infrared attention network (InfAttNet) is proposed to extract woodland from multispectral RSIs. InfAttNet has an extra infrared spectral encoder which makes use of the sensitivity of vegetation to near infrared and red edge spectrums. This extra encoder applies learning about vegetation to improve woodland segmentation. Several attention blocks are designed to enhance learning about vegetation features and so improve the performance. In addition, a new dataset is built, containing a large number of woodland RSIs and covering several typical woodland distribution regions in China. The experimental results demonstrate that, compared with other networks, InfAttNet has the highest accuracy and is capable of rapid extraction of the woodland in RSIs. Yuanyuan Gui, Wei Li 0032, Xiang-Gen Xia 0001, Ran Tao 0003, Anzhi Yue |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Hyperspectral Change Detection Based on Multiple Morphological ProfilesabstractWith the increasing availability of multitemporal hyperspectral imagery, hyperspectral change detection under heterogeneous backgrounds is a challenging task. Due to the complexity of background features, traditional change detection algorithms in the spectral domain cannot effectively detect changed features. A novel method using multiple morphological profiles (MMPs) is proposed for hyperspectral change detection to make full use of spatial information. In the designed framework, first, the max-tree/min-tree strategy is applied to extract different attributes of multitemporal hyperspectral images (HSIs), i.e., area attribute and height attribute. Second, a spectral angle weighted-based local absolute distance (SALA) method is designed to reconstruct the discriminative spectral domain. Then, the absolute distance (AD) is adopted to extract changes in constructed feature domain. Finally, a change map is obtained by guided filtering. Experiments conducted on four real hyperspectral datasets demonstrate that the proposed detector achieves better detection performance. Zengfu Hou, Wei Li 0032, Lu Li 0005, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Spatial-Spectral Weighted and Regularized Tensor Sparse Correlation Filter for Object Tracking in Hyperspectral VideosabstractHyperspectral video camera captures spatial, spectral and temporal information of moving objects. Traditional object tracking methods developed for color videos have been applied to hyperspectral videos after compressing hundreds of spectral bands into three, which does not fully utilize the wealth spectral information. In order to address this issue, we present a tensor sparse correlation filter with a spatial-spectral weighted regularizer for object tracking. First, tensor processing is employed to reduce the spectral differences of homogeneous background, thereby producing robust spectral structure features. Second, a spatial-spectral weighted regularizer is designed in the correlation filter framework to penalize filter template by suppressing spectral features dissimilar to the center pixel in tracking. Third, a sparse constraint term and tracking context information are incorporated to suppress unexpected peaks in the response map. Finally, a reformulated stacked HOG feature extractor and a two-dimensional adaptive scale search strategy are developed to further improve the tracker’s feature discrimination and scale adaptation capability. Experimental results demonstrate that the proposed method achieves superior tracking performance than traditional correlation filter-based trackers. Zengfu Hou, Wei Li 0032, Jun Zhou 0001, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | A Novel Nonlocal-Aware Pyramid and Multiscale Multitask Refinement Detector for Object Detection in Remote Sensing ImagesabstractObject detection (OD) is an important task of computer vision and has been widely used in many fields, including remote sensing (RS). However, the complex scenes, large-scale variation, and dense instances of RS bring huge challenges to OD. To meet these challenges, a novel Nonlocal-aware Pyramid and Multiscale Multitask Refinement Detector (NPMMR-Det) is proposed. Specifically, nonlocal-aware pyramid attention (NP-Attention) is designed for guiding a neural network model to focus more on efficient features and suppress background noise. Then a multiscale refinement feature pyramid network (MSR-FPN) is proposed to fuse the multiscale context features extracted by the NP-Attention guided neural network and adjust the optimal receptive field. In order to use these features more effectively, a multitask refinement head called MTR-Head, with offset sharing and a modulation mechanism, is developed to refine the feature misalignment between the localization task and the classification task. Extensive experiments performed on two public RS data sets demonstrate that the proposed NPMMR-Det achieves competitive performance compared with state-of-the-art methods. Zhanchao Huang, Wei Li 0032, Xiang-Gen Xia 0001, Xin Wu 0001, Zhaoquan Cai 0001, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | LO-Det: Lightweight Oriented Object Detection in Remote Sensing ImagesabstractA few lightweight convolutional neural network (CNN) models have been recently designed for remote sensing object detection (RSOD). However, most of them simply replace vanilla convolutions with stacked separable convolutions (SConvs), which may not be efficient due to a lot of precision losses and may not be able to detect oriented bounding boxes (OBBs). Also, the existing OBB detection methods are difficult to constrain the shape of objects predicted by CNNs accurately. In this article, we propose an effective lightweight oriented object detector (LO-Det). Specifically, a channel separation-aggregation (CSA) structure is designed to simplify the complexity of SConvs, and a dynamic receptive field (DRF) mechanism is developed to maintain high accuracy by customizing the convolution kernel and its perception range dynamically when reducing the network complexity. The CSA-DRF component optimizes efficiency while maintaining high accuracy. Then, a diagonal support constraint head (DSC-Head) component is designed to detect OBBs and constrain their shapes more accurately and stably. Extensive experiments on public data sets demonstrate that the proposed LO-Det can run very fast even on embedded devices with the competitive accuracy of detecting oriented objects. Zhanchao Huang, Wei Li 0032, Xiang-Gen Xia 0001, Hao Wang 0122, Feiran Jie, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Graph-Feature-Enhanced Selective Assignment Network for Hyperspectral and Multispectral Data ClassificationabstractDue to rich spectral and spatial information, the combination of hyperspectral and multispectral images (MSIs) has been widely used for Earth observation, such as wetland classification. However, mining of meaningful features and effective fusion of multisource remote sensing data are still urgent problems to be solved. In this article, graph-feature-enhanced selective assignment network (GSANet) is proposed. On the one hand, a graph feature extraction module (GFEM) is designed to extract topological structure information and combine with the rich spectral–spatial information. In particular, the features obtained by convolution are first mapped to the graph feature space, and the graph convolution operation is used to achieve propagation between nodes for preserving topological structure information. Moreover, to reduce the difference of graph features resulting from the mapping function and better explore the complementary properties of multisource data, a novel graph fusion strategy-graph dependence fusion is designed. A transition graph is generated to enhance the association and interaction between different graph features, so as to avoid the information loss caused by simple fusion operation. On the other hand, a selective feature assignment module (SFAM) is developed to adaptively assign weights to different discriminative features. SFAM assigns weights to different features to selectively emphasize informative features and suppress less useful ones. Extensive experiments are conducted on two multisource remote sensing datasets, and the improvement of at least 1.27% and 0.98% compared to other state-of-the-art work demonstrates the superiority of the proposed GSANet. Wei Li 0032, Yunhao Gao, Mengmeng Zhang 0005, Ran Tao 0003, Bing Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Multigraph-Based Low-Rank Tensor Approximation for Hyperspectral Image RestorationabstractLow-rank-tensor-approximation (LRTA)-based hyperspectral image (HSI) restoration has drawn increasing attention. However, most of the methods construct a hidden low-rank tensor by utilizing the non-local self-similarity (NLSS) and global spectral correlation (GSC) inherited by HSIs. Although achieving state-of-the-art (SOTA) restoration performance, NLSS and GSC have limitations. NLSS is introduced from natural image denoising to remove spatially independent identically distributed (i.i.d.) Gaussian and impulse noise. While GSC, which is naturally possessed by HSIs, is adopted to maintain the spectral integrity and remove spectrally, i.i.d., degradations. Therefore, NLSS and GSC may not be successfully used for complex HSI restoration tasks, such as destriping, cloud removal and recovery of atmospheric absorption bands. To solve the issue, borrowing the idea from manifold learning, the geometry information characterized by proximity relationship, is integrated with the LRTA to solve the above issue, named as multi-graph-based LRTA (MGLRTA). Different with most of the existing methods, the proposed MGLRTA directly models an HSI as a low-rank tensor and efficiently explores the extra proximity information on the defined graphs that are not only inherited by the low-rank constraints but also naturally possessed in HSIs. A well-posed iterative algorithm is designed to solve the restoration problem. Experimental results on different datasets that cover several severe degradation scenarios demonstrate that the proposed MGLRTA outperforms the SOTA HSI restoration methods. Na Liu 0014, Wei Li 0032, Ran Tao 0003, Qian Du 0001, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | STTM-SFR: Spatial-Temporal Tensor Modeling With Saliency Filter Regularization for Infrared Small Target DetectionabstractDetecting small infrared (IR) targets against low-altitude complex background is always a challenge for IR search and tracking (IRST) system due to limited small target characteristics, the moving background caused by camera motion, and extremely cluttered backgrounds. The existing methods usually cause high false alarm or do not work against the chaotic low-altitude complex background. In this article, a novel spatial–temporal tensor model with saliency filter regularization (STTM-SFR) is developed to detect small IR targets. First, the small target detection task is transformed into a sparse and low-rank tensor optimization problem using the spatial–temporal prior knowledge of background and target. The construction of the holistic STTM can retain the complete spatial–temporal information of the original IR image sequence. Then, the SFR term limited between background and foreground aims to promote target saliency learning. That is to say, the SFR term can avoid the offset approximation of the low-rank tensor, so as to recover a clean target image from the original IR tensor. Finally, an effective alternating direction method of multipliers (ADMM) algorithm framework is designed to solve the proposed STTM-SFR model. The effectiveness and robustness of the STTM-SFR model are verified in six real IR scenes. Experimental results show that our method outperforms other baseline methods. Moreover, the proposed STTM-SFR method is more robust than the existing state-of-the-art STTMs against low-altitude moving backgrounds. Dongdong Pang, Pengge Ma, Tao Shan, Wei Li 0032, Ran Tao 0003, Yueran Ma, Tianrun Wang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Facet Derivative-Based Multidirectional Edge Awareness and Spatial-Temporal Tensor Model for Infrared Small Target DetectionabstractInfrared (IR) small target detection in the complex background is an important but challenging research hotspot in the field of target detection. The existing methods usually cause high false alarms in the complex background and fail to make full use of the complete information of the image. In this article, a novel IR small target detection model that combines facet derivative-based multidirectional edge awareness with spatial–temporal tensor (FDMDEA-STT) is presented. First, we construct an STT model (STTM) to transform the target detection problem into a low-rank and sparse tensor optimization problem based on the prior information of the target and background in the spatial–temporal domain. Then, based on the facet derivative, we define a multidirectional edge awareness mapping and fuse it into the STTM as sparse prior information. Finally, an effective algorithm based on the alternating direction method of multipliers (ADMM) is designed to solve the above model. The effectiveness of the proposed method is verified on eight real IR image sequences. Experimental results demonstrate that the proposed method has better detection performance than the existing state-of-the-art methods. Dongdong Pang, Tao Shan, Wei Li 0032, Pengge Ma, Ran Tao 0003, Yueran Ma |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Human Activity Classification Based on Micro-Doppler Signatures SeparationabstractHuman activity classification based on micro-Doppler (m-D) signatures finds applications in surveillance, search and rescue operations, and healthcare. In this article, we propose a new approach for human activity classification. This approach deals with the situations of reduced limb movements that could be due to the presence of injury or an individual carrying objects. It applies a preprocessing step to separate human m-D signals of the limbs from the Doppler signal corresponding to the torso. The separated m-D signal is input to a two-layer convolutional principal component analysis network (CPCAN) for feature extraction and motion classification. The CPCAN comprises a simple network architecture for efficient training and implementation, and it automatically learns the highly discriminative features. Experiments involving multiple human subjects performing different activities show a high classification accuracy associated with small arm motions. Xingshuai Qiao, Moeness G. Amin, Tao Shan, Zhengxin Zeng, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Radar Point Clouds Processing for Human Activity Classification Using Convolutional Multilinear Subspace LearningabstractRadar-based human activity classification is crucial for applications such as healthcare monitoring, fall detection, and assisted living due to its superior sensing capabilities and privacy protection. Traditional classification methods generally retrieve features from the time-range domain or the time-frequency (TF) domain. Such 2-D representation neglects the underlying dependence between the three radar signal variables of time, range, and Doppler frequency, and cannot fully depict the dynamic human motion features. In this article, we propose a time-range-Doppler radar point clouds (RPCs)-based learning model for human activity classification using a frequency-modulated continuous waveform (FMCW) radar. The human echoes are first transformed into a series of 3-D point cloud cubes integrating the motion signatures in three domains, namely time-range, time-Doppler, and range-Doppler domains. The generated RPC cubes are then fed into a newly developed two-layer convolutional multilinear principal component analysis network (CMPCANet) for feature extraction and motion classification. The CMPCANet comprises a simple network architecture with small training parameters, and can be directly implemented on the 3-D tensor dataset to extract highly discriminative features. Experimental results demonstrate that proposed framework can achieve superior classification accuracy and noise robustness compared to other methods using multidomain information, even with small training samples. Xingshuai Qiao, Shengheng Liu, Tao Shan, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Human Activity Classification Based on Moving Orientation Determining Using Multistatic Micro-Doppler Radar SignalsabstractTraditional micro-Doppler (m-D)-based human activity classification system using monostatic radar suffers from the drawback that classification performance is vulnerable to the variation of human motion aspect angle. This leads to a performance degradation if the human movements are not directly toward or away with respect to the radar line of sight. The multistatic radar system has been suggested as an effective solution to solve the problem, as it can observe the target from multiple views and achieve favorable aspect angles to the targets. In this article, a novel human activity classification method based on motion orientation determining using multistatic m-D signals is proposed. First, the aspect angles of target motion direction with respect to each radar nodes are inferred by using the proposed motion orientation estimation method. The multistatic m-D data are then divided into several intervals based on the measured angle, and the data in the same interval are fused at the data level. Finally, the classification results are obtained through the adaptive weighted decision-level fusion. Compared with the traditional multistatic classification method, due to the consideration of the time-varying human motion aspect angle, the proposed method is more reasonable in data fusion and has better classification performance. Xingshuai Qiao, Gang Li 0008, Tao Shan, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Information Fusion for Classification of Hyperspectral and LiDAR Data Using IP-CNNabstractJoint use of multisensor information has attracted considerable attention in the remote sensing community. While applications in land-cover observation benefit from information diversity, multisensor integration technique is confronted with many challenges, including inconsistent size of data, different data structures, uncorrelated physical properties, and scarcity of training data. In this article, an information fusion network, named interleaving perception convolutional neural network (IP-CNN), is proposed for integrating heterogeneous information and improving joint classification performance of hyperspectral image (HSI) and light detection and ranging (LiDAR) data. Specifically, a bidirectional autoencoder is designed to reconstruct hyperspectral and LiDAR data together, and the reconstruction process is trained with no dependence upon annotated information. Both HSI-perception constraint and LiDAR-perception constraint are imposed on multisource structural information integration. Accordingly, fused data are fed into a two-branch CNN for final classification. To validate the effectiveness of the model, the experiments were conducted using three datasets (i.e., Muufl Gulfport data, Trento data, and Houston data). The final results demonstrate that the proposed framework can significantly outperform state-of-the-art methods even with small-size training samples. Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003, Heng-Chao Li 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Three-Order Tensor Creation and Tucker Decomposition for Infrared Small-Target DetectionabstractExisting infrared small-target detection methods tend to perform unsatisfactorily when encountering complex scenes, mainly due to the following: 1) the infrared image itself has a low signal-to-noise ratio (SNR) and insufficient detailed/texture knowledge; 2) spatial and structural information is not fully excavated. To avoid these difficulties, an effective method based on three-order tensor creation and Tucker decomposition (TCTD) is proposed, which detects targets with various brightness, spatial sizes, and intensities. In the proposed TCTD, multiple morphological profiles, i.e., diverse attributes and different shapes of trees, are designed to create three-order tensors, which can exploit more spatial and structural information to make up for lacking detailed/texture knowledge. Then, Tucker decomposition is employed, which is capable of estimating and eliminating the major principal components (i.e., most of the background) from three dimensions. Thus, targets can be preserved on the remaining minor principal components. Image contrast is further enhanced by fusing the detection maps of multiple morphological profiles and several groups with discontinuous pruning values. Extensive experiments validated on two synthetic data and six real data sets demonstrate the effectiveness and robustness of the proposed TCTD. Mingjing Zhao, Wei Li 0032, Lu Li 0005, Pengge Ma, Zhaoquan Cai 0001, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Hyperspectral Target Detection Based on Weighted Cauchy Distance Graph and Local Adaptive Collaborative RepresentationabstractHyperspectral target detection in complex backgrounds is a challenging and important research topic in the remote sensing field. Traditional target detectors consider the background spectrum to obey a Gaussian distribution. However, this distribution may not meet the requirements in real hyperspectral images. In addition, the background and spatial information of most existing target detection algorithms are rarely fully utilized. Therefore, a new weighted Cauchy distance graph (WCDG) and local adaptive collaborative representation detection (CGCRD) is proposed. First, a WCDG similarity measure is designed. In order to adjust the effect of target pixels on the graph model, a weighted Cauchy distance Laplace matrix is constructed, and then the matrix is applied to the matched filter detector. Second, local adaptive collaborative representation strategy is developed. The penalty coefficient is weighted by the local spatial Euclidean distance combined with the Pearson correlation coefficient, and then the detection result is obtained based on the residual. Finally, aforementioned two strategies are fused to fully utilize the spatial and spectral information. A 176-band hyperspectral image (BIT-HSI-I) dataset is collected for the target detection task. The related algorithms are performed on the BIT-HSI-I dataset, and the detection results demonstrate that the proposed algorithm has better detection performance than other state-of-the-art algorithms. Xiaobin Zhao, Wei Li 0032, Chunhui Zhao 0003, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Fractional Gabor Convolutional Network for Multisource Remote Sensing Data ClassificationabstractRemote sensing using multisensor platforms has been systematically applied for monitoring and optimizing human activities. Several advanced techniques have been developed to enhance and extract the spatially and spectrally semantic information in the hyperspectral image (HSI) and light detection and ranging (LiDAR) data processing and analysis. However, an abundance of redundant information and sometimes a lack of discriminative features reduce the efficiency and effectiveness of multisource classification methods. This article proposes a fractional Gabor convolutional network (FGCN), focusing on efficient feature fusion and comprehensive feature extraction. First, the proposed FGCN uses Octave convolution layers to perform multisource information fusion and preserve discriminative information. Second, fractional Gabor convolutional (FGC) layers are proposed to extract multiscale, multidirectional, and semantic change features. The completeness and discrimination of the multisource features using different FGC kernels are improved, which yield robust feature extraction against semantic changes. Finally, the fractional Gabor feature and spectral feature are combined with two weighting factors which can be learned during the network training. Experimental results and comparisons with state-of-the-art multisource classification methods indicate the effectiveness of the proposed FGCN. With the FGCN, we can obtain an 89.90% overall accuracy on the challenging Muufl Gulfport (MUUFL) data set, with an improvement of 3% over state-of-the-art methods. Xudong Zhao 0003, Ran Tao 0003, Wei Li 0032, Wilfried Philips, Wenzi Liao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A General Gaussian Heatmap Label Assignment for Arbitrary-Oriented Object DetectionabstractRecently, many arbitrary-oriented object detection (AOOD) methods have been proposed and attracted widespread attention in many fields. However, most of them are based on anchor-boxes or standard Gaussian heatmaps. Such label assignment strategy may not only fail to reflect the shape and direction characteristics of arbitrary-oriented objects, but also have high parameter-tuning efforts. In this paper, a novel AOOD method called General Gaussian Heatmap Label Assignment (GGHL) is proposed. Specifically, an anchor-free object-adaptation label assignment (OLA) strategy is presented to define the positive candidates based on two-dimensional (2D) oriented Gaussian heatmaps, which reflect the shape and direction features of arbitrary-oriented objects. Based on OLA, an oriented-bounding-box (OBB) representation component (ORC) is developed to indicate OBBs and adjust the Gaussian center prior weights to fit the characteristics of different objects adaptively through neural network learning. Moreover, a joint-optimization loss (JOL) with area normalization and dynamic confidence weighting is designed to refine the misalign optimal results of different subtasks. Extensive experiments on public datasets demonstrate that the proposed GGHL improves the AOOD performance with low parameter-tuning and time costs. Furthermore, it is generally applicable to most AOOD methods to improve their performance including lightweight models on embedded platforms. Zhanchao Huang, Wei Li 0032, Xiang-Gen Xia 0001, Ran Tao 0003 |
IEEE Trans. Image Process. | 4 |
| 2022 | Prior-Based Tensor Approximation for Anomaly Detection in Hyperspectral ImageryabstractThe key to hyperspectral anomaly detection is to effectively distinguish anomalies from the background, especially in the case that background is complex and anomalies are weak. Hyperspectral imagery (HSI) as an image–spectrum merging cube data can be intrinsically represented as a third-order tensor that integrates spectral information and spatial information. In this article, a prior-based tensor approximation (PTA) is proposed for hyperspectral anomaly detection, in which HSI is decomposed into a background tensor and an anomaly tensor. In the background tensor, a low-rank prior is incorporated into spectral dimension by truncated nuclear norm regularization, and a piecewise-smooth prior on spatial dimension can be embedded by a linear total variation-norm regularization. For anomaly tensor, it is unfolded along spectral dimension coupled with spatial group sparse prior that can be represented by the${l}_{2,1}$-norm regularization. In the designed method, all the priors are integrated into a unified convex framework, and the anomalies can be finally determined by the anomaly tensor. Experimental results validated on several real hyperspectral data sets demonstrate that the proposed algorithm outperforms some state-of-the-art anomaly detection methods. Lu Li 0005, Wei Li 0032, Ying Qu 0001, Chunhui Zhao 0003, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Multipixel Anomaly Detection With Unknown Patterns for Hyperspectral ImageryabstractIn this article, anomaly detection is considered for hyperspectral imagery in the Gaussian background with an unknown covariance matrix. The anomaly to be detected occupies multiple pixels with an unknown pattern. Two adaptive detectors are proposed based on the generalized likelihood ratio test design procedure and ad hoc modification of it. Surprisingly, it turns out that the two proposed detectors are equivalent. Analytical expressions are derived for the probability of false alarm of the proposed detector, which exhibits a constant false alarm rate against the noise covariance matrix. Numerical examples using simulated data reveal how some system parameters (e.g., the background data size and pixel number) affect the performance of the proposed detector. Experiments are conducted on five real hyperspectral data sets, demonstrating that the proposed detector achieves better detection performance than its counterparts. Jun Liu 0004, Zengfu Hou, Wei Li 0032, Ran Tao 0003, Danilo Orlando, Hongbin Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Waveform Design for the Joint MIMO Radar and Communications with Low Integrated Sidelobe Levels and Accurate Information EmbeddingabstractIn this paper, we focus on the multiple-waveform design for the joint multiple-input multiple-output radar and communications system, which aims to simultaneously attain low integrated sidelobe level (ISL) of waveforms and accurate fast-time modulation for information embedding (IE). We propose a novel strategy to exploit the attributes of waveform phases with extremely large degrees of freedom for embedding communication symbols, based on which we formulate the generalized waveform design into a nonconvex optimization problem. Our major contribution lies in converting both the ISL related objective and the fast-time modulation related constraints for IE into tractable quadratic forms. To achieve this, we introduce a novel diagonal matrix with Toeplitz blocks to reformulate and then relax the problem into a form that involves the outer product of the waveform vector in its objective. In order to solve this problem, we exploit the majorization-minimization technique to devise an algorithm that enables a closed-form solution at each iteration. Simulation results verify the effectiveness of our design. Yongzhe Li, Ran Tao 0003 |
ICASSP | 3 |
| 2021 | Feature Exchange for Multisource Data Classification in Wetland SceneabstractWetland classification is of great significance for monitoring. Recently, collaborative analysis of multisource data has received special attention considering the limitations of single source data. In this paper, a wetland classification method based on feature exchange is proposed. Firstly, the weighting shared residual blocks are utilized for feature extraction. Then, the scaling factors in batch normalization (BN) self-determine the redundancy of current channel, which is replaced by another channel when the scaling factor is less than the threshold. To eliminate unnecessary channels and improve the generalization, sparsity constraint is employed on partial scaling factors. Experimental results on multisource wetland dataset demonstrate that the proposed method outperforms other competitive works. Yunhao Gao, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003 |
IGARSS | 4 |
| 2021 | Hyperspectral Image Super-Resolution Based on Multiscale Residual Block and Multilevel Feature FusionabstractHyperspectral images have high spectral resolution, but this is often at the expense of spatial resolution. Although deep learning-based super-resolution (SR) algorithms have shown comparative performance for spatial resolution enhancement, most of them cannot effectively extract features of different size objects because of single scale convolution. In deep architectures, low level features also tend to disappear during transmission. In this paper, an efficient network (MRBMFF) for enhancing the spatial resolution of hyperspectral image is proposed. Based on the multiscale residual block (MRB), features at different scales can be effectively extracted and fused. Meanwhile, the multilevel feature fusion (MFF) is introduced to concatenate the low and high level features. Effective SR images could be recovered after inputting their low-resolution counterparts to the proposed network. Experimental results show that the proposed network achieves superior reconstruction performance compared with the state-of-the-art approaches. Feng Zhang 0011, Wei Li 0032, Ran Tao 0003 |
IGARSS | 5 |
| 2021 | Domain Adaptation Based on Graph and Statistical Features for Cross-Scene Hyperspectral Image ClassificationabstractCross-scene hyperspectral image (HSI) classification has gradually received widespread attention, because most models perform unsatisfactory classification performance on training and testing samples from two different scenes. At present, the domain adaptation technique is used to solve this problem, most of which only design models from the level of data statistical features, while ignore the potential topological relationships between the land cover classes. In order to make up for the above shortcoming, a domain adaptation based on graph and statistical features is proposed in the papaer. This method uses convolutional neural network (CNN) extracting features with rich semantic information to dynamically construct graphs, and further introduces graph optimal transport (GOT) to align topological relations to assist distribution alignment based on maximum mean discrepancy (MMD). The experimental results on two cross-scene HSI datasets demonstrate the effectiveness of the proposed method. Yuxiang Zhang 0005, Wei Li 0032, Ran Tao 0003 |
IGARSS | 3 |
| 2021 | Infrared Small-Target Detection Based on Three-Order Tensor Creation and Tucker DecompositionabstractRobust infrared small-target detection has always been a research hotspot in target search and tracking systems. However, the image itself has low signal-to-noise ratio (SNR), and the targets usually lack detailed/texture information. In addition, the background is complex and diverse. All the above factors make it easy for the targets to be submerged. In this paper, a novel method is proposed based on a three-order creation and the Tucker decomposition. First, the morphological profiles (i.e., area attribute and height attribute of max-tree) are applied to create a three-order tensor, which compensates for the lack of detailed information by supplementing spatial information in infrared images. Then, the Tucker decomposition is employed on the created tensor, in which most of the background can be estimated and eliminated from three dimensions. Finally, the target is detected on the remaining and the results of diverse morphological profiles are fused, which further enhances the target information. Experimental results demonstrate the effectiveness of the proposed method. Mingjing Zhao, Wei Li 0032, Lu Li 0005, Ran Tao 0003 |
IGARSS | 4 |
| 2021 | Joint feature extraction for multi-source data using similar double-concentrated network
Yixuan Zhu, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003, Qian Du 0001 |
Neurocomputing | 5 |
| 2021 | The hopping discrete fractional Fourier transform
Yu Liu 0033, Feng Zhang 0011, Hongxia Miao, Ran Tao 0003 |
Signal Process. | 4 |
| 2021 | A general fraction-of-time probability framework for chirp cyclostationary signals
Hongxia Miao, Feng Zhang 0011, Ran Tao 0003 |
Signal Process. | 3 |
| 2021 | Performance evaluation and parameter optimization of sparse Fourier transform
Hongchi Zhang, Tao Shan, Shengheng Liu, Ran Tao 0003 |
Signal Process. | 4 |
| 2021 | Low-Rank and Sparse Decomposition With Mixture of Gaussian for Hyperspectral Anomaly DetectionabstractRecently, the low-rank and sparse decomposition model (LSDM) has been used for anomaly detection in hyperspectral imagery. The traditional LSDM assumes that the sparse component where anomalies and noise reside can be modeled by a single distribution which often potentially confuses weak anomalies and noise. Actually, a single distribution cannot accurately describe different noise characteristics. In this article, a combination of a mixture noise model with low-rank background may more accurately characterize complex distribution. A modified LSDM, by modeling the sparse component as a mixture of Gaussian (MoG), is employed for hyperspectral anomaly detection. In the proposed framework, the variational Bayes (VB) algorithm is applied to infer a posterior MoG model. Once the noise model is determined, anomalies can be easily separated from the noise components. Furthermore, a simple but effective detector based on the Manhattan distance is incorporated for anomaly detection under complex distribution. The experimental results demonstrate that the proposed algorithm outperforms the classic Reed-Xiaoli (RX), and the state-of-the-art detectors, such as robust principal component analysis (RPCA) with RX. Lu Li 0005, Wei Li 0032, Qian Du 0001, Ran Tao 0003 |
IEEE Trans. Cybern. | 4 |
| 2021 | Hyperspectral Image Restoration Using Adaptive Anisotropy Total Variation and Nuclear NormsabstractRandom Gaussian noise and striping artifacts are common phenomena in hyperspectral images (HSI). In this article, an effective restoration method is proposed to simultaneously remove Gaussian noise and stripes by merging a denoising and a destriping submodel. A denoising submodel performs a multiband denoising, i.e., Gaussian noise removal, considering Gaussian noise variations between different bands, to restore the striped HSI from the corrupted image, in which the striped HSI is constrained by a weighted nuclear norm. For the destriping submodel, we propose an adaptive anisotropy total variation method to adaptively smoothen the striped HSI, and we apply, for the first time, the truncated nuclear norm to constrain the rank of the stripes to 1. After merging the above two submodels, an ultimate image restoration model is obtained for both denoising and destriping. To solve the obtained optimization problem, the alternating direction method of multipliers (ADMM) is carefully schemed to perform an alternative and mutually constrained execution of denoising and destriping. Experiments on both synthetic and real data demonstrate the effectiveness and superiority of the proposed approach. Wei Li 0032, Na Liu 0014, Ran Tao 0003, Feng Zhang 0011, Paul Scheunders |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Hyperspectral Restoration and Fusion With Multispectral Imagery via Low-Rank Tensor-ApproximationabstractTensor-based fusion that couples the high spatial resolution of a multispectral image (MSI) to the high spectral resolution of a hyperspectral image (HSI) is considered. The fusion problem is first formulated mathematically as a convex optimization of a tensor trace norm imposing low-rank spatially as well as spectrally, with an alternating-directions optimization featuring linearization providing the solution. Although prior tensor-based fusion approaches typically resort to tensor decomposition, the proposed algorithm exploits ideas from the field of tensor completion to directly impose a low-rank property spatially and spectrally while avoiding the computationally complex patch clustering and dictionary learning common to competing fusion techniques. Additionally, small modifications to the basic optimization permit a fusion process robust to missing hyperspectral values such as those that can result from dead stripes in real hyperspectral sensors. The experimental evaluations on both synthetic imagery as well as real imagery demonstrate that the resulting low-rank tensor-approximation (LRTA) fusion algorithm preserves both spatial details and texture, yielding significantly improved image quality when compared to other state-of-the-art fusion methods as well as effective restoration under conditions of missing stripes within the HSI. Na Liu 0014, Lu Li 0005, Wei Li 0032, Ran Tao 0003, James E. Fowler, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Cross-Scene Hyperspectral Image Classification With Discriminative Cooperative AlignmentabstractCross-scene classification is one of the major challenges for hyperspectral image (HSI) classification, especially for target scenes without label samples. Most traditional domain adaptive methods learn a domain invariant subspace to reduce statistical shift while ignoring the fact that there may not exist a shared subspace when marginal distributions of source and target domains are very different. In addition, it is important for HSI classification to preserve discriminant information in the original space. To solve this issue, discriminative cooperative alignment (DCA) of subspace and distribution is proposed to cooperatively reduce the geometric and statistical shift. In the proposed framework, both geometrical and statistical alignments are considered to learn subspaces of the two domains with preserving discrimination information. Furthermore, a reconstruction constraint is imposed to enhance the robustness of subspace projection. Experimental results on three cross-scene HSI data sets demonstrate that the proposed DCA is significantly better than some state-of-the-art domain-adaptive approaches. Yuxiang Zhang 0005, Wei Li 0032, Ran Tao 0003, Jiangtao Peng, Qian Du 0001, Zhaoquan Cai 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Infrared Small-Target Detection Based on Multiple Morphological ProfilesabstractInfrared small-target detection under heterogeneous background, as a challenging task, plays an important role in many applications. In practice, there are not only bright targets but also dim targets, e.g., rescue aircraft and vehicles in the forest fire scene. Considering that most existing infrared small-target detection methods are merely aimed at bright targets, a novel method using multiple morphological profiles (MMP) is proposed, which can detect various types of targets whose brightness varies greatly. In the designed morphological feature extraction, different attributes, i.e., area attribute and height attribute, are applied to extract spatial size and contrast information of small-target in the max-tree and min-tree, respectively. Furthermore, discontinuous pruning values are further utilized for different attributes, and a designed fusion strategy of different pruning values results in more robust detection performance. Experimental results validated on two synthetic data and six real data sets demonstrate that the proposed MMP can not only detect a variety of brightness of targets and different types of targets and kinds of spatial sizes of targets but also further improve the contrast between targets and background, and the background clutter is significantly suppressed. Mingjing Zhao, Lu Li 0005, Wei Li 0032, Ran Tao 0003, Liwei Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Discriminant Tensor-Based Manifold Embedding for Medical Hyperspectral ImageryabstractMedical hyperspectral imagery has recentlyattracted considerable attention. However, for identification tasks, the high dimensionality of hyperspectral images usually leads to poor performance. Thus, dimensionality reduction (DR) is crucial in hyperspectral image analysis. Motivated by exploiting the underlying structure information of medical hyperspectral images and enhancing the discriminant ability of features, a discriminant tensor-based manifold embedding (DTME) is proposed for discriminant analysis of medical hyperspectral images. Based on the idea of manifold learning, a new discriminant similarity metric is designed, which takes into account the tensor representation, sparsity, low-rank and distribution characteristics. Then, an inter-class tensor graph and an intra-class tensor graph are constructed using the new similarity metric to reveal intrinsic manifold of hyperspectral data. Dimensionality reduction is achieved by embedding this supervised tensor graphs into the low-dimensional tensor subspace. Experimental results on membranous nephropathy and white bloodcells identification tasks demonstrate the potential clinical value of the proposed DTME. Wei Li 0032, Tianhong Chen, Jun Zhou 0001, Ran Tao 0003 |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | Spatial-Spectral Density Peaks-Based Discriminant Analysis for Membranous Nephropathy Classification Using Microscopic Hyperspectral ImagesabstractThe traditional differential diagnosis of membranous nephropathy (MN) mainly relies on clinical symptoms, serological examination and optical renal biopsy. However, there is a probability of false positives in the optical inspection results, and it is unable to detect the change of biochemical components, which poses an obstacle to pathogenic mechanism analysis. Microscopic hyperspectral imaging can reveal detailed component information of immune complexes, but the high dimensionality of microscopic hyperspectral image brings difficulties and challenges to image processing and disease diagnosis. In this paper, a novel classification framework, including spatial-spectral density peaks-based discriminant analysis (SSDP), is proposed for intelligent diagnosis of MN using a microscopic hyperspectral pathological dataset. SSDP constructs a set of graphs describing intrinsic structure of MHSI in both spatial and spectral domains by employing density peak clustering. In the process of graph embedding, low-dimensional features with important diagnostic information in the immune complex are obtained by compacting the spatial-spectral local intra-class pixels while separating the spectral inter-class pixels. For the MN recognition task, a support vector machine (SVM) is used to classify pixels in the low-dimensional space. Experimental validation data employ two types of MN that are difficult to distinguish with optical microscope, including primary MN and hepatitis B virus-associated MN. Experimental results show that the proposed SSDP achieves a sensitivity of 99.36%, which has potential clinical value for automatic diagnosis of MN. Wei Li 0032, Ran Tao 0003, Nigel H. Lovell, Yue Yang 0041, Tianqi Tu, Wenge Li |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Convolutional Neural Network for Coastal Wetland Classification in Hyperspectral ImageabstractClassifying different land cover types with hyperspectral image (HSI) is significant for restoring and protecting natural resources and maintaining ecological services in coastal wetlands. This paper proposes a multi-domain features fusion convolutional neural network (MDF-CNN) based classification method for hyperspectral images of coastal wetlands. This method adopts inter-class sparsity based discriminative least square regression (ICS_DLSR) to learn a more compact and discriminative transformation, as well as fuse the high-level features of the original domain and the regression domain to obtain higher classification accuracy. Experimental results demonstrate the effectiveness of the proposed method when compared with some recent classifiers. The MDF-CNN achieved state-of-the-art performance on two latest GF-5 HSI datasets of Coastal Wetland. Mengmeng Zhang 0005, Wei Li 0032, Weiwei Sun 0005, Ran Tao 0003 |
IGARSS | 5 |
| 2020 | Hyperspectral Target Detection by Fractional Fourier TransformabstractTarget detection in hyperspectral images (HSI) is an important technique and many target detection algorithms have been developed in recent years. The most widely detection algorithms by the original spectral characteristics may lack the ability of target signal enhancement and background suppression. This paper presents an efficient algorithm for detecting hyperspectral targets based on fractional Fourier transform (FrFT). Firstly, fractional Fourier transform primary search is used as preprocessing to obtain the better intermediate domain features with complementary characteristics between the original reflection spectrum and the Fourier transform domain. Secondly, fractional Fourier transform secondary search and constrained energy minimization (FrFT-CEM) was adopted to find an optimal fractional order to distinguish the target from the background. The proposed method has been proved to be superior in two real hyperspectral data sets. Xiaobin Zhao, Wei Li 0032, Tao Shan, Lu Li 0005, Ran Tao 0003 |
IGARSS | 5 |
| 2020 | Sea-Ice Classification Based on Optical Image Using Morphological Profile FeaturesabstractSea-ice classification plays an important role in evaluating sea-ice hazards and ensuring maritime safety. In this paper, a method of sea-ice classification based on morphological feature extraction is proposed by using CEBRS-02B multispectral CCD image. The paper uses local contain profile (LCP) to extract morphological features of the multispectral image. Then SVM- and MRF- (SVMMRF) is adopted, which including probabilistic support vector machine (SVM) for the preliminary classification of multispectral image, the postprocessing by using Markov random field (MRF) based regularization. Experimental results demonstrate the validity of the classification model framework is applied to sea-ice classification. Compared with the traditional Local Binary Pattern (LBP) and Gabor, feature extraction by LCP can improve the accuracy of sea-ice classification. Yuchan Zhou, Wei Li 0032, Peng Ren 0001, Ran Tao 0003 |
IGARSS | 5 |
| 2020 | Collaborative Classification for Woodland Data Using Similar Multi-concentrated Network
Yixuan Zhu, Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003, Qiong Ran |
PRCV (2) | 4 |
| 2020 | The FrFT convolutional face: toward robust face recognition using the fractional Fourier transform and convolutional neural networks
Xin Wu 0001, Ran Tao 0003, Danfeng Hong, Yue Wang 0001 |
Sci. China Inf. Sci. | 2 |
| 2020 | Fourier-Based Rotation-Invariant Feature Boosting: An Efficient Framework for Geospatial Object DetectionabstractGeospatial object detection (GOD) of remote sensing imagery has been attracting increasing interest in recent years, due to the rapid development in spaceborne imaging. Most of the previously proposed object detectors are very sensitive to object deformations, such as scaling and rotation. To this end, we propose a novel and efficient framework for GOD in this letter, called Fourier-based rotation-invariant feature boosting (FRIFB). A Fourier-based rotation-invariant feature is first generated in polar coordinate. Then, the extracted features can be further structurally refined using aggregate channel features. This leads to a faster feature computation and more robust feature representation, which is good fitting for the coming boosting learning. Finally, in the test phase, we achieve a fast pyramid feature extraction by estimating a scale factor instead of directly collecting all features from the image pyramid. Extensive experiments are conducted on two subsets of NWPU VHR-10 data set, demonstrating the superiority and effectiveness of the FRIFB compared to the previous state-of-the-art methods. Xin Wu 0001, Danfeng Hong, Jocelyn Chanussot, Yang Xu 0006, Ran Tao 0003, Yue Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2020 | Mutual information rate of nonstationary statistical signals
Hongxia Miao, Feng Zhang 0011, Ran Tao 0003 |
Signal Process. | 3 |
| 2020 | Optimized sparse fractional Fourier transform: Principle and performance analysis
Hongchi Zhang, Tao Shan, Shengheng Liu, Ran Tao 0003 |
Signal Process. | 4 |
| 2020 | Clustered fractional Gabor transform
Zhichun Zhao, Ran Tao 0003, Gang Li 0008, Yue Wang 0001 |
Signal Process. | 2 |
| 2020 | Fractional Power Spectrum and Fractional Correlation Estimations for Nonuniform SamplingabstractThis letter proposes new estimations of fractional power spectral density (FrPSD) and fractional correlation function (FrCF) for nonuniform sampling of random signals with non-stationarity and limited bandwidths in the fractional Fourier domain. Unlike previous works, the developed FrPSD and FrCF estimations are capable of dealing with unknown sampling instants. In order to obtain them, we first formulate approximations of FrCF and FrPSD making use of uniform sampling instants. Then we convert the approximate FrPSD to a fractional filtered version of the FrPSD for the original random signal, which does not rely on the sampling instants. With such operations, we propose the FrPSD estimation to cancel the bias of FrPSD approximation by means of a fractional inverse filtering and thereby obtain a high accuracy of it. The FrCF estimation is proposed to be the inverse fractional Fourier transform of the FrPSD, and it serves as the fractional interpolation of the previously obtained approximation of the FrCF. Simulation results show the effectiveness of the proposed estimation methods. Ran Tao 0003, Yongzhe Li, Xuejing Kang |
IEEE Signal Process. Lett. | 2 |
| 2020 | Novel Second-Order Statistics of the Chirp Cyclostationary SignalsabstractIn communications and radar/sonar systems, one of the most popular nonstationary stochastic signal models is the chirp cyclostationary (CCS) signal. Recently, the second-order statistics of complex CCS signals are defined based on the conjugate correlation function, which have shown to be more proper in linear canonical transform domain than in the frequency domain. However, the information provided by unconjugate correlation function is missing, which leads to an incomplete study of second-order CCS signals. In this letter, the unified conjugate and unconjugate correlation function of CCS signals is researched. In detail, the definitions of the novel second-order statistics, chirp cyclic correlation and chirp cyclic spectrum, are expounded. Then the properties of these statistics are investigated. Finally, the additional information provided by the correlation function with conjugate operation and the usefulness of the proposed statistics are introduced and explained by simulations. Hongxia Miao, Feng Zhang 0011, Ran Tao 0003 |
IEEE Signal Process. Lett. | 3 |
| 2020 | Spatial-Spectral Feature Extraction via Deep ConvLSTM Neural Networks for Hyperspectral Image ClassificationabstractIn recent years, deep learning has presented a great advance in the hyperspectral image (HSI) classification. Particularly, long short-term memory (LSTM), as a special deep learning structure, has shown great ability in modeling long-term dependencies in the time dimension of video or the spectral dimension of HSIs. However, the loss of spatial information makes it quite difficult to obtain better performance. In order to address this problem, two novel deep models are proposed to extract more discriminative spatial-spectral features by exploiting the convolutional LSTM (ConvLSTM). By taking the data patch in a local sliding window as the input of each memory cell band by band, the 2-D extended architecture of LSTM is considered for building the spatial-spectral ConvLSTM 2-D neural network (SSCL2DNN) to model long-range dependencies in the spectral domain. To better preserve the intrinsic structure information of the hyperspectral data, the spatial-spectral ConvLSTM 3-D neural network (SSCL3DNN) is proposed by extending LSTM to the 3-D version for further improving the classification performance. The experiments, conducted on three commonly used HSI data sets, demonstrate that the proposed deep models have certain competitive advantages and can provide better classification performance than the other state-of-the-art approaches. Wen-Shuai Hu, Heng-Chao Li 0001, Lei Pan 0003, Wei Li 0032, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2020 | Discriminative Marginalized Least-Squares Regression for Hyperspectral Image ClassificationabstractLeast-squares regression (LSR)-based classifiers are effective in multiclassification tasks. However, most existing methods use limited projections, resulting in loss of much discriminant information; furthermore, they focus only on exactly fitting samples to target matrix while ignoring overfitting issue. To solve these drawbacks, discriminative marginalized LSR (DMLSR) is proposed to learn a more discriminative projection matrix with consideration of class separability and data-reconstruction ability simultaneously. In the proposed framework, an intraclass compactness graph is employed to avoid the overfitting problem and enhance class separability, and a data-reconstruction constraint is imposed to preserve discriminant information on limited projections. Experimental results on several hyperspectral data sets demonstrate that the proposed method significantly outperforms some state-of-the-art classifiers. Yuxiang Zhang 0005, Wei Li 0032, Heng-Chao Li 0001, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Joint Classification of Hyperspectral and LiDAR Data Using Hierarchical Random Walk and Deep CNN ArchitectureabstractEarth observation using multisensor data is drawing increasing attention. Fusing remotely sensed hyperspectral imagery and light detection and ranging (LiDAR) data helps to increase application performance. In this article, joint classification of hyperspectral imagery and LiDAR data is investigated using an effective hierarchical random walk network (HRWN). In the proposed HRWN, a dual-tunnel convolutional neural network (CNN) architecture is first developed to capture spectral and spatial features. A pixelwise affinity branch is proposed to capture the relationships between classes with different elevation information from LiDAR data and confirm the spatial contrast of classification. Then in the designed hierarchical random walk layer, the predicted distribution of dual-tunnel CNN serves as global prior while pixelwise affinity reflects the local similarity of pixel pairs, which enforce spatial consistency in the deeper layers of networks. Finally, a classification map is obtained by calculating the probability distribution. Experimental results validated with three real multisensor remote sensing data demonstrate that the proposed HRWN significantly outperforms other state-of-the-art methods. For example, the two branches CNN classifier achieves an accuracy of 88.91% on the University of Houston campus data set, while the proposed HRWN classifier obtains an accuracy of 93.61%, resulting in an improvement of approximately 5%. Xudong Zhao 0003, Ran Tao 0003, Wei Li 0032, Heng-Chao Li 0001, Qian Du 0001, Wenzi Liao, Wilfried Philips |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | MATNet: Motion-Attentive Transition Network for Zero-Shot Video Object SegmentationabstractIn this paper, we present a novel end-to-end learning neural network, i.e., MATNet, for zero-shot video object segmentation (ZVOS). Motivated by the human visual attention behavior, MATNet leverages motion cues as a bottom-up signal to guide the perception of object appearance. To achieve this, an asymmetric attention block, named Motion-Attentive Transition (MAT), is proposed within a two-stream encoder network to firstly identify moving regions and then attend appearance learning to capture the full extent of objects. Putting MATs in different convolutional layers, our encoder becomes deeply interleaved, allowing for close hierarchical interactions between object apperance and motion. Such a biologically-inspired design is proven to be superb to conventional two-stream structures, which treat motion and appearance independently in separate streams and often suffer severe overfitting to object appearance. Moreover, we introduce a bridge network to modulate multi-scale spatiotemporal features into more compact, discriminative and scale-sensitive representations, which are subsequently fed into a boundary-aware decoder network to produce accurate segmentation with crisp boundaries. We perform extensive quantitative and qualitative experiments on four challenging public benchmarks, i.e., DAVIS16, DAVIS17, FBMS and YouTube-Objects. Results show that our method achieves compelling performance against current state-of-the-art ZVOS methods. To further demonstrate the generalization ability of our spatiotemporal learning framework, we extend MATNet to another relevant task: dynamic visual attention prediction (DVAP). The experiments on two popular datasets (i.e., Hollywood-2 and UCF-Sports) further verify the superiority of our model. Our implementations have been made publicly available at https://github.com/tfzhou/MATNet. Tianfei Zhou, Jianwu Li, Shunzhou Wang, Ran Tao 0003, Jianbing Shen |
IEEE Trans. Image Process. | 4 |
| 2020 | Blood Cell Classification Based on Hyperspectral Imaging With Modulated Gabor and CNNabstractCell classification, especially that of white blood cells, plays a very important role in the field of diagnosis and control of major diseases. Compared to traditional optical microscopic imaging, hyperspectral imagery, combined with both spatial and spectral information, provides more wealthy information for recognizing cells. In this paper, a novel blood cell classification framework, which combines a modulated Gabor wavelet and deep convolutional neural network (CNN) kernels, named as MGCNN, is proposed based on medical hyperspectral imaging. For each convolutional layer, multi-scale and orientation Gabor operators are taken dot product with initial CNN kernels. The essence is to transform the convolutional kernels into the frequency domain to learn features. By combining characteristics of Gabor wavelets, the features learned by modulated kernels at different frequencies and orientations are more representative and discriminative. Experimental results demonstrate that the proposed model can achieve better classification performance than traditional CNNs and widely used support vector machine approaches, especially as training small-sample-size situations. Wei Li 0032, Baochang Zhang 0001, Qingli Li, Ran Tao 0003, Nigel H. Lovell |
IEEE J. Biomed. Health Informatics | 5 |
| 2019 | Hyperspectral Image Super-resolution Using Generative Adversarial Network and Residual LearningabstractDue to the limitation of image acquisition, hyperspectral remote sensing imagery is hard to reflect in both high spatial and spectral resolutions. Super-resolution (SR) is a technique which can improve the spatial resolution. Inspired by recent achievements in deep convolutional neural network (CNN) and generative adversarial network (GAN), a GAN based framework is proposed for hyperspectral image super-resolution. In the proposed method, residual learning is used to obtain a high metrics and spectral fidelity, and a shorter connection is set between the input layer and output layer. The gradient features from low-resolution (LR) image to high-resolution (HR) are utilized as auxiliary information to assist deep CNN to carry out counter training with discriminator. Experimental results demonstrate that the proposed SR algorithm achieves superior performance in spectral fidelity and spatial resolution compared with baseline methods. Wei Li 0032, Ran Tao 0003 |
ICASSP | 4 |
| 2019 | Multisource Remote Sensing Data Classification Using Deep Hierarchical Random Walk NetworksabstractCollaborative classification of hyperspectral imagery (HSI) and light detection and ranging (LiDAR) data is investigated using effective hierarchical random walk networks, denoted as HRWN. The proposed HRWN jointly optimizes dual-tunnel CNN, pixelwise affinity and seeds map via a novel random walk layer, which enforces spatial consistency in the deepest layers of the network. In designed random walk layer, the predicted distribution of dual-tunnel CNN serves as global prior while pixelwise affinity reflects local similarity of pixel pairs, which preserves boundary localization and spatial consistency well. Experimental results validated with two real multisource remote sensing data demonstrate that the proposed HRWN can significantly outperform other state-of-art methods. Xudong Zhao 0003, Ran Tao 0003, Wei Li 0032 |
ICASSP | 2 |
| 2019 | Collaborative Classification of Hyperspectral and Lidar Data With Information Fusion and Deep NetsabstractConvolutional neural network (CNN) receives extensive attention in hyperspectral image classification. While hyper-spectral images contain abundant spectral information but lack spatial information, which usually contributes to poor classification results. In this paper, a novel classification framework called information fusion based CNN (IF-CNN) is proposed to compensate for the shortcomings of hyper-spectral images. The proposed method merges hyperspectral images with abundant spectral information and LiDAR images with rich spatial information as the input of classification framework. Furthermore, the framework consists of two convolutional neural networks: one-dimensional CNN for extracting spectral features, and two-dimensional CNN for extracting spatial correlation features. Experimental results demonstrate that the proposed method achieves excellent performance compared with some existing methods. Chen Chen 0001, Xudong Zhao 0003, Wei Li 0032, Ran Tao 0003, Qian Du 0001 |
IGARSS | 4 |
| 2019 | Improved Multiresolution Analysis Method for Hyperspectral PansharpeningabstractThe fusion of Panchromatic (PAN) and Hyperspectral image (HSI) aims at improving resolution in spatial and spectral domain simultaneously. Multiresolution analysis is a widely used method for Mutispectral or HSI pansharpening. However, only detail information from PAN is considered while ignoring the detail information from HSI. In this paper, an improved approach based on multiresolution analysis is proposed, which extracts detail information from both PAN and HSI by choosing optimal multiresolution layers. Another contribution is that we discuss the weight when fusing the detail information. The experimental results demonstrate that the proposed method can provide better quality metrics and visual effects when compared with some existing methods. Xiuxiu Hu, Yan Shi 0012, Wei Li 0032, Ran Tao 0003 |
IGARSS | 4 |
| 2019 | LW-ODF: A Light-Weight Object Detection Framework for Optical Remote Sensing ImageryabstractIn this paper, we propose to extract the multi-scaled and rotation-insensitive deep features to address the issues of object multi-solutions and rotations in geospatial object detection. To this end, we develop a novel object detection framework where a rotation-insensitive convolution neural network is applied for extracting multi-scaled and direction-insensitive feature representation and then the learned features can be fed into the ensemble classifier learning with fast feature pyramid. Such a non-end-to-end learning strategy intuitively reduces the computational cost without the additional performance loss, yielding an effective and efficient light-weight object detection framework. Experimental results conducted on the NWPU VHR-10 dataset demonstrate that the proposed framework outperforms several state-of-the-art baselines. Xin Wu 0001, Danfeng Hong, Pedram Ghamisi, Wei Li 0032, Ran Tao 0003 |
IGARSS | 5 |
| 2019 | A Weakly-Supervised Deep Network for DSM-Aided Vehicle DetectionabstractWith the breakthrough of the spatial resolution of optical remote sensing images at the sub-meter level and the explosive development of deep learning, geospatial object detection has achieved a growing interest in remote sensing community. However, labeling large training datasets in object level is still an expensive and tedious procedure. This might lead to the poor model generalization and degraded network learning ability. To this end, a weakly-supervised deep network (WSDN) is developed for geospatial object detection by applying a digital surface model (DSM)-aided auto-labeling and a pre-trained network learned from the task-independent dataset. Experimental results conducted on the stereo aerial imagery of a large camping site are performed to demonstrate that the proposed WSDN yields better detection results, with 62.78% precision and 55.13% recall. Xin Wu 0001, Danfeng Hong, Jiaojiao Tian, Ralph Kiefl, Ran Tao 0003 |
IGARSS | 5 |
| 2019 | Variation of a signal in Schwarzschild spacetime
Huan Liu 0015, Xiang-Gen Xia 0001, Ran Tao 0003 |
Sci. China Inf. Sci. | 3 |
| 2019 | Analysis and comparison of discrete fractional fourier transforms
Xinhua Su, Ran Tao 0003, Xuejing Kang |
Signal Process. | 2 |
| 2019 | Sliding 2D Discrete Fractional Fourier TransformabstractThe two-dimensional discrete fractional Fourier transform (2D DFrFT) has been shown to be a powerful tool for 2D signal processing. However, the existing discrete algorithms aren't the optimal for real-time applications, where the input signals are stream data arriving in a sequential manner. In this letter, a new sliding algorithm is proposed to solve this problem, termed as the 2D sliding DFrFT (2D SDFrFT). The proposed 2D SDFrFT algorithm directly computes the 2D DFrFT in current window using the results of previous window, which greatly reduces the computations. During the derivation, we find that the (m + δ, n)th DFrFT bin in previous window is needed for computing the (m, n)th DFrFT bin in current window, where the increment δ isn't always an integer. Further, a method is proposed to convert the increment δ to a certain integer by determining appropriate sampling interval. The theoretical analysis demonstrates that when compute the new 2D DFrFT in a shifted window in sliding process, our proposed algorithm has the lowest computational cost among existing 2D DFrFT algorithms. Yu Liu 0033, Hongxia Miao, Feng Zhang 0011, Ran Tao 0003 |
IEEE Signal Process. Lett. | 4 |
| 2019 | Multi-Branch Spatial-Temporal Network for Action RecognitionabstractHuman action recognition based on deep-learning methods have received increasing attention and developed rapidly. However, current methods suffer from the confusion caused by convolving over time and space independently, processing shorter sequences, restricted to single temporal scale modeling and so on. The key objective of precisely classifying actions is to capture the appearance and motion throughout entire videos. Based on this purpose, a multi-branch spatial-temporal network (MSTN) is proposed. It consists of a multi-branch deep network and a long-term feature (LTF) layer. Benefits of the proposed MSTN include: (a) the multi-branch spatial-temporal network aims at encoding spatial and temporal information simultaneously, and (b) the LTF layer is used to aggregate the video-level representation with multiple temporal scales. Evaluations on two action datasets and comparison with several state-of-the-art approaches demonstrate the effectiveness of the proposed network. Wei Li 0032, Ran Tao 0003 |
IEEE Signal Process. Lett. | 3 |
| 2019 | Reality-Preserving Multiple Parameter Discrete Fractional Angular Transform and Its Application to Color Image EncryptionabstractIn this paper, we first define a reality-preserving multiple parameter fractional angular transform (RPMPDFrAT), which is a useful tool for image encryption. Then, we propose a new color image encryption algorithm based on the defined RPMPDFrAT. The encryption process consists of two phases: encryption in the spatial domain and RPMPDFrAT domain. In the spatial domain, three color components of the plain image are mapped by dual cylindrical transform, which can nonlinearly hide the original color information. Then, the intermediate output is scrambled by a coupled logistic map to reduce the correlation of adjacent pixels and uniformly distribute the image energy of different color components. Thereafter, the scrambled image is transformed by the proposed RPMPDFrAT, which can ensure that we obtain the real-value output. Finally, a process similar to the spatial domain is performed in the RPMPDFrAT domain to further improve the security of the cryptosystem. Numerical simulations are performed and demonstrate that the proposed image encryption algorithm is effective and sensitive to keys. Moreover, some potential attacks are tested to verify the robustness of the proposed method, and the performance of our method outperforms previously published ones. Xuejing Kang, Anlong Ming, Ran Tao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Color Image Encryption Using Pixel Scrambling Operator and Reality-Preserving MPFRHTabstractTo ensure the confidentiality of color images during their storage or transmission on insecure networks, a number of encryption methods based on fractional transforms have been proposed and widely investigated. However, most of their outputs are complex values that are inconvenient for record and transmission. Also, those methods always deal with a whole color component with the same fractional-order and ignore the textural features that are contained in different image parts. In this paper, we first define a reality-preserving multiple-parameter fractional Hartley transform (RPMPFRHT), whose output is real value, and then a novel color image encryption method is proposed that divides the RGB components into different blocks and uses a constructed pixel scrambling operator to mix and hide the original color information. The outputs are transformed to different RPMPFRHT domains and further scrambled by a non-adjacent coupled map lattices system. Numerical simulations are performed to demonstrate that the proposed encryption algorithm is feasible, secure, sensitive to keys, and robust to potential attacks. Xuejing Kang, Ran Tao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Nonconvex Truncated Nuclear Norm Minimization Based on Adaptive Bisection MethodabstractThe explosive growth in high-dimensional visual data requires effective regularization techniques to utilize the underlying low-dimensional structure. We consider low-rank matrix recovery, and many existing approaches are based on the nuclear norm regularization. Recently, truncated nuclear norm (TNNR) has been proposed to achieve a better approximation to the rank function than that of the traditional nuclear norm. TNNR was defined by the nuclear norm by subtracting the sum of the largest r singular values. However, the estimation of r is not trivial. In addition, the original algorithm based on TNNR only considers the matrix completion cases and requires double loops, which is not quite computationally efficient. Correspondingly, in this paper, we propose the adaptive bisection method to adaptively estimate r, which can efficiently reduce the cost of computation. Moreover, to further accelerate computing, we apply iteratively reweighted nuclear norm to solve the nonconvex TNNR directly, and the convergence can also be guaranteed. Finally, we extend the applications of TNNR from the matrix completion problems to the general low-rank matrix recovery. Extensive experiments validate the superiority of the proposed algorithm over the state-of the-art methods. Xinhua Su, Xuejing Kang, Ran Tao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Microwave Radiometer Data Superresolution Using Image Degradation and Residual NetworkabstractMicrowave radiometers are the key sensors to globally monitor environmental parameters; however, it suffers from its low and nonuniform spatial resolution. In this paper, a superresolution (SR) technique based on image degradation and residual network is proposed to enhance the spatial resolution of microwave radiometer data. Specifically, an improved degradation model is proposed to construct pairs of high-resolution (HR) and low-resolution (LR) data for training and testing. In addition, a new residual network connected by the SR main and gradient auxiliary branches in parallel is designed to achieve SR reconstructions, where eight-channel gradient maps extracted from LR data are input into the auxiliary branch to help to reconstruct. SR results are eventually generated by the trained SR network. Experiments executed on both simulated and actual data demonstrate the soundness and the superiority of the proposed SR technique. Feng Zhang 0011, Wei Li 0032, Weidong Hu, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | Structure-Aware Collaborative Representation for Hyperspectral Image ClassificationabstractRecently, collaborative representation (CR) has drawn increasing attention in hyperspectral image classification due to its simplicity and effectiveness. However, existing representation-based classifiers do not explicitly utilize class label information of training samples in estimating representation coefficients. To solve this issue, a structure-aware CR with Tikhonov regularization (SaCRT) method is proposed to consider both class label information of training samples and spectral signatures of testing pixels to estimate more discriminative representation coefficients. In the proposed framework, marginal regression is employed; furthermore, an interclass row-sparsity structure is designed to preserve the compact relationship among intraclass pixels and more separable interclass pixels, thereby enhancing class separability. The experimental results evaluated using three hyperspectral data sets demonstrate that the proposed method significantly outperforms some state-of-the-art classifiers. Wei Li 0032, Yuxiang Zhang 0005, Na Liu 0014, Qian Du 0001, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | Wavelet-Domain Low-Rank/Group-Sparse Destriping for Hyperspectral ImageryabstractPushbroom acquisition of hyperspectral imagery is prone to striping artifacts in the along-track direction. A hyperspectral destriping algorithm is proposed such that the subbands of a 3-D wavelet transform most affected by pushbroom stripes-namely, those with spatially vertical orientation-are the exclusive focus of destriping. The proposed method features an iterative image decomposition composed of a low-rank model for the stripes coupled with a group-sparse prior on the wavelet coefficients of the subbands in question. While low-rank stripe models have been widely used in the past, they typically have been deployed in conjunction with a total-variation prior on the image that is prone to oversmoothing and residual stripe artifacts. On the other hand, the proposed group-sparse prior not only captures the well-known sparse nature of wavelet coefficients but also capitalizes on their vertical clustering in the subbands in question. In addition, while many prior destriping methods are wavelet-based, they employ 2-D transforms band by band. In contrast, the proposed 3-D wavelet transform provides a greater concentration of stripe information into fewer wavelet coefficients, leading to more effective destriping. Experimental results on both synthetically striped imagery as well as real striped imagery from an actual hyperspectral sensor demonstrate superior image quality for the proposed method as compared with other state-of-the-art methods. Na Liu 0014, Wei Li 0032, Ran Tao 0003, James E. Fowler |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | ORSIm Detector: A Novel Object Detection Framework in Optical Remote Sensing Imagery Using Spatial-Frequency Channel FeaturesabstractWith the rapid development of spaceborne imaging techniques, object detection in optical remote sensing imagery has drawn much attention in recent decades. While many advanced works have been developed with powerful learning algorithms, the incomplete feature representation still cannot meet the demand for effectively and efficiently handling image deformations, particularly objective scaling and rotation. To this end, we propose a novel object detection framework, called Optical Remote Sensing Imagery detector (ORSIm detector), integrating diverse channel features extraction, feature learning, fast image pyramid matching, and boosting strategy. An ORSIm detector adopts a novel spatial-frequency channel feature (SFCF) by jointly considering the rotation-invariant channel features constructed in the frequency domain and the original spatial channel features (e.g., color channel and gradient magnitude). Subsequently, we refine SFCF using learning-based strategy in order to obtain the high-level or semantically meaningful features. In the test phase, we achieve a fast and coarsely scaled channel computation by mathematically estimating a scaling factor in the image domain. Extensive experimental results conducted on the two different airborne data sets are performed to demonstrate the superiority and effectiveness in comparison with the previous state-of-the-art methods. Xin Wu 0001, Danfeng Hong, Jiaojiao Tian, Jocelyn Chanussot, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2018 | Distributed bias-compensated normalized least-mean squares algorithms with noisy input
Lijuan Jia, Ran Tao 0003, Yue Wang 0001 |
Sci. China Inf. Sci. | 3 |
| 2018 | Single-channel sampling and multi-channel reconstruction AIC via multiple chirp noise sequences
Hao Huan, Ran Tao 0003 |
Sci. China Inf. Sci. | 3 |
| 2018 | Improved RIP-based performance guarantees for multipath matching pursuit
Juan Zhao 0001, Xia Bai, Ran Tao 0003 |
Sci. China Inf. Sci. | 3 |
| 2018 | Adaptive EIV-FIR filtering against coloured output noise by using linear prediction techniqueabstractThe problem of finite impulse response (FIR) filtering in errors‐in‐variables (EIV) system is studied. Due to the input noise, traditional recursive least‐squares (RLS) algorithms are biased in EIV system. Most existing bias‐compensated approaches are proposed in the case that both the input–output noises are white Gaussian random processes. However, taking account of the situation where the output is corrupted by coloured noise, there are rare existing algorithms work well. Two bias‐compensated RLS algorithms with acceptable computational complexity are proposed, which can obtain unbiased real‐time filtering in non‐stationary system when the input noise is white while the output noise is coloured. Under the assumption that the input signal is a coloured process, linear prediction technique is used to estimate the sample of the input signal. Exploiting the statistical properties of the cross‐correlation function between the least‐squares error and the forward/backward prediction error, the input noise variance can be estimated and the bias can be compensated. Simulation results illustrate the good performance of the proposed algorithms. Lijuan Jia, Jian Lou 0005, Ran Tao 0003, Yue Wang 0001 |
IET Signal Process. | 4 |
| 2018 | Optimal design of orders of DFrFTs for sparse representationsabstractThis study proposes an optimal design of the orders of the discrete fractional Fourier transforms (DFrFTs) and construct an overcomplete transform using the DFrFTs with these orders for performing the sparse representations. The design problem is formulated as an optimisation problem with an ‐norm non‐convex objective function. To avoid all the orders of the DFrFTs to be the same, the exclusive OR of two constraints are imposed. The constrained optimisation problem is further reformulated to an optimal frequency sampling problem. A method based on solving the roots of a set of harmonic functions is employed for finding the optimal sampling frequencies. As the designed overcomplete transform can exploit the physical meanings of the signals in terms of representing the signals as the sums of the components in the time–frequency plane, the designed overcomplete transform can be applied to many applications. Xiao-Zhi Zhang, Bingo Wing-Kuen Ling, Ran Tao 0003, Zhijing Yang, Wai Lok Woo, Saeid Sanei, Kok Lay Teo |
IET Signal Process. | 3 |
| 2017 | Fast FOCUSS method based on bi-conjugate gradient and its application to space-time clutter spectrum estimation
Gatai Bai, Ran Tao 0003, Juan Zhao 0001, Xia Bai, Yue Wang 0001 |
Sci. China Inf. Sci. | 2 |
| 2017 | Distributed incremental bias-compensated RLS estimation over multi-agent networks
Jian Lou 0005, Lijuan Jia, Ran Tao 0003, Yue Wang 0001 |
Sci. China Inf. Sci. | 3 |
| 2017 | Complementary peak reducing signals for TDCS PAPR reductionabstractTransform domain communication systems (TDCSs) are cognitive anti‐interference multi‐carrier communication systems with dynamic spectrum access. The inherent high peak‐to‐average power ratio (PAPR) of TDCS reduces the efficiency of the power amplifier. Adaptive waveform generation of the TDCS also causes the PAPR to vary according to the spectral conditions on hand. In this study, a complementary peak reducing signal (CPRS) method is proposed and analysed. It uses all unoccupied frequency bins to transmit data and uses all interfered frequency bins to generate CPRSs. Every data signal and its corresponding CPRS are orthogonal and complementary, so as to fully occupy all frequency bins. Therefore, the PAPR reduction of the composite signal with all frequency bins is considered. Once the optimal pseudo‐random phase sequence is determined, the sequence can adapt to all spectral conditions without side information, and the computational complexity for diverse spectral conditions is greatly reduced. Moreover, the orthogonality between the data signal and its CPRS in the frequency domain eliminates distortions and spectral spreading. As a component of the transmitting signal, CPRS may cause bit error rate (BER) loss. This study also proposes a signal power adjustment mechanism to achieve a compromise between PAPR reduction and BER loss. Hao Huan, Jianmin Guo, Ran Tao 0003 |
IET Commun. | 4 |
| 2017 | Long-term integration based on two-stage differential acquisition for weak direct sequence spread spectrum signalabstractFast acquisition in space telemetry and telecontrol communication system is challenged significantly by the high dynamic and low signal‐to‐noise ratio. Increasing the dwell time to obtain better acquisition sensitivity, based on two‐dimensional time‐frequency search, is impossible because of resource limitation in receiver. To tackle these problems, this study proposes a two‐stage differential acquisition (TSDA) algorithm. The differential coherent integration is employed by the first stage of TDSA to extend integration time, during which keystone transform and fast Fourier transform are used to compensate the linear code phase drift and obtain high accumulation gain. The carrier‐frequency offset is estimated in terms of the code phase difference between the two parts of received signal divided by the second stage of TSDA in which the two‐dimensional searching process is avoided. The proposed TSDA algorithm and the traditional acquisition methods are compared under the same acquisition time and typical parameters of S‐band space telemetry and telecontrol communication system. Simulation results demonstrate that the proposed TSDA algorithm attains a higher acquisition probability and better acquisition sensitivity than the traditional acquisition methods. Yichao Guo, Hao Huan, Ran Tao 0003, Yue Wang 0001 |
IET Commun. | 3 |
| 2017 | Anti-eavesdropping FrFT-OFDM system exploiting multipath channel characteristicsabstractThis study exploits wireless channel response and fractional Fourier transform (FrFT) to achieve anti‐eavesdropping downlink transmission over a multipath channel. The proposed system works in time‐division duplexing mode, and only the transmitter knows the main channel response through uplink services. With the use of the main channel information, a time‐varying spread‐spectrum orthogonal frequency division multiplexing (TS‐OFDM) subsystem is proposed to realise confidential transform order (TO) transmission, and a security‐enhanced FrFT‐OFDM (SE‐FrFT‐OFDM) subsystem is presented for high‐rate data transmission. The optimal power allocation algorithms that optimise the average secrecy outage capacity of the TS‐OFDM subsystem and the bit error rate performance of the SE‐FrFT‐OFDM subsystem are also provided. Given channel reciprocity and spatial decorrelation, the legitimate user can achieve optimal anti‐multipath fading performance, but the intercepted data symbols will be randomised by both channel fading and TO deviations. To avoid randomisation, the eavesdropper has to test all possible TO combinations of multiple SE‐FrFT‐OFDM symbols, which is usually an impractical task. Compared to other typical methods, the transmitting power of the proposed system can be used more efficiently because the data symbol is not transmitted along with additional interference. Computer simulations are carried out to validate this scheme. Teng Wang 0005, Hao Huan, Ran Tao 0003, Yue Wang 0001 |
IET Commun. | 3 |
| 2017 | Parameter-searched OMP method for eliminating basis mismatch in space-time spectrum estimation
Gatai Bai, Ran Tao 0003, Juan Zhao 0001, Xia Bai |
Signal Process. | 2 |
| 2017 | Signal reconstruction from recurrent samples in fractional Fourier domain and its application in multichannel SAR
Na Liu 0014, Ran Tao 0003, Robert Wang 0001, Yunkai Deng, Ning Li 0002, Shuo Zhao 0002 |
Signal Process. | 2 |
| 2017 | Multichannel Consistent Sampling and Reconstruction Associated With Linear Canonical TransformabstractMultichannel sampling is fundamental in the theory of multichannel parallel analog-to-digital converters and multiplexing wireless communication. This letter investigates the multichannel consistent sampling theory associated with the linear canonical transform, which can combine the reconstruction of the signal and the correction of sensor distortions. The consistency requests that a resampling of the reconstruction by the original sampling operation yields the same measurements as before. The consistent sampling is applicable to an arbitrary finite energy signal and without the constraint such as band-limiting. Under this circumstance, two effective reconstruction methods based on the shift-invariant reconstruction space and the linear canonical basis functions space are developed. Finally, the reconstruction experiment of the radar echo signal is presented. Liyun Xu, Ran Tao 0003, Feng Zhang 0011 |
IEEE Signal Process. Lett. | 2 |
| 2017 | Motion-State-Adaptive Video Summarization via Spatiotemporal AnalysisabstractWith the explosive growth of video data, managing and browsing videos in a timely and effective manner has become an urgent problem, particularly in surveillance applications. Video summarization as a feasible solution is considerably attracting more attention. In this paper, we propose a novel motion-state-adaptive video summarization method based on spatiotemporal analysis. To overcome the low efficiency of traditional video summarization, the proposed method utilizes spatiotemporal slices to analyze object motion trajectories and selects motion state changes as a metric to summarize videos. Initially, a motion-active segment is detected using motion power. Subsequently, motion state changes are modeled as a collinear segment on a spatiotemporal slice (STS-CS) and an attention curve based on the STS-CS model is formed to extract the key frames. Finally, a visually distinguishing mechanism is employed to refine the key frames. The experimental results demonstrate that the proposed method outperforms the existing state-of-the-art methods in terms of both computational efficiency and detailed video motion dynamic maintenance. This is accomplished with a comparable subjective performance. Yunzuo Zhang, Ran Tao 0003, Yue Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Automatic human fall detection in fractional fourier domain for assisted livingabstractFast and accurate detection of elderly falls can significantly reduce the rate of morbidity and mortality. In the past decade, extensive research has been performed to achieve real-time fall monitoring solutions. In this paper, we consider the radar-based modality and utilize the family of fractional Fourier transform to enhance the motion Doppler signature of falls. Compare with the conventional time-frequency analysis approaches, the proposed method achieves higher signal energy concentration and thus yields improved fall detection in low signal-to-noise ratio scenarios. Experimental results are used to validate the theoretical analysis and to demonstrate the feasibility of the proposed approach. Shengheng Liu, Zhengxin Zeng, Yimin Zhang 0001, Tao Shan, Ran Tao 0003 |
ICASSP | 6 |
| 2016 | Bi-conjugate gradient based computation of weight vector in space-time adaptive processing
Gatai Bai, Ran Tao 0003, Yue Wang 0001 |
Sci. China Inf. Sci. | 2 |
| 2016 | Generalized spatial representation for digital modulation and its potential application
Hao Huan, Ran Tao 0003 |
Sci. China Inf. Sci. | 4 |
| 2016 | Reconstruction of uniformly sampled signals from non-uniform short samples in fractional Fourier domainabstractSignal reconstruction from non‐uniform samples, especially for non‐stationary signals, is an important issue in the area of digital signal processing. As a type of signal processing tool, the fractional Fourier transform has been proved to be effective for solving problems in non‐stationary signal processing. For non‐stationary discrete‐time signals, the reconstruction of uniformly sampled signals from non‐uniform samples in the fractional Fourier domain is first derived in this study. Since only finite non‐uniform samples are collected in practical applications, for preferable reconstruction, two types of symmetric extensions are considered in the reconstruction to overcome the discontinuity problem that exists in the periodisation of short non‐stationary sequences, which is more critical than that of long sequences. In addition, the average signal‐to‐noise ratio is used to evaluate the performance of the reconstruction with two types of symmetric extensions. Simulations and two applications are given to verify the effectiveness of the proposed reconstruction method. Feng Zhang 0011, Liyun Xu, Ran Tao 0003, Yue Wang 0001 |
IET Signal Process. | 4 |
| 2016 | Application of linear canonical transform correlation for detection of linear frequency modulated signalsabstractLinear canonical transform (LCT), which can be deemed to be a generalisation of the fractional Fourier transform, has been used in several areas, including signal processing and optics. Motivated by the operator theory, a new unitary operator associated with the LCT is introduced. This new operator generalises the unitary fractional operator which is proposed by Akay et al . recently. Via operator manipulations, the authors also derive a new definition, the LCT correlation operation, and present an alternative and efficient implementation of it. It is shown that the proposed LCT autocorrelation corresponds to radial slices of the ambiguity function in the ambiguity plane. On the basis of this relationship, an application of the fast LCT autocorrelation for detection and parameter estimation with respect to the chirp rates of linear frequency modulated signals corrupted by noise is proposed. Finally, the validity of the proposed method is verified by simulation results. Yuanchao Li, Feng Zhang 0011, Ran Tao 0003 |
IET Signal Process. | 4 |
| 2016 | Rotation Feature Extraction for Moving Targets Based on Temporal Differencing and Image Edge DetectionabstractA rotation parameter extraction method based on temporal differencing and image edge detection from range-Doppler images is presented in this letter. The proposed method first detects the motion trail of the moving pixels caused by the rotating parts in temporal differential range-Doppler images using image edge detection. A Doppler-slow-time image is then generated from the edge pixels on the motion trail. Finally, the rotation parameters are extracted from the Doppler-slow-time image. The proposed method is simple, rapid, and practical. Computer simulations and experimental results demonstrate its effectiveness in terms of computation time compared with existing methods. Zhangfeng Li, Houjun Sun, Ran Tao 0003, Xiaojing Huang 0001, Y. Jay Guo |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2016 | A Novel Two-Dimensional Sparse-Weight NLMS Filtering Scheme for Passive Bistatic RadarabstractIn passive bistatic radars, weak target echoes may often be masked by direct path interference, multipath components, and strong target echoes, making weak target detection a challenging problem. The conventional 1-D adaptive cancelation algorithms, such as the normalized least mean square (NLMS), cannot effectively suppress strong target echoes when their Doppler frequencies spread. In addition, the continuous distribution of the NLMS weight vector does not match the sparse characteristics of strong multipath components and target echoes, thus resulting in degraded cancelation performance. Motivated by this fact, a novel 2-D sparse-weight NLMS filtering scheme is proposed by extending the NLMS to a 2-D structure, in which the weight vector is sparsely distributed and adaptively adjusted based on the sparse strong multipath components and target echoes. Yahui Ma, Tao Shan, Yimin Zhang 0001, Moeness G. Amin, Ran Tao 0003 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2016 | Regularized smoothed ℓ0 norm algorithm and its application to CS-based radar imaging
Hongxia Bu, Ran Tao 0003, Xia Bai, Juan Zhao 0001 |
Signal Process. | 2 |
| 2016 | Randomized nonuniform sampling and reconstruction in fractional Fourier domainabstractThe fractional Fourier transform (FRFT) is one of the most useful tools for the nonstationary signal processing. In this paper, the randomized nonuniform sampling and approximate reconstruction of the nonstationary random signals in the fractional Fourier domain (FRFD) are developed. The nonuniform samples are treated as random perturbations from a uniform grid. The samples used for the sinc interpolation reconstruction are placed on another nonuniform grid which is not necessarily equal to the samples originally acquired. When considering the second-order random statistic characters, the nonuniform sampling is equivalent to the uniform sampling of the signal after a pre-filter in the FRFD, where the frequency response is related to the characteristic function (with its argument scaled by csc α ) of the perturbations. The effectiveness of the reconstruction is analyzed and the mean square error (MSE) is computed by utilizing the equivalent filter system. Furthermore, the randomized reconstruction of the chirp period stationary random signal is proposed. At last, the minimum MSE on the special cases of the randomized sampling and reconstruction is discussed. The effectiveness of the proposed reconstruction method is verified by the simulation. Liyun Xu, Feng Zhang 0011, Ran Tao 0003 |
Signal Process. | 3 |
| 2015 | Chirp Noise Waveform Aided Fast Acquisition Approach for Large Doppler Shifted TT&C SystemabstractIn the next generation Ka-band telemetry, track, and command (TT&C) system, the received TT&C signal will suffer from more than 800 kHz of Doppler shift. Even with code phase-Doppler parallel search algorithm based on FFT, the mean acquisition time (MAT) will be more than several seconds. By exploiting the Doppler tolerance of chirp noise waveform (CNW), the authors propose a novel CNW aided spread spectrum signal fast acquisition algorithm for large Doppler shifted TT&C system. Through this article, the authors first describe the ambiguity function (AF) of CNW, and then design a simplified quadrature modulation model, in which the data symbol and the CNW are transmitted simultaneously. At the receiving end, a fast acquisition method based on dual match filters (MFs) is proposed. Due to the delay-Doppler coupling feature, the receiver can achieve both code synchronization and Doppler shift estimation through only 1-Dimensional (1D) delay search, and then MAT can be reduced to a few milliseconds. Simulation results demonstrate the effectiveness of the proposed method. Teng Wang 0005, Hao Huan, Chuying Feng, Ran Tao 0003 |
GLOBECOM | 4 |
| 2015 | Waveform design for higher-level 3D constellation mappings and its construction based on regular tetrahedron cells
Hao Huan, Ran Tao 0003 |
Sci. China Inf. Sci. | 3 |
| 2015 | Coherence-based analysis of modified orthogonal matching pursuit using sensing dictionaryabstractCompressed sensing (CS) has attracted considerable attention in signal processing because of its advantage of recovering sparse signals with lower sampling rates than the Nyquist rates. Greedy pursuit algorithms such as orthogonal matching pursuit (OMP) are well‐known recovery algorithms in CS. In this study, the authors study a modified OMP proposed by Schnass et al ., which uses a special sensing dictionary to identify the support of a sparse signal while maintaining the same computational complexity. The performance guarantee of this modified OMP in recovering the support of a sparse signal is analysed in the framework of mutual (cross) coherence. Furthermore, they discuss the modified OMP in the case of bounded noise and Gaussian noise, and show that the performance of the modified OMP in the presence of noise relies on the mutual (cross) coherence and the minimum magnitude of the non‐zero elements of the sparse signal. Finally, simulations are constructed to demonstrate the performance of the modified OMP. Juan Zhao 0001, Xia Bai, Shi-He Bi, Ran Tao 0003 |
IET Signal Process. | 4 |
| 2015 | A Novel SAR Imaging Algorithm Based on Compressed SensingabstractTo reduce the amount of measurements, compressed sensing (CS) has been introduced to synthetic aperture radar (SAR). In this letter, a novel CS-SAR imaging algorithm is proposed, which consists of 2-D undersampling, range reconstruction, range-azimuth decoupling, and azimuth reconstruction. In the proposed algorithm, the range profile is reconstructed in the fractional Fourier domain, and range-azimuth decoupling in the case of azimuth undersampling is realized by using the reference function multiplication and chirp-z transform. Comparisons with the existing 2-D undersampling CS-SAR imaging algorithms are also presented. Experimental results from both simulated and real data demonstrate that the proposed algorithm can efficiently realize high-quality imaging with limited measurements. Hongxia Bu, Ran Tao 0003, Xia Bai, Juan Zhao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2015 | Multichannel Random Discrete Fractional Fourier TransformabstractWe propose a multichannel random discrete fractional Fourier transform (MRFrFT) with random weighting coefficients and partial transform kernel functions. First, the weighting coefficients of each channel are randomized. Then, the kernel functions, selected based on a choice scheme, are randomized using a group of random phase-only masks (RPOMs). The proposed MRFrFT can be carried out both electronically and optically, and its main features and properties have been given. Numerical simulation about one-dimensional signal demonstrates that the MRFrFT has an important feature that the magnitude and phase of its output are both random. Moreover, the MRFrFT of two-dimensional image can be viewed as a security enhanced image encryption scheme due to the large key space and the sensitivity to the private keys. Xuejing Kang, Feng Zhang 0011, Ran Tao 0003 |
IEEE Signal Process. Lett. | 3 |
| 2014 | ISAR Imaging of a Ship Target Based on Parameter Estimation of Multicomponent Quadratic Frequency-Modulated SignalsabstractHigh-resolution inverse synthetic aperture radar (ISAR) imaging of a ship target is a challenging task because of fluctuation with the ocean waves. The images obtained with a standard range-Doppler algorithm are usually blurred. Consequently, the range-instantaneous-Doppler (RID) technique should be used to improve the image quality. In this paper, the received signal in a range cell is modeled as a multicomponent quadratic frequency-modulated (QFM) signal after range compression and motion compensation, and then a new RID ISAR imaging algorithm is proposed that introduces a new method for estimating parameters of the QFM signal. By defining a new function and using the scaled Fourier transform (SCFT) with respect to the time axis, the coherent integration of auto-terms can be realized via the subsequent Fourier transformation with respect to the lag-time axis, and a peak can be obtained in the 2-D frequency plane, which is appropriate for parameter estimation of the QFM signal to reconstruct RID images. The proposed algorithm is accurate and fast since the defined function has moderate order nonlinearity and the SCFT can be performed via chirp z-transform. Experiments demonstrate the performance of the new algorithm. Comparisons with existing algorithms are also given, which show that the proposed algorithm can efficiently produce a focused image with less fake scatterers. Xia Bai, Ran Tao 0003, Zhijiao Wang, Yue Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2013 | Multi-channel filter banks associated with linear canonical transform
Juan Zhao 0001, Ran Tao 0003, Yue Wang 0001 |
Signal Process. | 2 |
| 2012 | Compressed sensing SAR imaging based on sparse representation in fractional Fourier domain
Hongxia Bu, Xia Bai, Ran Tao 0003 |
Sci. China Inf. Sci. | 3 |
| 2012 | Modeling and characteristic analysis of underwater acoustic signal of the accelerating propeller
Ran Tao 0003, Yue Wang 0001 |
Sci. China Inf. Sci. | 2 |
| 2012 | Fractional Fourier domain equalization for single carrier broadband wireless systems
Kewu Huang, Ran Tao 0003, Yue Wang 0001 |
Sci. China Inf. Sci. | 2 |
| 2012 | Research on resolution between multi-component LFM signals in the fractional Fourier domain
Feng Liu 0028, Huifa Xu, Ran Tao 0003, Yue Wang 0001 |
Sci. China Inf. Sci. | 3 |
| 2012 | Ambiguity function based on the linear canonical transformabstractThe ambiguity function (AF) as an important tool for time–frequency analysis, has widely been used in radar signal processing, sonar technology and so on. It does well in analysing chirp signals. However, it fails in estimating the cubic phase (CP) signal, which is required in many applications. As the generalisation of the Fourier transform (FT), the linear canonical transform (LCT) has received much attention, and has found many applications in filter design, pattern recognition, optics and so on. In this study, the authors define a new kind of AF – the AF based on the LCT (LCTAF), which is proposed to estimate the CP signal. Some important properties of LCTAF are discussed, such as symmetry and conjugation property, shifting property and Moyal formula. The relationships between the LCTAF and other time–frequency analysis distributions are derived, including the classical AF, the Wigner distribution function (WDF) based on LCT, the short-FT and the wavelet transform. The linear canonical AF (LCAF) is another kind of AF. The authors also discuss its relation with AF and WDF and give some new properties of the LCAF. At last, the LCTAF is applied for estimating the CP signal. The simulation indicates that the LCTAF is useful and effective. Ran Tao 0003, Yu-E. Song, Zhijiao Wang, Yue Wang 0001 |
IET Signal Process. | 1 |
| 2012 | Artifact-Free Despeckling of SAR Images Using ContourletabstractContourlets have been attracting increasing attention in despeckling of synthetic aperture radar (SAR) images in recent years. However, contourlets produce undesirable artifacts while despeckling. This letter presents a novel contourlet-based regularization method to remove speckle without introducing extra artifacts. In this method, a nonlocal regression function is constructed for the regularization term through patch weighting on the contourlet reconstruction. Moreover, by employing the intrinsic dependence of contourlet coefficients, a parameterized shrinkage algorithm is proposed to resolve the contourlet reconstruction. This contourlet despeckling method enables significant elimination of speckle and resultant artifacts with well-protected strong scatters and structures of the scene. Experiments demonstrate such superiority on real SAR images. Ran Tao 0003, Yue Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2012 | Fractional Weierstrass Model for Rough Ocean Surface and Analytical Derivation of Its Scattered Field in a Closed FormabstractAnalytical derivation of scattered field from rough ocean surface can be accomplished by a two-step analysis involving ocean surface modeling and scattered-field calculation. In the first step, the simplest linear model is constructed by assuming a rough ocean state as composition of infinite number of linear waves with identical crests and troughs. Unfortunately, regardless of ocean circumstances, ocean surface waves always exceed the linear range and yield a typical nonlinear asymmetric waveform with sharper peaks and shallow troughs. In this paper, by using curve-fit method and the concept of fractal geometry, we propose a new fractal function named fractional Weierstrass function which may be viewed as a generalization of classical Weierstrass function, and then, a fractional Weierstrass model for ocean surface is proposed by combining the fractional Weierstrass function with the Pierson-Moskowitz spectrum. An important advantage of our proposed model is the ability to represent nonlinearity with different degrees of ocean surface depending on a new parameterc. Some simulations demonstrate ratiocas an invariant scale of fractal ocean surface. The second step is strongly related to the surface modeling result of the first step. Within Kirchhoff approximation, an analytical derivation of scattered field in a closed form is obtained from our proposed fractional Weierstrass model surface. Moreover, some related discussions, including the normalization constant of the model, the correlation length, the scattered-field structure, and the influence of fractal and electromagnetic parameters over scattered field, are given. Finally, edge effects caused by the truncation of scatterer surface are also analyzed in detail. Ran Tao 0003, Xia Bai |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2011 | Time-frequency filtering-based autofocus
Ran Tao 0003 |
Signal Process. | 1 |
| 2011 | Sampling random signals in a fractional Fourier domain
Ran Tao 0003, Feng Zhang 0011, Yue Wang 0001 |
Signal Process. | 1 |
| 2010 | The discrete multiple-parameter fractional Fourier transform
Ran Tao 0003, Yue Wang 0001 |
Sci. China Inf. Sci. | 2 |
| 2010 | Research on distributed intrusion detection system based on multi-living agent
Yue Wang 0001, Ran Tao 0003 |
Sci. China Inf. Sci. | 2 |
| 2010 | On signal moments and uncertainty relations associated with linear canonical transform
Juan Zhao 0001, Ran Tao 0003, Yue Wang 0001 |
Signal Process. | 2 |
| 2010 | Comments on "A Convolution and Product Theorem for the Linear Canonical Transform"abstractA recent letter proposed a new convolution structure for the linear canonical transform, and claimed that theirs was clearly easier to implement in the designing of filters than the one suggested earlier in. However, we find that the two kinds of filtering methods are essentially the same through the theoretic deduction. That is to say, the two kinds of filtering methods can obtain the same effect through the same filtering steps. Taking digital signal processing into account, we analyze further the computation complexity according to the steps of multiplicative filter. Bing Deng, Ran Tao 0003, Yue Wang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2010 | Image Encryption With Multiorders of Fractional Fourier TransformsabstractThe original information in the existing security system based on the fractional Fourier transform (FRFT) is essentially protected by only a certain order of FRFT. In this paper, we propose a novel method to encrypt an image by multiorders of FRFT. In the image encryption, the encrypted image is obtained by the summation of different orders inverse discrete FRFT of the interpolated subimages. And the original image can be perfectly recovered using the linear system constructed by the fractional Fourier domain analysis of the interpolation. The proposed method can be applied to the double or more image encryptions. Applying the transform orders of the utilized FRFT as secret keys, the proposed method is with a larger key space than the existing security systems based on the FRFT. Additionally, the encryption scheme can be realized by the fast-Fourier-transform-based algorithm and the computation burden shows a linear increase with the extension of the key space. It is verified by the experimental results that the image decryption is highly sensitive to the deviations in the transform orders. Ran Tao 0003, Xiangyi Meng, Yue Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2009 | Analysis and side peaks identification of Chinese DTTB signal ambiguity functions for passive radar
ZhiWen Gao, Ran Tao 0003, Yue Wang 0001 |
Sci. China Ser. F Inf. Sci. | 2 |
| 2009 | Realizations of BELS as WIV method in both direct and indirect closed-loop system identification
Lijuan Jia, Ran Tao 0003, Yue Wang 0001, Kiyoshi Wada |
Sci. China Ser. F Inf. Sci. | 2 |
| 2009 | Forward/backward prediction solution for adaptive noisy FIR filtering
Lijuan Jia, Ran Tao 0003, Yue Wang 0001, Kiyoshi Wada |
Sci. China Ser. F Inf. Sci. | 2 |
| 2009 | Using the multi-living agent concept to investigate complex information systems
Yue Wang 0001, Ran Tao 0003, Bing-Zhao Li 0001 |
Sci. China Ser. F Inf. Sci. | 2 |
| 2009 | Oversampling analysis in fractional Fourier domain
Feng Zhang 0011, Ran Tao 0003, Yue Wang 0001 |
Sci. China Ser. F Inf. Sci. | 2 |
| 2009 | The Poisson sum formulae associated with the fractional Fourier transform
Bing-Zhao Li 0001, Ran Tao 0003, Tian-Zhou Xu, Yue Wang 0001 |
Signal Process. | 2 |
| 2009 | New Results on H∞ Filtering for Fuzzy Time-Delay SystemsabstractThis paper is concerned with theHinfinfuzzy filtering problem for a class of continuous-time fuzzy systems with time-varying delays. The objective is to design a stable filter guaranteeing the asymptotic stability and a prescribedHinfinperformance of the filtering error system. Motivated by the parallel distributed compensation technique, we design a new filter model in this paper. Filter parameter matrices can be obtained from the solution of a convex optimization problem in terms of linear matrix inequalities (LMIs). When these LMIs are feasible, an explicit expression of a desiredHinfinfuzzy filter is given. Two numerical examples are provided to demonstrate the effectiveness and less conservativeness of the proposed design approach. Jinhui Zhang 0003, Yuanqing Xia, Ran Tao 0003 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2008 | The multiple-parameter fractional Fourier transform
Ran Tao 0003, Qi-Wen Ran, Yue Wang 0001 |
Sci. China Ser. F Inf. Sci. | 2 |
| 2008 | Research progress on discretization of fractional Fourier transform
Ran Tao 0003, Feng Zhang 0011, Yue Wang 0001 |
Sci. China Ser. F Inf. Sci. | 1 |
| 2008 | MIMO-OFDM system based on fractional Fourier transform and selecting algorithm for optimal order
Ran Tao 0003, Yue Wang 0001, Enqing Chen |
Sci. China Ser. F Inf. Sci. | 2 |
| 2008 | Sampling rate conversion for linear canonical transform
Juan Zhao 0001, Ran Tao 0003, Yue Wang 0001 |
Signal Process. | 2 |
| 2008 | Generalization of the Fractional Hilbert TransformabstractIn this letter, we generalize the fractional Hilbert transform of a real signal to get an analytic version which contains no negative spectrum while maintaining the essential information of the real signal. We also present a secure single-sideband (SSB) modulation system in which the angle of the fractional Fourier transform and the phase of the fractional Hilbert transform are used as double keys for demodulation. Ran Tao 0003, Yue Wang 0001 |
IEEE Signal Process. Lett. | 1 |
| 2007 | Fractional Fourier domain analysis of decimation and interpolation
Xiangyi Meng, Ran Tao 0003, Yue Wang 0001 |
Sci. China Ser. F Inf. Sci. | 2 |
| 2007 | New sampling formulae related to linear canonical transform
Bing-Zhao Li 0001, Ran Tao 0003, Yue Wang 0001 |
Signal Process. | 2 |
| 2006 | Convolution theorems for the linear canonical transform and their applications
Bing Deng, Ran Tao 0003, Yue Wang 0001 |
Sci. China Ser. F Inf. Sci. | 2 |
| 2006 | Research progress of the fractional Fourier transform in signal processing
Ran Tao 0003, Bing Deng, Yue Wang 0001 |
Sci. China Ser. F Inf. Sci. | 1 |
| 2004 | Nonlinear Dynamic Method to Suppress Reverberation Based on RBF Neural Networks
Bing Deng, Ran Tao 0003 |
ISNN (2) | 2 |
| 2004 | Detection and parameter estimation of multicomponent LFM signal based on the fractional fourier transform
Ran Tao 0003, Siyong Zhou, Yue Wang 0001 |
Sci. China Ser. F Inf. Sci. | 2 |
| 2003 | Detection and imaging of slowly moving target of airborne SAR based on the GMCWD-Hough transformabstractIn this paper, the features of airborne SAR moving target echoes are analysed, the Generalized-marginal Choi-Williams Distribution-Hough transform (GMCWD-HT) is also introduced. According to the echoes model of airborne SAR, a new method of detecting and imaging the slowly moving target of airborne SAR based on the (GMCWD-HT) is proposed. Computer simulation results show that this method can be used to perform the slowly moving target detection and imaging of airborne SAR in the low signal to clutter ratio; its detecting performance is better than the common method based on Wigner-Ville distribution. Ran Tao 0003, Siyong Zhou |
ICASSP (6) | 2 |
| 2002 | Efficient video coding with hybrid spatial and fine-grain SNR scalabilities
Feng Wu 0001, Shipeng Li 0001, Ran Tao 0003, Yue Wang 0001 |
VCIP | 4 |
| 2002 | An improved self-organizing CPN-based fuzzy system with adaptive back-propagation algorithm
Yue Wang 0001, Ran Tao 0003, Siyong Zhou |
Fuzzy Sets Syst. | 3 |