Xiaobin Zhao

dblp:156/4583 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A Cooperative Meta-Learning and Spectral Diversity Adaptation Framework for Hyperspectral Target Detection
abstract
Meta-learning has demonstrated significant potential in addressing the limited annotation data challenge in hyperspectral target detection. However, existing meta-learningbased methods face two major challenges: 1) weak inter-task correlation leading to unstable optimization directions, and 2) meta-knowledge adaptation based on single prior spectral information fails to effectively characterize spectral variation properties of targets in new scenarios. To overcome these challenges, this letter proposes a cooperative meta-learning framework with spectral diversity adaptation for hyperspectral target detection. The framework introduces a co-learner through cooperative learning to dynamically capture cross-task knowledge and stabilize optimization directions, while designing a spectral diversity-based meta-knowledge adaptation strategy to enhance the model’s ability to understand the spectral variation characteristics of targets in new scenarios and precisely distinguish the spectral features between targets and backgrounds. Experimental results on two public datasets demonstrate that the proposed method outperforms state-of-the-art hyperspectral target detection algorithms. The code is available at https://github.com/Li-ZK/CMLSDA.
Yan Wang 0087, Bo Yuan 0013, Zhaokui Li, Xiaobin Zhao, Jiaxu Guo
IEEE Geosci. Remote. Sens. Lett.4
2025 Hyperspectral Target Detection Using Diffusion Model and Convolutional Gated Linear Unit
abstract
Deep learning can effectively extract latent information from data to enhance target-background separation in hyperspectral target detection (HTD). However, these models typically require extensive labeled samples, while available target spectra in hyperspectral images (HSI) are scarce. Additionally, existing deep models struggle with target detection in complex backgrounds due to subtle spectral differences. To address these issues, we propose a novel HTD method based on diffusion model and convolutional gated linear unit (HTD-DMCG). First, the diffusion model is integrated with MixUp for data augmentation to generate a diverse and sufficiently large sample set. Next, a Transformer architecture utilizing a convolutional gated linear unit is designed to effectively capture global dependencies and local feature correlations, leading to more discriminative feature representations. Additionally, a new target aggregation and background separation loss is introduced, which emphasizes target sample aggregation while increasing the distance between targets and background samples to enhance separability. The HTD-DMCG method is compared against classical and state-of-the-art HTD methods on four real HSI datasets. Extensive experiments show that it can effectively outperform existing methods in target detection performance. The code is available at https://github.com/Li-ZK/HTD-DMCG.
Zhaokui Li, Xiaobin Zhao, Cuiwei Liu, Xuewei Gong, Wei Li 0032, Qian Du 0001, Bo Yuan 0013
IEEE Trans. Geosci. Remote. Sens.3
2025 Efficient Mamba-Attention Network for Remote Sensing Image Super-Resolution
abstract
Lightweight remote sensing image super-resolution (RSISR) methods aim to reconstruct remote sensing images (RSIs) while reducing computational complexity. Previous lightweight model development has primarily focused on the design of convolutional neural networks (CNNs). While CNNs excel at capturing local features, they are limited in establishing long-range dependencies. Mamba, as a model for long-range modeling, has linear computational complexity, making it a viable option for lightweight models. Based on these considerations, this paper proposes an efficient mamba-attention network (EMAN) that can efficiently capture the intricate details and broader semantic information in RSIs. Specifically, we designed a multi-scale detail extraction unit (MDEU) and a multi-dimensional mamba-attention (MDMA). In MDEU, we introduced a multi-scale mechanism and local variance to focus on structural information in RSIs. In MDMA, we integrated spatial expansion and an atrous-based selective scan mechanism to design an efficient scanning method. This method ensures the lightweight nature of the model while establishing global correlations. Additionally, MDMA establishes inter-channel correlations to enhance information exchange. We conducted a comprehensive evaluation of the proposed method on two remote sensing datasets and five benchmark super-resolution (SR) datasets. Extensive experiments demonstrate that our method can achieve superior performance while maintaining a model complexity similar to other lightweight models.
Tianren Wu, Rundong Zhao, Ming Lv, Zhenhong Jia, Liangliang Li 0001, Minqin Liu, Xiaobin Zhao, Hongbing Ma, Gemine Vivone
IEEE Trans. Geosci. Remote. Sens.7
2025 PRF-Net: A Progressive Remote Sensing Image Registration and Fusion Network
abstract
Most of the existing fusion algorithms are not robust to unregistered input images. Even after image registration, nonlinear nonregistration may persist in the local areas of the images, leading to poor quality in the fused image. So, as to tackle these challenges, a progressive remote sensing image registration and fusion network is proposed in this article, and named PRF-Net, which is particularly useful when two images are from different platforms. First, a registration network is designed to register the input image patches, which includes a global spatial transform network (GSTN) and a local spatial warp network (LSWN). The GSTN is primarily used for coarse registration, applying rigid transformation to globally align the input images. After coarse registration, the preliminarily registered moving image is input into the LSWN for local fine-tuning to maximize correlation between the input image patches. Subsequently, the fine registered images are degraded and input into the fusion network to generate the fused image. To maintain sufficient spectral and spatial information of the fused image, a multiscale feature extraction (MSFE) block with a highly interpretable spatial details attention (SDA) block is designed, which can enhance the ability of fusion network to extract and preserve spatial details and spectral information. Three groups of experiments conducted on four types of remote sensing images give evidence of that the proposed PRF-Net exhibits excellent performance in both reduced and full resolutions, showcasing its outstanding registration and fusion quality.
Zhangxi Xiong, Wei Li 0032, Xiaobin Zhao, Baochang Zhang 0001, Ran Tao 0003, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Grid-Forming Control Based on Capacitor Voltage Synchronization for Modular Multilevel Converter Under Weak Grid Condition
abstract
Considering the voltage source characteristics, grid-forming (GFM) control is regarded as an effective method to enhance the voltage support capability of inverters, especially under weak grid condition. However, relatively mature grid-forming control methods such as virtual synchronous generator control typically assume a constant voltage source on the dc side, which is not suitable for the modular multilevel converter (MMC) that controls the dc voltage. To address this issue, a GFM control strategy based on the decoupled control is proposed. This strategy achieves grid synchronization based on the dynamic variation of submodule capacitor voltage rather than the dc voltage, by linking the swing equation of the synchronous generator with the energy dynamic equation of the MMC. Additionally, the energy stored in submodule capacitors can be flexibly adjusted to provide inertial responses for the connected grid. With the proposed GFM control, the MMC can provide voltage and reactive power support under weak grid condition. Meanwhile, the dc voltage can be regulated independently through the decoupled control of MMC. The effectiveness of the proposed control method is verified through the simulation results in MATLAB/Simulink.
Qiang Song 0002, Kailun Wang, Xiaobin Zhao, Biyue Huang, Qingming Xin
IECON4
2024 Global Feature-Injected Blind-Spot Network for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) poses the challenge of distinguishing anomalous targets from the majority of background objects without prior knowledge. Most existing deep learning (DL) models struggle to account for both local and global spatial-spectral features in the image, limiting their performance. In this letter, we introduce PUNNet, which integrates the patch-shuffle downsampling technique and nonlinear activation-free network (NAFNet) block with dilated convolution into an advanced blind-spot network for HAD. Specifically, PUNNet utilizes the patch-shuffle downsampling operation to extend its receptive field and exploits channel attention in the NAFNet block with dilated convolution to capture global contextual information in the image. Meanwhile, PUNNet satisfies the blind-spot requirement, meaning its receptive field excludes the center pixel’s information. This allows for reliable and precise background reconstruction in a self-supervised learning paradigm, further weakening anomalous feature expression and increasing the reconstruction error of anomalies. Experimental results demonstrate that PUNNet achieves a leading position in HAD performance. The code is available athttps://github.com/DegangWang97/IEEE_GRSL_PUNNet.
Lina Zhuang, Lianru Gao, Xu Sun 0005, Xiaobin Zhao
IEEE Geosci. Remote. Sens. Lett.5
2024 Sliding Dual-Window-Inspired Reconstruction Network for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) aims to identify anomalous objects that deviate from surrounding backgrounds in an unlabeled hyperspectral image (HSI). Most available neural networks that make use of the reconstruction error to perform HAD tend to fit both backgrounds and anomalies, resulting in small reconstruction errors for both and not being effective in separating targets from background. To address this issue, we develop DirectNet, a new background reconstruction network for HAD that seamlessly integrates a sliding dual-window model into a blind-block architecture. Concretely, DirectNet establishes an inner window within the network’s receptive field by erasing the center block information, so that the content of the inner window remains invisible during the reconstruction of the central pixel. Additionally, the depth of our reconstruction network is adaptive to the size of the input image patch, ensuring that the network’s receptive field aligns with the dimensions of the input patch. The receptive field outside the inner window is considered an outer window. This weakens the impact of anomalies on the reconstruction process, causing the reconstructed pixels to converge towards the background distribution in the outer window region. Consequently, the reconstructed HSI can be regarded as a pure background HSI, leading to further amplification of reconstruction errors for anomalous targets. This enhancement improves the discriminatory ability of DirectNet. Specifically, DirectNet solely utilizes the outer window information to predict/reconstruct the central pixel. As a result, when reconstructing pixels inside anomalous targets of different sizes, the targets primarily fall within the inner window. Comprehensive experiments (conducted on four datasets) demonstrate that DirectNet achieves competitive performance compared to other state-of-the-art detectors.
Lina Zhuang, Lianru Gao, Xu Sun 0005, Xiaobin Zhao, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.5
2024 A Hierarchical Hybrid Learning Framework for Multi-Agent Trajectory Prediction
abstract
Accurate trajectory prediction for neighboring agents is crucial for autonomous vehicles navigating complex scenes. Recent deep learning (DL) methods excel in encoding complex interactions but often generate invalid predictions due to difficulties in modeling transient and contingency interactions. This paper proposes a hierarchical hybrid framework that combines DL and reinforcement learning (RL) for multi-agent trajectory prediction, capturing multi-scale interactions that shape future motion. In the DL stage, Transformer-style graph neural network (GNN) is employed to encode heterogeneous interactions at intermediate and global scales, predicting multi-modal intentions as key future positions for agents. In the RL stage, we divide the scene into local scenes based on DL predictions. A Transformer-based Proximal Policy Optimization (PPO) model, incorporated with vehicle kinematics, generates future trajectories in the form of motion planning shaped by microscopic interactions and guided by a multi-objective reward for balanced agent-centric accuracy and scene-wise compatibility. Experimental results on the Argoverse benchmark and driver-in-loop simulations demonstrate that our framework enhances trajectory prediction feasibility and plausibility in interactive scenes.
Yujun Jiao, Mingze Miao, Zhishuai Yin, Chunyuan Lei, Xiaobin Zhao, Linzhen Nie, Bo Tao 0003
IEEE Trans. Intell. Transp. Syst.6
2023 Hyperspectral Time-Series Target Detection Based on Spectral Perception and Spatial-Temporal Tensor Decomposition
abstract
The detection of camouflaged targets in the complex background is a hot topic of current research. Existing hyperspectral target detection algorithms do not take advantage of spatial information and rarely use temporal information. It is difficult to obtain the required targets, and the detection performance in hyperspectral sequences with complex background will be low. Therefore, a hyperspectral time-series target detection method based on spectral perception and spatial-temporal tensor decomposition (SPSTT) is proposed. Firstly, a sparse target perception strategy based on spectral matching is proposed. To initially acquire the sparse targets, the matching results are adjusted by using the correlation mean of the prior spectrum, the pixel to be measured and the four-neighborhood pixel spectra. The separation of target and background is enhanced by making full use of local spatial structure information through local topology graph representation of the pixel to be measured. Secondly, in order to obtain a more accurate rank and make full use of temporal continuity and spatial correlation, a spatial-temporal tensor model based on the Gamma norm andL2,1norm is constructed. Furthermore, an excellent alternating direction method of multipliers is proposed to solve this model. Finally, spectral matching is fused with spatial-temporal tensor decomposition in order to reduce false alarms and retain more right targets. A 176-band hyperspectral image sequence (BIT-HSIS-I) dataset is collected for the hyperspectral target detection task. It is found by testing on the collected dataset that the proposed SPSTT has superior performance over the state-of-the-art algorithms.
Xiaobin Zhao, Kun Gao 0001, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.1
2022 Hyperspectral Target Detection Based on Weighted Cauchy Distance Graph and Local Adaptive Collaborative Representation
abstract
Hyperspectral target detection in complex backgrounds is a challenging and important research topic in the remote sensing field. Traditional target detectors consider the background spectrum to obey a Gaussian distribution. However, this distribution may not meet the requirements in real hyperspectral images. In addition, the background and spatial information of most existing target detection algorithms are rarely fully utilized. Therefore, a new weighted Cauchy distance graph (WCDG) and local adaptive collaborative representation detection (CGCRD) is proposed. First, a WCDG similarity measure is designed. In order to adjust the effect of target pixels on the graph model, a weighted Cauchy distance Laplace matrix is constructed, and then the matrix is applied to the matched filter detector. Second, local adaptive collaborative representation strategy is developed. The penalty coefficient is weighted by the local spatial Euclidean distance combined with the Pearson correlation coefficient, and then the detection result is obtained based on the residual. Finally, aforementioned two strategies are fused to fully utilize the spatial and spectral information. A 176-band hyperspectral image (BIT-HSI-I) dataset is collected for the target detection task. The related algorithms are performed on the BIT-HSI-I dataset, and the detection results demonstrate that the proposed algorithm has better detection performance than other state-of-the-art algorithms.
Xiaobin Zhao, Wei Li 0032, Chunhui Zhao 0003, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.1
2020 Hyperspectral Target Detection by Fractional Fourier Transform
abstract
Target detection in hyperspectral images (HSI) is an important technique and many target detection algorithms have been developed in recent years. The most widely detection algorithms by the original spectral characteristics may lack the ability of target signal enhancement and background suppression. This paper presents an efficient algorithm for detecting hyperspectral targets based on fractional Fourier transform (FrFT). Firstly, fractional Fourier transform primary search is used as preprocessing to obtain the better intermediate domain features with complementary characteristics between the original reflection spectrum and the Fourier transform domain. Secondly, fractional Fourier transform secondary search and constrained energy minimization (FrFT-CEM) was adopted to find an optimal fractional order to distinguish the target from the background. The proposed method has been proved to be superior in two real hyperspectral data sets.
Xiaobin Zhao, Wei Li 0032, Tao Shan, Lu Li 0005, Ran Tao 0003
IGARSS1