Yice Cao

dblp:225/6789 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
11since 2021 · last 2025
0009-0009-4734-4592ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Few-Shot Radar Active Jamming Recognition Under Label Prior Information Reuse
abstract
In the current fully open and highly dynamic electronic countermeasures environment, acquiring sufficient jamming label priors is particularly challenging. Especially when the labeled samples are extremely limited, existing intelligent recognition models for radar active jamming struggle to effectively learn discriminative features, resulting in inaccurate and unstable recognition results. Therefore, this letter proposes a visual-text alignment network (VTANet) that introduces the text modality to reuse the label prior information, thereby leveraging labeling knowledge to enhance jamming recognition accuracy under few-shot conditions. During training, VTANet utilizes a text feature encoding module (TFEM) to encode text constructed based on label priors. Through a contrastive learning strategy, these obtained text features are then used to guide the visual feature encoding module (VFEM) to learn more discriminative visual time-frequency (TF) representations. Experimental results show that reusing label priors through the text modality significantly improves the intraclass compactness and interclass separability of visual TF features. With the jamming-to-noise ratio (JNR) of 5 dB and only two labeled samples per jamming type, VTANet achieves a recognition accuracy of over 90%, improving by more than 5% compared to existing methods, thus demonstrating its superiority.
Tengxin Wang, Yice Cao, Wenjie Guo, Lixia Yang
IEEE Geosci. Remote. Sens. Lett.3
2025 A Biclassifier Network With Intermediate Domain for Unsupervised Domain Adaptation PolSAR Image Classification
abstract
Due to significant data distribution differences among polarimetric synthetic aperture radar (PolSAR) images and the expensive and time-consuming nature of data labeling, existing methods are challenged to classify newly acquired unlabeled data. To address these issues, this letter proposes an unsupervised domain adaptation (UDA) network that leverages an intermediate domain-assisted biclassifier. An adversarial UDA network incorporating a biclassifier is introduced as the fundamental structure. Then, by interacting with the semantic information of the features extracted from the source and target domains, the cross-domain feature enhancement module (CFEM) is integrated to improve intraclass cohesion and interclass separation. In addition, to achieve a more stable domain alignment process, an intermediate domain is created through the weighted fusion of features from both the source and target domains, serving as a conduit for cross-domain knowledge transfer. Experimental results conducted on three datasets captured with different systems or regions show that the proposed method achieves an average overall accuracy (OA) improvement of over 1.5% compared with other state-of-the-art methods.
Dayi Zhu, Yice Cao, Lixia Yang
IEEE Geosci. Remote. Sens. Lett.3
2024 SARGap: A Full-Link General Decoupling Automatic Pruning Algorithm for Deep Learning-Based SAR Target Detectors
abstract
Synthetic aperture radar (SAR) target detectors based on deep learning have difficulty finding a good balance between accuracy and speed. Current pruning methods are usually used for backbone consistent pruning and seldom directly for the whole structure of deep learning target detectors; therefore, for edge-end applications, this article proposes a new full-link general automatic pruning algorithm for SAR target detectors, referred to as SARGap. First, SARGap automatically analyzes the network structure by creating a dependency graph, divides the pair-coupled network structure into the same group, and prunes the same channel for the same group of network structures so that the algorithm can be applied to a variety of complex target detectors. Second, an automatic pruning rate search method (APRS) is designed to search for the optimal pruning rate of each group of network structures in the target detector. Finally, to find a good balance between precision and speed in the automatic search of the pruning rate, a multiobjective optimization loss function (MOOL) is constructed as the APRS objective function. A series of experiments based on SSDD and HRSID, two large-scale SAR target detection datasets, are carried out to prove the superiority of this method. Using Yolov5s as the baseline, SARGap can compress parameters by 84.29%/82.86% and flops by 80.50%/81.93% on two datasets with almost no loss of accuracy. In addition, SARGap can be applied to any deep learning target detector and match hardware computing resources to achieve optimal full-link pruning.
Jingqian Yu, Jie Chen 0035, Huiyao Wan, Yice Cao, Zhixiang Huang, Yingsong Li 0001, Bocai Wu, Baidong Yao
IEEE Trans. Geosci. Remote. Sens.5
2024 VFL3D: A Single-Stage Fine-Grained Lightweight Point Cloud 3D Object Detection Algorithm Based on Voxels
abstract
In this work, we propose a voxel-based single-stage fine-grained and efficient point cloud 3D object detection algorithm to address the inadequate granularity in point cloud feature extraction tasks and the imbalance between efficiency and accuracy in single-stage point cloud 3D object detection scenarios. We develop a lightweight multibranch cross-sparse convolution network (LMCCN) that is designed to preserve the feature granularity of the original point cloud while achieving enhanced extraction efficiency. Additionally, we introduce a compact fine-grained self-attention augmented bird’s eye view (BEV) feature extraction module (CFSAM). This module aims to further refine BEV features, enabling the acquisition of both locally and globally enhanced features and thereby augmentingthe perceptual capabilities of the constructed model. Without bells and whistles, the proposed method attains excellent performance on many autonomous driving benchmarks, with detection accuracies of up to 81.67% on KITTI, 72.74% on ONCE, and 84.00% on nuScenes. Moreover, it reaches a peak detection speed of 46.08 FPS, effectively balancing accuracy with speed.
Bing Li 0033, Jie Chen 0035, Xinde Li, Yice Cao, Jun Wu 0024, Yingsong Li 0001, Paulo S. R. Diniz
IEEE Trans. Intell. Transp. Syst.6
2023 High-Confidence Sample Augmentation Based on Label-Guided Denoising Diffusion Probabilistic Model for Active Deception Jamming Recognition
abstract
Accurate recognition of the types of mainlobe active deception jamming is essential for radar systems to take anti-jamming countermeasures. A bunch of deep learning (DL)-based recognition methods that require largescale datasets for training have shown promising results. However, capturing a sufficient number of diverse deception jamming samples is particularly intricate in actual dynamic and complex battlefields, the yielding limited or unbalanced datasets presents a significant challenge in training and generalizing DL models. This letter proposes a deep generative model, called label-guided denoising diffusion probabilistic model (LG-DDPM), to address the issue of limited or class-imbalanced active deception jamming samples through data augmentation. By embedding label information into the diffusion process, the proposed model can generate and expand the active deception jamming samples specific to a pre-defined class even under low jamming-to-noise ratios (JNR) scenarios. The proposed method demonstrates superior performance in terms of both the fidelity and diversity of the generated jamming samples as well as the recognition accuracy of the DL recognizer when compared to state-of-the-art methods. Experimental results demonstrate the effectiveness and robustness of the proposed method.
Yice Cao, Tengxin Wang, Lixia Yang
IEEE Geosci. Remote. Sens. Lett.4
2023 Multifrequency PolSAR Image Fusion Classification Based on Semantic Interactive Information and Topological Structure
abstract
Compared with the rapid development of single-frequency polarimetric SAR (PolSAR) image classification technology, there is less research on the land cover classification of multi-frequency PolSAR (MF-PolSAR) images. And the deep learning methods among them are mainly based on convolutional neural networks (CNNs), only local spatiality is considered but the nonlocal relationship is ignored. Therefore, this paper proposes the MF semantics and topology fusion (MF-STF) model based on semantic interaction and nonlocal topological structure to improve MF-PolSAR classification performance. During MF-STF optimization, the semantic information-based classification (SIC) and topological property-based classification (TPC) work collaboratively, not only fully leveraging the complementarity of bands, but also combining local and nonlocal spatial information to improve the discrimination of different categories. For SIC, the designed cross-band interactive feature extraction (CIFE) module is embedded to explicitly model the deep semantic correlation among bands, thereby leveraging the complementarity of bands to make ground objects more separable. In TPC, the graph sample and aggregate network (GraphSAGE) is employed to dynamically capture the representation of nonlocal topological relations between land cover categories. In this way, the robustness of classification can be further improved by combining nonlocal spatial information. Finally, a MF weighted fusion (MFWF) strategy is proposed to merge inference from different bands, so as to make the MF joint classification decisions of SIC and TPC. Notably, its weights are adjusted based on the total model loss. The effectiveness of the proposed modules is proved by ablation experiments on three measured MF-PolSAR datasets. In addition, the comparative experiments show that MF-STF can achieve more competitive classification performance than some state-of-the-art methods.
Yice Cao, Yan Wu 0003, Ming Li 0004, Mingjie Zheng 0001, Peng Zhang 0003, Jili Wang
IEEE Trans. Geosci. Remote. Sens.1
2023 SARNas: A Hardware-Aware SAR Target Detection Algorithm via Multiobjective Neural Architecture Search
abstract
Most of the existing deep learning-based SAR target detection algorithms rely on manual experience to repeatedly adjust structures and parameters to design models suitable for specific scenarios or tasks. The implementation of the above methods is complicated, the design efficiency is low, and it is difficult to ensure the balance between accuracy and complexity. We innovatively propose a hardware-aware SAR target detection algorithm via multiobjective neural architecture search (NAS), referred to as SARNas. First, we design a flexible and efficient search space, a supernet search strategy and a subnet contribution evaluation strategy. Furthermore, we construct a new NAS loss function, called SARMI-Loss, to guide the learning of a SAR object detector that balances accuracy and computational complexity. Our SAR-Nas method can address the resource limitations of edge devices and automatically search for the optimal SAR target detector in an end-to-end manner for any deep learning-based SAR baseline model. A series of comparative experiments on three SAR image object detection datasets (SSDD, HRSID and MSAR) demonstrate the superiority of our method. The experimental results with YOLOV5 as the benchmark model show that the detection accuracy of the target detection networks automatically found by using the SARNas method on the SSDD, HRSID, and MSAR datasets can reach 98.5%, 92.8%, and 91.8% in mean average precision (mAP) with only 2.31M, 1.99M, 2.21M parameters, respectively. The number of model parameters is reduced by 88.9%, 90.46%, and 68.5%, respectively, and the inference speed is increased by 51.6%, 46.1%, and 13.9% without losing accuracy.
Wentian Du, Jie Chen 0035, Chaochen Zhang, Po Zhao, Huiyao Wan, Yice Cao, Zhixiang Huang, Yingsong Li 0001, Bocai Wu
IEEE Trans. Geosci. Remote. Sens.7
2022 Adaptive multiple kernel fusion model using spatial-statistical information for high resolution SAR image classification
Wenkai Liang, Yan Wu 0003, Ming Li 0004, Yice Cao
Neurocomputing4
2022 High Resolution SAR Image Classification Using Global-Local Network Structure Based on Vision Transformer and CNN
abstract
High-resolution (HR) synthetic aperture radar (SAR) image classification is a challenging task for the limitation of its complex semantic scenes and coherent speckles. Convolutional neural networks (CNNs) have been proven the superior local spatial features representation capability for SAR images. However, it is hard to capture global information of images by convolutions. To solve such issues, this letter proposes an end-to-end network named global–local network structure (GLNS) for HR SAR classification. In the GLNS framework, a lightweight CNN and a compact vision transformer (ViT) are designed to learn local and global features, and two types of features are fused in quality to mine complementary information through the fusion net. Then, our research devolves the twofold loss function to reduce the interclass distance of SAR images, which brings more compactness to classification features and less interference of coherent speckles. Experimental results on real HR SAR images indicate that the proposed method has more strong feature extraction capability and noise resistance performance. This method achieves the highest classification accuracy on both datasets compared with other related approaches based on CNN.
Yan Wu 0003, Wenkai Liang, Yice Cao, Ming Li 0004
IEEE Geosci. Remote. Sens. Lett.4
2022 DFAF-Net: A Dual-Frequency PolSAR Image Classification Network Based on Frequency-Aware Attention and Adaptive Feature Fusion
abstract
Multifrequency (MF) polarization synthetic aperture radar (PolSAR) systems can obtain more abundant and continuous earth resource information than single-frequency ones and have been widely used in the remote sensing community. However, there are relatively a few researches for the fine MF PolSAR image classification, which is an important part of remote sensing image interpretation. The main focus currently is on the single-frequency part. Therefore, for dual-frequency PolSAR image classification, this article proposes the dual-frequency attention fusion network (DFAF-Net). It is based on frequency-aware attention and adaptive feature fusion to improve classification performance. First, the dual-frequency PolSAR data is input into the joint feature extraction (JFE) module to obtain the joint feature representation. Meanwhile, two frequency-aware attention block (FAB) modules with the same structure are constructed, which, respectively, generate frequency-specific attention masks based on the guidance of different frequency data. Subsequently, these masks are used to weigh the joint features to highlight the description of different frequency-aware importance. The activated frequency-aware features can fully mine and utilize the complementary information provided by different frequencies, thereby enhancing the discrimination of similar landcover categories. Finally, the adaptive feature fusion block (AFFB) module is utilized to adaptively aggregate different frequency-aware features multiple times, which can effectively eliminate information differences. The obtained fusion features are more compact within classes and separable between classes, thereby effectively improving the classification performance. Experiments on three measured spaceborne and airborne dual-frequency PolSAR datasets verify that DFAF-Net can better perceive frequency characteristics and fully mine the complementary. Therefore, the classification accuracy is effectively enhanced, and the inaccuracy of single-frequency classification can be eliminated. Meanwhile, due to the introduction of the attention module, the classification performance of the proposed DFAF-Net is more competitive than the related deep learning networks. Quantitatively, the overall accuracy of DFAF-Net on the three datasets is respectively 98.30%, 97.42%, and 99.45%, which is better than other methods.
Yice Cao, Yan Wu 0003, Ming Li 0004, Wenkai Liang
IEEE Trans. Geosci. Remote. Sens.1
2022 A Feature Fusion-Net Using Deep Spatial Context Encoder and Nonstationary Joint Statistical Model for High-Resolution SAR Image Classification
abstract
The nonstationary and non-Gaussian distribution of the high-resolution (HR) synthetic aperture radar (SAR) image provides much valuable information. However, the current methods, especially deep learning models, directly learn spatial features from HR SAR data while ignoring global statistical information. Combining the local spatial features and global statistical properties of HR SAR images is urgently needed to capture complete HR SAR characteristics. In this paper, a feature fusion network (Fusion-Net) using both deep spatial context encoder and nonstationary joint statistical model (NS-JSM) is proposed for the first time. Fusion-Net realizes the fusion description of local spatial and global statistical features in an end-to-end supervised classification framework. First, a deep spatial context encoder network (DSCEN) is designed based on multiscale group convolution (MSGC) module and channel attention (CA) module. The DSCEN expands the scope of context information extraction with few parameters and increases the interaction between high- level feature channels. Then, the NS-JSM is adopted to capture the unique SAR statistical information. Specifically, the SAR image is transformed into the Gabor wavelet domain. The produced sub-band magnitudes and phases are modeled by the log-normal and uniform distribution. The covariance matrix (CM) is calculated for mapped sub-band data to capture the interscale and intrascale nonstationary correlation. Finally, the group compression and smooth normalization units are introduced into Fusion-Net to fuse the statistical features and spatial features, which not only exploits the complementary information between different features but also optimizes the fusion feature representation. Experiments on four real HR SAR images validate the superiority of the proposed method over other related algorithms.
Wenkai Liang, Yan Wu 0003, Ming Li 0004, Yice Cao
IEEE Trans. Geosci. Remote. Sens.4
2020 High-Resolution SAR Image Classification Using Context-Aware Encoder Network and Hybrid Conditional Random Field Model
abstract
A pixel-wise classification for high-resolution (HR) synthetic aperture radar (SAR) images is a challenging task, due to the limited availability of labeled SAR data, as well as the difficulty of exploring context information affected by coherent speckle. In this article, we propose a novel supervised classification method for HR SAR images, which combines a context-aware encoder network (CAEN) and a hybrid conditional random field (HCRF) model. First, a new CAEN architecture is developed based on the intrinsic property of HR SAR pixel-wise labeling. The proposed architecture follows an encoder-decoder structure, wherein the residual context encoder (RCE) block and the global context-aware (GCA) block are proposed in the encoder module to capture local to global semantic contexts. The multiscale skip connections and feature compression structures are designed in the decoder module to preserve precise object structures while improving computational efficiency. Then, a patch sampling strategy is adopted to ensure that the training and test data are completely separated. It can achieve a less-biased estimate of the test error of CAEN. In addition, the overlapped sampling and data augmentation techniques are used to solve the problem of limited labeled data. Finally, the HCRF model is constructed and combined with the previous CAEN to further enhance the spatial label consistency. Our HCRF integrates pixel-level and region-level potentials into a unified Bayesian framework, making two spatial supports come to a more accurate decision on pixel categories. Experiments on four HR SAR images validate the superiority of the proposed method over other related algorithms.
Wenkai Liang, Yan Wu 0003, Ming Li 0004, Yice Cao
IEEE Trans. Geosci. Remote. Sens.4
2018 SAR and Optical Image Registration Using Nonlinear Diffusion and Phase Congruency Structural Descriptor
abstract
The registration of synthetic aperture radar (SAR) and optical images is a challenging task due to the potential nonlinear intensity differences between the two images. In this paper, a novel image registration method, which combines nonlinear diffusion and phase congruency structural descriptor (PCSD), is proposed for the registration of SAR and optical images. First, to reduce the influence of speckle noise on feature extraction, a uniform nonlinear diffusion-based Harris (UND-Harris) feature extraction method is designed. The UND-Harris detector is developed based on nonlinear diffusion, feature proportion, and block strategy, and explores many more well-distributed feature points with potential of being correctly matched. Then, according to the property that structural features are less sensitive to modality variation, a novel structural descriptor, namely, the PCSD, is constructed to robustly describe the attributes of the extracted points. The proposed PCSD is built on a PC structural image in a grouping manner, which effectively increases the discriminability and robustness of the final structural descriptor. Experimental results conducted on SAR and optical image pairs demonstrate that the proposed method is more robust against speckle noise and nonlinear intensity differences and improves the registration accuracy effectively.
Jianwei Fan, Yan Wu 0003, Ming Li 0004, Wenkai Liang, Yice Cao
IEEE Trans. Geosci. Remote. Sens.5