Yue Ma 0008

dblp:08/6794-8 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0002-5422-315XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RDAM: Domain adaptation under small and class-imbalanced samples
Youquan Fu, Zhixi Feng, Yue Ma 0008
Knowl. Based Syst.4
2025 Hybrid-View Self-Supervised Framework for Automatic Modulation Recognition
abstract
Applying self-supervised deep learning improves the processing speed and accuracy of automatic modulation recognition (AMR). It reduces the dependence of previous deep networks on many labeled samples. However, affected by an incomplete signal representation modes set, previous models do not fully utilize the multiview property of signals in self-supervised learning. To deal with this issue, a hybrid-view contrastive model for AMR is proposed in this article based on self-supervised learning framework. First, star video is proposed to complete the set of signal representation modes. Next, a self-supervised learning framework based on hybrid-view contrastive learning, hybrid-view self-supervised framework (HVSF), is established to fully extract the signal features, where signals are augmented across views, including the discrete sequence, image, and video format. Considering the view-exclusive information loss and the model complexity, a weakly contrastive strategy and a Transformer-based view-shared feature extractor are finally constructed. Evaluation on four standard datasets demonstrates that the proposed model, HVSF, outperforms both the self-supervised models and supervised models, affirming its superior performance and stability.
Youquan Fu, Yue Ma 0008, Zhixi Feng, Shuyuan Yang 0001, Yixing Wang
IEEE Internet Things J.2
2025 Cross-sensor contrastive learning-based pre-training for machinery fault diagnosis under sample-limited conditions
Yue Ma 0008, Ruoxue Li, Zhixi Feng, Shuyuan Yang 0001, Shaoyi Du, Yue Gao 0002
Knowl. Based Syst.2
2025 DAE-GSP: Discriminative Autoencoder With Gaussian Selective Patch for Multimodal Remote Sensing Image Classification
abstract
In the field of multimodal remote sensing image (MRSI) classification, self-supervised learning (SSL) algorithms have demonstrated significant advantages, particularly in scenarios with limited labeled samples. Existing SSL methods typically use auxiliary tasks within either contrastive or generative frameworks, focusing on discriminative or structural information separately. In this article, we propose a novel hybrid SSL paradigm, discriminative autoencoder with Gaussian selective patch (DAE-GSP) for MRSI classification. The DAE framework integrates contrastive learning with the masked image modeling (MIM) technique, allowing for simultaneous learning of structural information and discriminative representations from images. Furthermore, a cross-attention-based data-level fusion strategy is introduced during pretraining stage to enhance intermodal interactions, thereby improving the effectiveness of modality fusion. In addition, we propose a novel Gaussian selective patch (GSP) strategy, addressing the limitations of traditional square patch selection methods. Combined with self-supervised auxiliary tasks, this strategy facilitates the improved integration of multiple modalities and encourages the model to capture essential semantic information. Extensive experiments conducted on three public datasets (Houston2013, Augsburg, and Berlin) demonstrate the effectiveness of the proposed approach. With only ten labeled training samples per class, the proposed method achieves overall accuracy (OA) of 90.15%, 82.64%, and 71.03% on the Houston2013, Augsburg, and Berlin datasets, respectively, indicating improvements of 1.31%, 1.22%, and 1.48% over state-of-the-art methods.
Mengchang Li, Zhixi Feng, Shuyuan Yang 0001, Yue Ma 0008, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2025 Learning Temporal-Spectral Feature Fusion Representation for Radio Signal Classification
abstract
With the rapid development of wireless communications, industrial electromagnetic environments are facing challenges in terms of spectrum scarcity and cyberspace threats. Moreover, the coexistence of various types of radio signals within the same frequency band may cause signal distortion and degrade the quality and efficiency of communication. To effectively address these challenges, a novel temporal–spectral feature fusion network (TSFFN) for radio signal classification (RSC) is proposed. TSFFN adopts a Cutmix-based temporal–spectral fusion and an attention-based multiview feature fusion mechanism. These mechanisms automatically learn and merge spectrogram, temporal–spectral, and time-domain features by effectively combining temporal and spectral information into high-dimensional representations. This augmentation enhances the network's ability to discriminate different radio signals, enabling accurate spectrum sensing and signal identification for effective spectrum management. Experimental results on five datasets demonstrate the effectiveness of our approach in enhancing RSC performance across diverse industrial scenarios.
Zhixi Feng, Yue Ma 0008, Yachen Gao, Shuyuan Yang 0001
IEEE Trans. Ind. Informatics3
2025 A Generative Self-Supervised Framework for Cognitive Radio Leveraging Time-Frequency Features and Attention-Based Fusion
abstract
With the advancement of cognitive radio technology (CRT) in radio communication networks, deep learning (DL) has become instrumental in enhancing spectrum efficiency. However, supervised DL methods demand extensive labeled data and incur high manual costs. Consequently, practical applications of CRT increasingly necessitate techniques capable of learning robust representations from large volumes of unlabeled data. Although recent DL advancements have driven the use of self-supervised learning (SSL) in CRT through time-domain contrastive methods, these approaches fall short in extracting high-level spectral representations due to their neglect of time-frequency features. To address these limitations, a generative SSL framework is proposed for CRT applications. First, SSL pretraining is conducted in the time-frequency domain by reconstructing masked spectrograms using a Masked Autoencoder. Then, to recover the spectrogram under extreme radio conditions, mutual information maximization is employed to extract high-level spectral information obscured by noise patterns. Additionally, an attention-based channel-spectrum fusion module is designed to automatically extract and integrate features from the channel and spectral domains. The feasibility of the proposed framework is evaluated across multiple downstream tasks on four public datasets. Experimental results demonstrate that the proposed framework significantly outperforms existing methods in various downstream tasks.
Zhixi Feng, Shuyuan Yang 0001, Yue Ma 0008, Zhuoyue Qi
IEEE Trans. Wirel. Commun.4
2024 A Novel Cross-Sensor Self-Supervised Learning Method for Rotating Machinery Fault Diagnosis
abstract
Fault diagnosis is crucial in mechanical prognostics and health management. However, fault features extracted from single-sensor data are limited in complex operating environments. Extracting complementary and robust fault features from multi-sensor monitoring data is essential, especially under limited labeled samples. Leveraging the advantages of self-supervised learning, we propose a novel cross-sensor self-supervised learning (CSSL) method for rotating machinery fault diagnosis under limited sample conditions. Our method employs contrastive learning across multiple sensors, including both intra-sensor and inter-sensor contrastive learning, to derive robust cross-sensor fault representations. The efficacy of our approach is substantiated on two benchmark datasets, revealing superior classification performance. Furthermore, the experimental results under various operating conditions demonstrate outstanding performance and solid robustness.
Zhixi Feng, Ruoxue Li, Yue Ma 0008, Shuyuan Yang 0001
ICASSP4
2024 Multi-Scale Sparse Transformer for Remote Sensing Scene Classification
abstract
Vision Transformer (ViT) has achieved great success in the field of computer vision since it was proposed, and there have been many works applying ViT based models to remote sensing scene classification (RSSC) tasks. The proposal of Pyramid Vision Transformer (PVT) greatly reduces the calculation amount of the ViT while maintaining accuracy. But PVT did not utilize multi-scale information in remote sensing (RS) scenes, which is crucial for RSSC. This paper proposes a multi-scale sparse transformer (MST) based on PVT. MST enables the network to learn multi-scale representations of RS scenes through spatial reduction implementations at different scales. In addition, we employ sparse operations to adaptively guide the model’s attention towards semantically relevant regions during self-attention computation, thereby reducing interference from semantically irrelevant areas. Experiments conducted on the UCM and AID datasets demonstrate the outstanding performance of the proposed MST.
Xu Tang 0004, Zhixi Feng, Yue Ma 0008, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao
IGARSS4
2024 Center Mask Self-Attention Network for Hyperspectral Image Classification
abstract
Benefiting from the thousands of continuous band information in hyperspectral images (HSIs), the task of HSI classification has become an indispensable part of the field of remote sensing. With the development of deep learning, deep learning techniques such as convolutional neural networks have been widely introduced into HSI classification research. However, most of these methods do not fully consider the potential relationship between the central pixel and surrounding neighborhoods. Therefore, we introduce a novel center mask self-attention network (CMSAN) to enable the model to effectively capture the association between the central pixel and its neighbors for better feature extraction. We conduct experiments on two publicly available HSI datasets. The positive results on both datasets fully demonstrate the effectiveness of our proposed method.
Yizhou Zou, Xu Tang 0004, Yue Ma 0008, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao
IGARSS3
2024 D3R-Net: Denoising Diffusion-Based Defense Restore Network for Adversarial Defense in Remote Sensing Scene Classification
abstract
Deep learning models (algorithms) have demonstrated their superior performance in interpreting Earth science and remote sensing data. However, adversarial examples generated with perturbations imperceptible to humans could render deep learning algorithms ineffective. This significant vulnerability of deep learning models, thus, inspires the exploration of defense methods resistible to adversarial examples. Although numerous countermeasures against adversarial examples have been proposed, the design of a universally applicable defense method across multiple scenarios still remains to be explored. In this study, we propose an effective denoising diffusion-based defense restore network (D3R-Net) based on the denoising diffusion model from the perspective of adversarial restoration, which transforms the adversarial examples into clean samples. Utilizing a highly effective denoising diffusion probabilistic model (DDPM), our D3R-Net transforms input adversarial examples into a state of noise, where diverse forms of adversarial noise transition into Gaussian noise. Subsequently, it captures semantic information through a series of iterative denoising steps. The pixel distribution of adversarial examples is restored in the proposed network to match the original distribution, enabling the classifier to identify adversarial examples correctly. Furthermore, we introduce a combined filtering module to preserve the semantic information of the original image, thereby further enhancing the defensive performance. Instead of modifying the model structure or excluding suspected samples, the proposed method restores the adversarial examples, making it simple yet effective and applicable to a broader range of scenarios. Extensive experiments are conducted on four benchmark datasets, and the results demonstrate that D3R-Net has significant defense capabilities against known and unknown attacks. Our source code is available athttps://github.com/SIM-xidian/D3R-Net.
Xuehu Liu, Zhixi Feng, Yue Ma 0008, Shuyuan Yang 0001, Zhihao Chang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2024 Generative Model With Sinkhorn-Knopp Loss for Unsupervised Signal Modulation Clustering
abstract
Modulation types clustering (MC) is crucial for adaptive high-frequency communication between devices in the Industrial Internet of Things. The strength of MC resides in its self-supervised framework, enabling it to extract modulation features efficiently without any manual labeling. However, the misalignment of proxy tasks and erroneous pseudolabeling constrain the performance of prevalent MC feature extraction techniques that utilize time series signals. In this article, we compare the saliency maps on time–frequency image (TFI) with that on time series signal, highlighting the consistency of TFI reconstruction with modulation feature extraction. Subsequently, in order to address the sensitivity of K-means to outliers, Sinkhorn–Knopp labeling (SKLb) is proposed to balance the scale of clusters and neighboring distances. Moreover, in consideration of the potential instability of the SKLb iteration result in backpropagation, the Sinkhorn–Knopp loss is proposed to ensure stable training of the model. Finally, two models, SK-IDC and SK-STDC, were tested on four datasets. Experimental results on these datasets present that our approach outperforms original signal representation and prevalent deep clustering methods, achieving State-of-the-Art performance.
Zhixi Feng, Shuyuan Yang 0001, Yue Ma 0008
IEEE Trans. Ind. Informatics5
2018 Discriminant sparse and collaborative preserving embedding for bearing fault diagnosis
Yue Ma 0008
Neurocomputing1