Yaoxiang Yu

dblp:302/7345 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 A federated class-incremental learning framework with dynamic client participation and evolution for machine fault diagnosis
Yaoxiang Yu, Xueyi Li 0004, Guangyao Zhang, Wenyang Hu, Tianyang Wang 0001, Shaoze Yan, Fulei Chu
Adv. Eng. Informatics1
2025 Learning hierarchical scene graph and contrastive learning for object goal navigation
Bo Cai 0003, Yaoxiang Yu, Aihua Ke
Knowl. Based Syst.4
2025 Dual-function discriminator for semantic image synthesis in variational GANs
Aihua Ke, Yaoxiang Yu
Pattern Recognit.5
2025 Clustering Weighted Envelope Spectrum for Rolling Bearing Fault Diagnosis
abstract
Spectral coherence (SCoh) is a powerful tool to reveal the hidden periodicities of signals, which has been widely used for rolling bearing fault diagnosis. However, most SCoh-based methods focus on searching a single demodulation band, which results in their inability to compound fault diagnosis and discrete frequency band localization. Moreover, many studies are conducted based on prior fault characteristic frequencies (FCFs), which limits their application in limited vision cases. To solve such issues, a prior knowledge-needless method namely clustering weighted envelope spectrum (CWES) is proposed for rolling bearing fault diagnosis. Firstly, based on the algorithms of peak searching and multiple relation checking, the potential FCFs (PFCFs) of each spectral frequency slice (SFS) of SCoh are automatically identified without any prior knowledge. The PFCFs of each SFS are regarded as its fault type label and are used to design a weight to evaluate its fault information abundance. Then, the SFSs with similar labels are clustered and other SFSs are ignored. Each cluster is considered to be associated with a potential cyclostationary component, and the importance of all clusters is sorted based on their maximum weights. Finally, to further enhance the fault characteristics, CWESs are defined as the weighted average of the SFSs in each top-ranked cluster. By using this method, the discrete informative frequency bands of multiple faults can be quickly located without prior FCFs and iterative optimization. The advantages of CWES over the state-of-the-art methods are validated by the experimental data of bearing single and compound faults. The results indicate that CWES has the best completeness in fault information extraction and the highest accuracy of fault diagnosis compared with other methods. Moreover, the robustness and computational efficiency of the proposed method are also advantageous.Note to Practitioners—This paper is motivated by the problems of discrete frequency band localization and compound fault separation in the field of rolling bearing fault diagnosis. Different from other prior FCF-oriented methods, we design a prior knowledge-needless algorithm to identify the PFCFs of each SFS of the SCoh. The PFCFs of each SFS can not only indicate the fault type but also quantify the abundance of fault information. Based on the identified PFCFs, several CWESs can be generated for fault diagnosis through the clustering algorithm and the weighted mechanism. Our experimental results show the proposed method has higher diagnostic accuracy than the existing methods.
Tao Chen 0024, Liang Guo 0001, Hongli Gao, Tingting Feng, Yaoxiang Yu
IEEE Trans Autom. Sci. Eng.5
2024 APT: Adaptive Prefix-Tuning on Pretrained Models for Code Intelligence
abstract
In the field of code intelligence, pretrained models exhibit impressive performance. However, it is imperative to modify all model parameters and maintain full copies for various tasks. Moreover, the effectiveness of fine-tuning a pretrained model depends on the availability of data, which can be constrained in practical settings. Prefix-tuning, a novel approach in NLP for addressing the aforementioned issue, has demonstrated promising results in numerous tasks. This paper aims to investigate the potential of prefix-tuning in the field of code intelligence, an area that has been relatively understudied. We also acknowledge the critical role of initialization in the performance of prefixtuning, which has been neglected in prior research. To address this gap, we introduce a novel method termed "adaptive prefixtuning". This approach entails training the prefix parameters on adapting tasks, followed by subsequent tuning on downstream tasks. We perform prefix-tuning and adaptive prefix-tuning using well-known pretrained models, CodeT5 and CodeGPT, and conduct experiments on four code intelligence tasks, namely, defect detection, code completion, code summarization, and code translation. Our experiments revealed that, in situations with restricted data availability, adaptive prefix-tuning yields substantial performance enhancements compared to fine-tuning, whereas prefix-tuning does not exhibit evident advantages. Furthermore, adaptive prefix-tuning has demonstrated superior performance relative to prefix-tuning in scenarios with comprehensive datasets. In certain cases, it has even outperformed traditional fine-tuning methods. Our findings indicate that within the field of code intelligence, adaptive prefix-tuning can serve as an effective substitute for fine-tuning, especially in situations with constrained data availability.
Yaoxiang Yu, Zhengming Yuan, Siwei Huang
IJCNN2
2024 Prefix Tuning for Few-shot Malware Classification with Supervised Contrastive Cross-Entropy Learning
abstract
The accurate detection of malware is of paramount importance in today’s society, where the number of computers is vast. Traditional methods for classifying malicious code face three main challenges. Firstly, previous deep learning models encounter difficulties in handling excessively long input sequences due to their structural and computational resource limitations. Secondly, the rapid changes and rapid propagation of malicious samples pose certain challenges in collecting a sufficient number of samples and ensuring the accuracy of sample labels. Thirdly, in scenarios with limited sample availability, prior approaches failed to consider the augmentation of the model’s generalization performance when dealing with a limited number of samples. To address the first challenge, we propose a novel method that segments malicious samples into subroutines to extract opcode sequences. These sequences are then converted into two-dimensional opcode sequences and fed into a pretrained model. Additionally, we transform the relevant features of malicious sample opcodes, registers, and assembly language comments into one-dimensional vectors as prefixes. This approach guides the model to better learn sample feature representations, ultimately enhancing the classification performance. To solve the latter two challenges, we introduce a combination of supervised contrastive learning and cross-entropy loss, utilizing support set computation to generate prototype vectors. This enables the model to effectively tackle emerging variations of malicious code and enhance its performance in few-shot scenarios. Experimental results demonstrate that our method outperforms existing malicious code classification models and achieves outstanding performance in few-shot experiments.
Zhengming Yuan, Yaoxiang Yu, Siwei Huang
IJCNN2
2024 Learning multimodal adaptive relation graph and action boost memory for visual navigation
Bo Cai 0003, Yaoxiang Yu, Aihua Ke
Adv. Eng. Informatics3
2023 Prompt Learning for Developing Software Exploits
abstract
A software exploit is a sequence of commands that exploits software vulnerabilities or security flaws, written either by security researchers as a Proof-Of-Concept (POC) threat or by malicious attackers for use in their operations. Writing exploits is a challenging task since it is time-consuming and costly. Pre-trained Language Models (PLMs) for code can benefit automatic exploit generation, achieving state-of-the-art performance. However, the typical paradigm for using these PLMs is fine-tuning, which leads to a significant gap between the language model pre-training and the target task fine-tuning process since they are in different forms. In this paper, we propose a prompt learning approach PT4Exploits based on the PLM, i.e., CodeT5 to automatically generate desired exploits by modifying the original English description inputs. More specifically, PT4Exploits can better elicit knowledge from the PLM by inserting trainable prompt tokens into the original input to construct the same form as the pre-training tasks. Experimental results show that our approach can significantly outperform baseline models on both automatic evaluation and human evaluation. Additionally, we conduct extensive experiments to investigate the performance of prompt learning on few-shot settings and the scenario of different prompt templates, which also show the competitive effectiveness of PT4Exploits.
Xiaoming Ruan, Yaoxiang Yu, Bo Cai 0003
Internetware2
2023 Pre-trained Model Based Feature Envy Detection
abstract
Code smells slow down software system development and makes them harder to maintain. Existing research aims to develop automatic detection algorithms to reduce the labor and time costs within the detection process. Deep learning techniques have recently been demonstrated to enhance the performance of recognizing code smells even more than metric-based heuristic detection algorithms. As large-scale pre-trained models for Programming Languages (PL), such as CodeT5, have lately achieved the top results in a variety of downstream tasks, some researchers begin to explore the use of pre-trained models to extract the contextual semantics of code to detect code smells. However, little research has employed contextual code semantics relationship between code snippets obtained by pre-trained models to identify code smells. In this paper, we investigate the use of the pre-trained model CodeT5 to extract semantic relationships between code snippets to detect feature envy, which is one of the most common code smells. In addition, to investigate the performance of these semantic relationships extracted by pre-trained models of different architectures on detecting feature envy, we compare CodeT5 with two other pre-trained models CodeBERT and CodeGPT. We have performed our experimental evaluation on ten open-source projects, our approach improves F-measure by 29.32% on feature envy detection and 16.57% on moving destination recommendation. Using semantic relations extracted by several pre-trained models to detect feature envy outperforms the state-of-the-art. This shows that using this semantic relation to detect feature envy is promising. To enable future research on feature envy detection, we have made all the code and datasets utilized in this article open source.
Yaoxiang Yu, Xiaoming Ruan
MSR2
2023 Capture More Structured Context by Vision Transformers for Free-Hand Sketch Recognition
abstract
Free-hand sketch recognition presents unique challenges that require careful consideration of spatial and temporal properties. Convolutional neural networks, which are widely used for sketch recognition, have limited ability to extract structured context, leading to insufficient learning of spatial properties. In this study, we introduce Vision Transformers to the sketch recognition domain to address this issue. To the best of our knowledge, this is the first attempt to evaluate the applicability and effectiveness of Vision Transformers for sketch recognition. In addition, we also address the problem of semantically different categories appearing similar in sketches. To overcome this challenge, we propose a novel loss function for sketch recognition, termed Sketch Online Label Smoothing (SketchOLS). Our experiments on QuickDraw dataset demonstrate that the proposed approach is a robust and effective solution for sketch recognition. Remarkably, our method outperforms existing models even without the injection of additional temporal information. These results highlight the potential of our approach to enhance free-hand sketch recognition tasks.
Weirong Guo, Sirui Liu 0007, Yaoxiang Yu
SMC3
2023 Enhancing Code Search with Token-Level Information Flow Graphs Generated from Aligned Structural and Textual Features
abstract
Large-scale code search is a crucial task in software engineering, yet existing deep learning based models often embed Abstract Syntax Trees (ASTs) and code sequences separately, limiting their ability to learn the correlation between structural and textual features. To address this limitation, we propose a novel code search model that automatically generates Token-Level Information Flow Graphs (TL-IFGs) from aligned AST nodes and source code tokens. Our model includes an aligner that establishes a one-to-one correspondence between AST leaves and code tokens, which we make publicly available, along with a processed dataset to facilitate further research. The model automatically generates a TL-IFG for each code snippet from the aligned datas by predicting the information flow at the token-level, which ensures that structural and textual features are highly correlated during the embedding process. We also generate TL-IFGs for descriptions and embed them using a similar process. Experimental results demonstrate that our model outperforms state-of-the-art code search models, indicating the effectiveness of our approach. Furthermore, an ablation study shows that the generated TL-IFGs for both code and description positively impact model performance.
Sirui Liu 0007, Weirong Guo, Yaoxiang Yu
SMC3
2023 CSSAM: Code Search via Attention Matching of Code Semantics and Structures
abstract
Code search greatly improves developers’ coding efficiency by retrieving reusable code segments with natural language queries. Despite the continuous efforts in improving both the effectiveness and efficiency of code search, two issues remained unsolved. First, programming languages have inherent strong structural linkages, and feature mining of code as text form would omit the structural information contained inside it. Second, there is a potential semantic relationship between code and query, it is challenging to align code and text across sequences so that vectors are spatially consistent during similarity matching.To tackle both issues, in this paper, a code search model named CSSAM (Code Semantics and Structures Attention Matching) is proposed. By introducing semantic and structural matching mechanisms, CSSAM effectively extracts and fuses multidimensional code features. Specifically, the cross and co-attention layer was developed to facilitate high-latitude spatial alignment of code and query at the token level. By leveraging the residual interaction, a matching module is designed to preserve more code semantics and descriptive features, which enhances the relevance between the code and its corresponding query text. Besides, to improve the model’s comprehension of the code’s inherent structure, a code representation structure named CSRG (Code Semantic Representation Graph) is proposed for jointly representing abstract syntax tree nodes and the data flow of the codes. According to the experimental results on two publicly available datasets containing 475k and 330k code segments, CSSAM significantly outperforms the baselines in terms of achieving the highest SR@1/5/10, MRR, and NDCG@50 on both datasets, respectively. Moreover, the ablation study is conducted to quantitatively measure the impact of each key component of CSSAM on the efficiency and effectiveness of code search, which offers insights into the improvement of advanced code search solutions.
Yaoxiang Yu
SANER2
2023 Scene sketch semantic segmentation with hierarchical Transformer
Jie Yang 0076, Aihua Ke, Yaoxiang Yu, Bo Cai 0003
Knowl. Based Syst.3
2022 Online Remaining Useful Life Prediction of Milling Cutters Based on Multisource Data and Feature Learning
abstract
A milling cutter is one of the most important parts of machine tools. Its working status significantly influences the precision of workpiece. Due to the complex wear mechanism, the single sensor may be difficult to acquire the complete degradation information of milling cutters. Therefore, in this article, a feature learning based method is proposed to automatically extract features from multisource data and predict the remaining useful life of cutting tools in real time. First, a statistic-based method is constructed to detect and delete the outliers hidden in the monitoring data. Second, the clean data are input into a multiscale convolutional attention network (MSAN) to learn features and fuse multisource data. At last, the fused data are used to predict the remaining useful life of cutting tools in a regression layer. Compared with traditional tool life prediction methods, the proposed method is able to fuse multisource data through an attention feature learning model to conduct the life prediction of tools. Additionally, the data cleaning and model optimization methods are also proposed to promote engineering practicability. To validate the effectiveness of such method, the life testing experiments on milling cutters are conducted to obtain run-to-failure data. In those experiments, multisensor monitor data are acquired, which are used to conduct validation experiments testing the effectiveness of the proposed method. The results indicate the superiority of the proposed method in remaining useful life prediction milling cutters.
Liang Guo 0001, Yaoxiang Yu, Hongli Gao, Tingting Feng, Yuekai Liu
IEEE Trans. Ind. Informatics2
2022 Pareto-Optimal Adaptive Loss Residual Shrinkage Network for Imbalanced Fault Diagnostics of Machines
abstract
In the industrial applications of mechanical fault diagnosis, machines work in normal condition at most time. In other words, most of the collected datasets are highly imbalanced. Although deep learning has been widely applied in intelligent diagnosis, it is unsuitable for such imbalanced situation. In addition, few studies attempted to determine the parameters in the diagnosis models. For solving such problems, Pareto-optimal adaptive loss residual shrinkage network (PALRSN) is proposed. First, a fixed length-based encoding method is implemented to represent the candidate architectures of PALRSN. Then, multiply accumulate operations and Gmean value representing the model complexity and identification performance, respectively, on imbalanced datasets are selected as the optimization targets to search for the optimal PALRSN architecture. In the training process, an adaptive loss function assigns different misclassification costs on all categories according to their number discrepancy to highlight the minority samples. The proposed method is validated by bearing data and milling cutter data with different imbalanced ratio. The experimental results demonstrate that such approach outperforms the state-of-the-art methods in imbalanced classification.
Yaoxiang Yu, Liang Guo 0001, Hongli Gao, Yuekai Liu, Tingting Feng
IEEE Trans. Ind. Informatics1
2021 Control chart recognition based on the parallel model of CNN and LSTM with GA optimization
Yaoxiang Yu, Min Zhang 0036
Expert Syst. Appl.1