EDBT 2026 Demo / reviewers in the wild / expert
Yaoshen Yu
dblp:248/3780
· DBLP profile ↗
16ranked-venue papers
3as first author
12since 2021 · last 2025
0000-0003-0609-7985ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CPTCP: Incorporating Convolutional Neural Network into Text-Vector Based Test Case Prioritization for Compilers
Chunqi Li, Weiwei Li 0001, Yaoshen Yu |
ICIC (10) | 3 |
| 2025 | An empirical study of best practices for code pre-trained models on software engineering classification tasks
Lina Gong, Yaoshen Yu, Mingqiang Wei |
Expert Syst. Appl. | 3 |
| 2025 | LevDetectCode: Zero-Shot Detection of AI-Generated Code Using Levenshtein DistanceabstractAI-generated code is spreading rapidly, making it harder to trace the true origins of code. Text-based tools such as DetectGPT perform well on texts but struggle with code. We propose a new approach: under masked reconstruction, code produced by LLMs exhibits a much wider spread in Levenshtein distance than human-written code. Our key innovation is to treat this Levenshtein distance dispersion as an indicator for token-probability differences, leading us to develop LevDetectCode. Our zero-shot detector relies solely on string-level Levenshtein distance, requiring neither GPUs nor access to model internals. Experiments on Python snippets from MBPP-train and HumanEval, as well as Java data from CodeContest_Java_test, using GPT-3.5, GPT-4, and Deepseek-V3, show a 5–10 point AUROC improvement over other zero-shot baselines, with an additional case study for the C[Formula: see text] dataset, which also demonstrates strong effectiveness. Furthermore, it runs nearly 3000 times faster and significantly reduces memory usage. These results demonstrate that Levenshtein distance can outperform other zero-shot methods that rely on model-based computations, offering a practical and portable solution for detecting AI-generated code. Jiazhou Fu, Guohua Shen, Yaoshen Yu |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2024 | ASTSDL: predicting the functionality of incomplete programming code via an AST-sequence-based deep learning model
Yaoshen Yu, Guohua Shen, Weiwei Li 0001, Yichao Shao |
Sci. China Inf. Sci. | 1 |
| 2024 | FuEPRe: a fusing embedding method with attention for post recommendation
Xinbo Zhang, Guohua Shen, Yaoshen Yu |
Serv. Oriented Comput. Appl. | 4 |
| 2023 | Shrinking the Semantic Gap: Spatial Pooling of Local Moment Invariants for Copy-Move Forgery DetectionabstractCopy-move forgery is a manipulation of copying and pasting specific patches from and to an image, with potentially illegal or unethical uses. Recent advances in the forensic methods for copy-move forgery have shown increasing success in detection accuracy and robustness. However, for images with high self-similarity or strong signal corruption, the existing algorithms often exhibit inefficient processes and unreliable results. This is mainly due to the inherent semantic gap between low-level visual representation and high-level semantic concept. In this paper, we present a very first study of trying to mitigate the semantic gap problem in copy-move forgery detection, with spatial pooling of local moment invariants for midlevel image representation. Our detection method expands the traditional works on two aspects: 1) we introduce the bag-of-visual-words model into this field for the first time, may meaning a new perspective of forensic study; 2) we propose a word-to-phrase feature description and matching pipeline, covering the spatial structure and visual saliency information of digital images. Extensive experimental results show the superior performance of our framework over state-of-the-art algorithms in overcoming the related problems caused by the semantic gap. Chao Wang 0028, Yaoshen Yu, Guohua Shen, Yushu Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Improved Methods of Pointer Mixture Network for Code CompletionabstractCode completion is an efficient software development technique in modern integrated development environments (IDEs), which can predict the most likely code token(s) based on the context of the code to be completed, so as to improve the work efficiency of developers. The Pointer Mixture Network proposed in recent years has achieved good results in code completion, the contribution of this paper is to improve the Pointer Mixture Network’s method. We used one-hot encoding in the data preprocessing phase, which makes the distance between the tokens of calculation more reasonable, and also has an effect on the expansion characteristics of the code. Besides, we add label smoothing to avoid the overfitting of neural language networks and improve the generalization ability of the model. In neural language networks, we apply the three-layer LSTM, so that the hidden layers of LSTM can fully learn the context information. In terms of the optimizer, we choose NAdam whose performance is better than Adam used in the Pointer Mixture Network, which greatly accelerates the training speed of the model. Experiments show that our work exceeds the results obtained in the Pointer Mixture Network, which is in code completion tasks in Python and JavaScript programming languages. Yaoshen Yu |
QRS | 3 |
| 2022 | Multi-Modal Code Summarization with Retrieved SummaryabstractA high-quality code summary describes the functionality and purpose of a code snippet concisely, which is key to program comprehension. Automatic code summarization aims to generate natural language summaries from code snippets automatically, which can save developers time and improve efficiency in development and maintenance. Recently, researchers mainly use neural machine translation (NMT) based approaches to fill this task. They apply a neural model to translate code snippets into natural language summaries. However, the performance of existing NMT-based approaches is limited. Although a summary and a code snippet are semantically related, they may not share common lexical tokens or language structures. Such a semantic gap between codes and summaries hinders the effect of NMT-based models. Only using code tokens to represent a code snippet cannot help NMT-based models overcome this gap. To solve this problem, in this paper, we propose a code summarization approach that incorporates lexical, syntactic and semantic modalities of codes. We treat code tokens as the lexical modality and the abstract syntax tree (AST) as the syntactic modality. To obtain the semantic modality, inspired by translation memory (TM) in NMT, we use the information retrieval (IR) technique to retrieve a relevant summary for a code snippet to describe its functionality. We propose a novel approach based on contrastive learning to build a retrieval model to retrieve semantically similar summaries. Our approach learns and fuses those different modalities using Transformer. We evaluate our approach on a large Java dataset, experiment results show that our approach outperforms the state-of-the-art approaches on automatic evaluation metrics BLEU, ROUGE and METEOR by 10%, 8% and 9%. Lile Lin, Yaoshen Yu |
SCAM | 3 |
| 2022 | Fast code recommendation via approximate sub-tree matchingabstractSoftware developers often write code that has similar functionality to existing code segments. A code recommendation tool that helps developers reuse these code fragments can significantly improve their efficiency. Several methods have been proposed in recent years. Some use sequence matching algorithms to find the related recommendations. Most of these methods are time-consuming and can leverage only low-level textual information from code. Others extract features from code and obtain similarity using numerical feature vectors. However, the similarity of feature vectors is often not equivalent to the original code’s similarity. Structural information is lost during the process of transforming abstract syntax trees into vectors. We propose an approximate sub-tree matching based method to solve this problem. Unlike existing tree-based approaches that match feature vectors, it retains the tree structure of the query code in the matching process to find code fragments that best match the current query. It uses a fast approximation sub-tree matching algorithm by transforming the sub-tree matching problem into the match between the tree and the list. In this way, the structural information can be used for code recommendation tasks that have high time requirements. We have constructed several real-world code databases covering different languages and granularities to evaluate the effectiveness of our method. The results show that our method outperforms two compared methods, SENSORY and Aroma, in terms of the recall value on all the datasets, and can be applied to large datasets. Yichao Shao, Weiwei Li 0001, Yaoshen Yu |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2022 | ASTENS-BWA: Searching partial syntactic similar regions between source code fragments via AST-based encoded sequence alignment
Yaoshen Yu, Guohua Shen, Weiwei Li 0001, Yichao Shao |
Sci. Comput. Program. | 1 |
| 2021 | Supporting Requirements to Code Traceability Creation by Code CommentsabstractRequirements-to-code tracing is an important and costly task that creates trace links from requirements to source code. These trace links help engineers reduce the time and complexity of software maintenance. Code comments play an important role in software maintenance tasks. However, few studies have focused intensively on the impact of code comments on requirements-to-code trace links creation. Different types of comments have different purposes, so how different types of code comments provide different improvements for requirements-to-code trace links creation? We focus on learning whether code comments and different types of comments can improve the quality of trace links creation. This paper presents a study to evaluate the contribution of code comments and different types of code comments to the creation of trace links. More specifically, this paper first experimentally evaluates the impact of code comments on requirements-to-code trace links creation, and then divides code comments into six categories to evaluate its impact on trace links creation. The results show that the precision increases by an average of 15% (based on the same recall) after adding code comments (even for different trace links creation techniques), and the type of Purpose comments contributes more to the tracing task than the other five. This empirical study provides evidence that code comments are effective in tracing links creation, and different types of code comments contribute differently. Purpose comments can be used to improve the accuracy of requirements-to-code trace links creation. Guohua Shen, Haijuan Wang, Yaoshen Yu |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2021 | Analyzing close relations between target artifacts for improving IR-based requirement traceability recoveryabstractRequirement traceability is an important and costly task that creates trace links from requirements to different software artifacts. These trace links can help engineers reduce the time and complexity of software maintenance. The information retrieval (IR) technique has been widely used in requirement traceability. It uses the textual similarity between software artifacts to create links. However, if two artifacts do not share or share only a small number of words, the performance of the IR can be very poor. Some methods have been developed to enhance the IR by considering relations between target artifacts, but they have been limited to code rather than to other types of target artifacts. To overcome this limitation, we propose an automatic method that combines the IR method with the close relations between target artifacts. Specifically, we leverage close relations between target artifacts rather than just text matching from requirements to target artifacts. Moreover, the method is not limited to the type of target artifacts when considering the relations between target artifacts. We conduct experiments on five public datasets and take account of trace links between requirements and different types of software artifacts. Results show that under the same recall, the precisions on the five datasets improve by 40%, 8%, 20%, 4%, and 6%, respectively, compared with the baseline method. The precision on the five datasets improves by an average of 15.6%, showing that our method outperforms the baseline method when working under the same conditions. Haijuan Wang, Guohua Shen, Yaoshen Yu |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2020 | ASPDup: AST-Sequence-based Progressive Duplicate Code Detection Tool for Onsite Programming CodeabstractDuplicate code is an example of bad smells, which are usually been refactored after the detection to improve the quality of programs. Locate the duplicate code at the programming phase may reduce the cost of maintenance, but the challenge is it need to detect duplicate code between an incomplete code fragment with complete files, which the existing tools are hard to be applied to this scenario. In this paper, we propose an AST-sequence-based duplicate code detection approach for onsite programming code. The abstract syntax tree (AST) is extracted from source code and then is transformed into an encoded sequence. A local sequence alignment algorithm is used to find highly similar subsequences. After the post-processing, similar regions will be found between two code fragments according to the subsequences. We have developed a prototype tool as a plugin for Visual Studio Code. Experimental results indicate that our approach is effective in finding highly similar regions between cross-granularity code fragments, which can facilitate duplicate code detection for incomplete onsite programming code. Yaoshen Yu, Yu Zhou 0010, Weiwei Li 0001, Yichao Shao |
Internetware | 1 |
| 2020 | A topology and risk-aware access control framework for cyber-physical space
Yan Cao 0005, Yaoshen Yu, Changbo Ke |
Frontiers Comput. Sci. | 3 |
| 2020 | Automatic traceability link recovery via active learningabstractTraceability link recovery (TLR) is an important and costly software task that requires humans establish relationships between source and target artifact sets within the same project. Previous research has proposed to establish traceability links by machine learning approaches. However, current machine learning approaches cannot be well applied to projects without traceability information (links), because training an effective predictive model requires humans label too many traceability links. To save manpower, we propose a new TLR approach based on active learning (AL), which is called the AL-based approach. We evaluate the AL-based approach on seven commonly used traceability datasets and compare it with an information retrieval based approach and a state-of-the-art machine learning approach. The results indicate that the AL-based approach outperforms the other two approaches in terms of F-score. Tianbao Du, Guohua Shen, Yaoshen Yu, Dexiang Wu |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2019 | SENSORY: Leveraging Code Statement Sequence Information for Code Snippets RecommendationabstractSoftware developers often have to implement unfamiliar programming tasks. When faced with these problems, developers often search online for code snippets as references to learn how to solve the unfamiliar tasks. In recent years, some researchers propose several approaches to use programming context to recommend code snippets. Most of these approaches use information retrieval based techniques and treat code snippets as a set of tokens. However, in code, the smallest meaningful unit is code statement, in general, the line of code. Since these studies did not consider this issue, there is still room for improvement in the code snippets recommendation. In this paper, we propose a code Statement sEquence iNformation baSed cOde snippets Recommendation sYstem (SENSORY). Different from existing token based approaches, SENSORY performs code snippets recommendation at code statement granularity. It uses the Burrows Wheeler Transform algorithm to search relevant code snippets, and uses the structure information to re-rank the results. To evaluate the effectiveness of our proposed method, we construct a code database with 1000000 real world code snippets which contain more than 15000000 lines of code. The experimental results show that SENSORY outperforms the two strong baseline work in terms of precision and NDCG. Lei Ai, Weiwei Li 0001, Yu Zhou 0010, Yaoshen Yu |
COMPSAC (1) | 5 |