Siyue Feng

dblp:337/0846 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-6694-9126ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Eler: Ensemble Learning-Based Automated Verification of Code Clones
Shihan Dou, Siyue Feng, Yueming Wu 0001, Deqing Zou
IEEE Trans. Software Eng.3
2025 MalScan: Android Malware Detection Based on Social-Network Centrality Analysis
abstract
Malware scanning of an app market is expected to be scalable and effective. However, existing approaches use syntax-based features that can be evaded by transformation attacks or semantic-based features which are usually extracted by expensive program analysis. Therefore, to address the scalability challenges of traditional heavyweight static analysis, we propose a graph-based lightweight approachMalScanfor Android malware detection.MalScanconsiders the function call graph as a complex social network and employs centrality analysis on sensitiveapplication program interfaces(APIs) to express the semantic characteristics of the graph. On this basis, machine learning algorithms and ensemble learning algorithms are applied to classify the extracted features. We evaluateMalScanon datasets of 104,892 benign apps and 108,640 malwares, and the results of experiments indicate thatMalScanoutperforms six state-of-the-art detectors and can quickly detect Android malware with an f-value as high as 99%. In addition, there are also significant improvements in the robustness of Android app evolution and robustness to obfuscation. Finally, we conduct an exhaustive statistical study of over one million applications in the Google-Play app market and successfully identify 498 zero-day malware, which further validates the feasibility ofMalScanon market-wide malware scanning.
Yueming Wu 0001, Wenqi Suo, Siyue Feng, Deqing Zou, Wei Yang 0013, Yang Liu 0003, Hai Jin 0001
IEEE Trans. Dependable Secur. Comput.3
2025 Fine-Grained Code Clone Detection by Keywords-Based Connection of Program Dependency Graph
abstract
Code clone detection is intended to identify functionally similar code fragments, a matter of escalating significance in contemporary software engineering. Numerous methodologies have been proffered for the detection of code clones, among which graph-based approaches exhibit efficacy in addressing semantic code clones. However, they all only consider the feature extraction of a single sample and ignore the semantic connection between different samples, resulting in the detection effect being unsatisfactory. Simultaneously, the majority of existing methods can only ascertain the presence of clones, lacking the capability to provide nuanced insights into which lines of code exhibit greater similarity. In this article, we advocate a novel PDG-based semantic clone detection method, namely,Keyborwhich can locate specific cloned lines of code by providing a fine-grained analysis of clone pairs. The highlight of the approach is to consider keywords as a bridge to connect PDG nodes of the target program to retain more semantic information about the functional code. To examine the effectiveness ofKeybor, we assess it on a widely usedBigCloneBenchdataset. Experimental results indicate thatKeyboris superior to 14 advanced code clone detection tools (i.e.,CCAligner,SourcererCC,Siamese,NIL,NiCad,LVMapper,CCFinder,CloneWorks,Oreo,Deckard,CCGraph,Code2Img,GPT-3.5-turbo, andGPT-4).
Yueming Wu 0001, Wenqi Suo, Siyue Feng, Cong Wu 0003, Deqing Zou, Hai Jin 0001
IEEE Trans. Reliab.3
2024 Machine Learning is All You Need: A Simple Token-based Approach for Effective Code Clone Detection
abstract
As software engineering advances and the code demand rises, the prevalence of code clones has increased. This phenomenon poses risks like vulnerability propagation, underscoring the growing importance of code clone detection techniques. While numerous code clone detection methods have been proposed, they often fall short in real-world code environments. They either struggle to identify code clones effectively or demand substantial time and computational resources to handle complex clones. This paper introduces a code clone detection method namely Toma using tokens and machine learning. Specifically, we extract token type sequences and employ six similarity calculation methods to generate feature vectors. These vectors are then input into a trained machine learning model for classification. To evaluate the effectiveness and scalability of Toma, we conduct experiments on the widely used BigCloneBench dataset. Results show that our tool outperforms token-based code clone detectors and most tree-based clone detectors, demonstrating high effectiveness and significant time savings.
Siyue Feng, Wenqi Suo, Yueming Wu 0001, Deqing Zou, Yang Liu 0003, Hai Jin 0001
ICSE1
2024 FIRE: Combining Multi-Stage Filtering with Taint Analysis for Scalable Recurring Vulnerability Detection
Siyue Feng, Yueming Wu 0001, Wenjie Xue, Sikui Pan, Deqing Zou, Yang Liu 0003, Hai Jin 0001
USENIX Security Symposium1
2024 Goner: Building Tree-Based N-Gram-Like Model for Semantic Code Clone Detection
abstract
Code clone detection refers to the detection of code fragments that are functionally similar. As software engineering progresses, the significance of code clone detection continues to grow. A number of code clone detection techniques have been designed. Among these methods, tree-based code clone detection approaches can discover semantic code clones. However, given the intricate nature of tree structures, they consume plenty of time to complete the tree analysis, thus cannot scale to large-scale code scanning. In this paper, we propose a novel tree-based scalable semantic code clone detection method by transforming the heavy-weight tree processing into efficient N-gram-like subtrees analysis. Specifically, we build a variant of N-gram model to partition the original complex tree into small subtrees. After collecting all subtrees, we divide them into different groups according to the positions of the subtree nodes, and then calculate the similarity of the same group between two functions one by one. Similarity scores of all groups are made up of a feature vector. Given feature vectors, we train a machine learning model for semantic code clone detection. We implementGonerand conduct evaluations on two extensively utilized datasets, namely BigCloneBench and Google Code Jam. The experimental results indicate thatGoneroutperforms our comparative systems (i.e.SourcererCC,RtvNN,Deckard,ASTNN,TBCNN,CDLH,Amain,FCCA,DeepSim, andSCDetector). Additionally, in the context of scalability,Gonerdemonstrates remarkable speed, being approximately 56 times faster than another advanced tree-based tool, namelyASTNN, when it comes to identifying semantic code clones.
Yueming Wu 0001, Siyue Feng, Wenqi Suo, Deqing Zou, Hai Jin 0001
IEEE Trans. Reliab.2
2023 Tritor: Detecting Semantic Code Clones by Building Social Network-Based Triads Model
abstract
Code clone detection refers to finding the functional similarities between two code fragments, which is becoming increasingly important with the evolution of software engineering. It is reasonable because code cloning can increase maintenance costs and even cause the propagation of vulnerabilities, which can have a negative impact on software security. Numbers of code clone detection methods have been proposed, including tree-based methods that are capable of detecting semantic code clones. However, since tree structure is complex, these methods are difficult to apply to large-scale clone detection. In this paper, we propose a scalable semantic code clone detector based on semantically enhanced abstract syntax tree. Specifically, we add the control flow and data flow details into the original tree and regard the enhanced tree as a social network. Then we build a social network-based triads model to collect the similarity features between the two methods by analyzing different types of triads within the network. After obtaining all features, we use them to train a machine learning-based code clone detector (i.e., Tritor). Our comparative experimental results show that Tritor is superior to SourcererCC, RtvNN, Deckard, ASTNN, TBCNN, CDLH, and SCDetector, are equally good with DeepSim and FCCA. As for scalability, Tritor is about 39 times faster than another current state-of-the-art tree-based code clone detector ASTNN.
Deqing Zou, Siyue Feng, Yueming Wu 0001, Wenqi Suo, Hai Jin 0001
ESEC/SIGSOFT FSE2
2022 Detecting Semantic Code Clones by Building AST-based Markov Chains Model
abstract
Code clone detection aims to find functionally similar code fragments, which is becoming more and more important in the field of software engineering. Many code clone detection methods have been proposed, among which tree-based methods are able to handle semantic code clones. However, these methods are difficult to scale to big code due to the complexity of tree structures. In this paper, we design Amain, a scalable tree-based semantic code clone detector by building Markov chains models. Specifically, we propose a novel method to transform the original complex tree into simple Markov chains and measure the distance of all states in these chains. After obtaining all distance values, we feed them into a machine learning classifier to train a code clone detector. To examine the effectiveness of Amain, we evaluate it on two widely used datasets namely Google Code Jam and BigCloneBench. Experimental results show that Amain is superior to nine state-of-the-art code clone detection tools (i.e., SourcererCC, RtvNN, Deckard, ASTNN, TBCNN, CDLH, FCCA, DeepSim, and SCDetector).
Yueming Wu 0001, Siyue Feng, Deqing Zou, Hai Jin 0001
ASE2