Peng Dai 0007

dblp:08/3547-7 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0003-1919-2498ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 A Reinforcement Learning-Based Approach for Determining Infeasible Paths of Programs
abstract
Program path analysis is an essential component of software defect detection and quality assurance. Accurately identifying infeasible paths can prevent false positives caused by invalid paths, enabling developers to pinpoint actual defects more efficiently and enhancing overall software quality and reliability. This paper proposes an integrated approach for determining infeasible paths based on program path features and constraint-based reinforcement learning. First, a loop-structure path search and reduction algorithm is proposed to systematically simplify path explosion induced by loops. Then, a global subgraph-based path reduction algorithm is introduced to effectively remove redundant and irrelevant paths. Subsequently, we propose a path set generation algorithm guided by control and implication relationships to construct an optimized path set. Path constraints and symbolic path constraints are used to enhance semantic representation. Finally, a reinforcement learning-based model utilizing reachability rewards and exploration rewards to dynamically determine path reachability. Experimental results show that our proposed approach significantly reduces path explosion, accurately identifies infeasible paths and outperforms existing methods in terms of accuracy and computational efficiency.
Peng Dai 0007, Tang He, Zebo Peng, Chen Zhao 0015, Yunzhan Gong
Int. J. Softw. Eng. Knowl. Eng.1
2025 A Static Analysis Framework for Investigating Tainted Data Sources in Software Systems
abstract
One of the most effective methods for detecting software security vulnerabilities is taint analysis. Some software defects originate from certain external input data. Analyzing the taint sources and the data flow propagation from these sources to defect points through static analysis can help us understand the causes of software defects and reduce the difficulty of debugging them. This paper combines intraprocedural and interprocedural analysis methods to obtain global taint source information. A novel propagation path calculation algorithm is proposed, incorporating predecessor node computation and alias analysis, effectively reducing the negative impact of irrelevant code on the performance of taint analysis. This method not only helps detect errors that lead to vulnerabilities but also analyzes the impact of vulnerable input data on the system. Based on the global taint source analysis algorithm, we developed a static taint source analysis prototype tool for C programs, called AWsTS. Experiments conducted on five open-source projects show that AWsTS improves the accuracy of analysis results without increasing the required analysis time. The average precision for intra-procedural taint source analysis is 93.4%, and the average recall is 90.2%. Similarly, for interprocedural taint source analysis, the average precision is 87.6%, and the average recall is 84.9%. Additionally, AWsTS can output taint propagation paths, providing valuable support for further taint analysis.
Peng Dai 0007, Xiaoqin Ma, Zebo Peng, Chen Zhao 0015
Int. J. Softw. Eng. Knowl. Eng.1
2025 FedFM: A federated few-shot learning method by comparison network and model calibration
Chen Zhao 0015, Shu-Di Bao, Meng Chen 0013, Zhipeng Gao 0001, Kaile Xiao, Peng Dai 0007
Knowl. Based Syst.6
2022 An improving approach to analyzing change impact of C programs
Peng Dai 0007, Dahai Jin, Yunzhan Gong
Comput. Commun.1
2022 Improving Large-Gap Clone Detection Recall Using Multiple Features
abstract
Code clone refers to two or more identical or similar source code fragments. Research on code clone detection has lasted for decades. Investigation and evaluation of existing clone detection techniques indicate that they are resilient to function-level clone detection. Still, there may be room for further research in block-level clone detection. Particularly, type-3 clones that include large gaps, are ongoing challenges. To solve these problems, we propose a clone detection method based on multiple code features. It aims to improve the recall rate of code block clone detection and overcome large-gap and hard-to-detect type-3 clones. This method first splits the source code files based on the program’s structural features and context features to obtain code blocks. The collection of code blocks obtained in this way is complete, and the large gaps in clone pairs will also be removed. In addition, we only need to compute the similarity between code blocks with the same structural features, which can also significantly save time and resources. The similarity is obtained by calculating the proportion of the same tokens between two code blocks. Moreover, since different types of tokens have different weights in similarity calculation, we use supervised learning to obtain a classifier model between token features and code clone. We divide the tokens into 13 types and train the machine learning model with the manually confirmed clone or non-clone pair. Finally, we develop a prototype system and compare our tools with existing tools under the Mutation Framework and in several actual C projects. The experimental results also demonstrate the advancement and practicality of our prototype.
Peng Dai 0007, Dahai Jin, Yunzhan Gong
Int. J. Softw. Eng. Knowl. Eng.1