Yujun Dai

dblp:257/5627 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2024 Enhancing human-machine pair inspection with risk number and code inspection diagram
abstract
Abstract Software inspection is a widely-used approach to software quality assurance. Human-Machine Pair Inspection (HMPI) is a novel software inspection technology proposed in our previous work, which is characterized by machine guiding programmers to inspect their own code during programming. While our previous studies have shown the effectiveness of HMPI in telling risky code fragments to the programmer, little attention has been paid to the issue of how the programmer can be effectively guided to carry out inspections. To address this important problem, in this paper we propose to combine Risk Number with Code Inspection Diagram (CID) to provide accurate guidance for the programmer to efficiently carry out inspections of his/her own programs. By following the Code Inspection Diagram, the programmer will inspect every checking item shown in the CID to efficiently determine whether it actually contain bugs. We describe a case study to evaluate the performance of this method by comparing its inspection time and number of detected errors with our previous work. The result shows that the method is likely to guide the programmer to inspect the faulty code earlier and be more efficient in detecting defects than the previous HMPI established based on Cognitive Complexity.
Yujun Dai, Shaoying Liu, Guangquan Xu
Softw. Qual. J.1
2023 Utilizing Risk Number and Program Slicing to Improve Human-Machine Pair Inspection
abstract
Human-Machine Pair Inspection (HMPI) is a novel code inspection technology proposed in our previous work, which is the style that machine will intelligently guide the programmer to carry out inspections of the program code during programming. For large-scale software projects, the efficiency of HMPI needs to be improved due to the inaccurate measurement of the code structure and the excessive inspection scope. In this paper, to alleviate the above deficiencies, we propose the Risk Number, a code evaluation metric generated based on historical error data. The Risk Number is calculated by a statistical tool called regression analysis, which more accurately indicates the relationship between the nested structure of the code and the likelihood of containing bugs than Cognitive Complexity. Additionally, HMPI is supported by utilizing Risk Number to point out high-risk code and program slicing techniques to extract statements that have dependencies on the code to generate checklists, thereby reducing the scope of inspection. We describe a case study to evaluate the performance of this method by comparing its inspection time and number of detected errors with our previous work. The result shows that the method is likely to guide the programmer to inspect the faulty code earlier and be more efficient in detecting defects than HMPI based on Cognitive Complexity.
Yujun Dai, Shaoying Liu, Guangquan Xu, Ai Liu
ICECCS1
2019 Mining Maximal Clique Summary with Effective Sampling
abstract
Maximal clique enumeration (MCE) is a fundamental problem in graph theory and is used in many applications, such as social network analysis, bioinformatics, intelligent agent systems, cyber security. Most existing MCE algorithms focus on improving the efficiency rather than reducing the size of the output, which could consist of a large number of maximal cliques. In this paper, we study how to report a summary of less overlapping maximal cliques. The problem was studied before, however, after examining the pioneer approach, we consider it still not satisfactory. To advance the research along this line, this paper attempts to make two contributions: (a) We propose a more effective sampling strategy, which produces a much smaller summary but still ensures that the summary can somehow witness all the maximal cliques and the expectation of each maximal clique witnessed by the summary is above a predefined threshold. (b) To verify experimentally, we tested ten real benchmark datasets that have a variety of graph characteristics. The results show that our new sampling strategy consistently outperforms the state-of-the-art method by producing smaller summaries and running faster on all the datasets.
Xiaofan Li 0004, Rui Zhou 0001, Yujun Dai, Lu Chen 0008, Chengfei Liu, Qiang He 0001, Yun Yang 0001
ICDM3