VLDB 2026 Research / reviewers in the wild / expert
Yunzhan Gong
dblp:83/6569
· DBLP profile ↗
24ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-9105-4797ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 5 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Systems, architecture and hardware · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Reinforcement Learning-Based Approach for Determining Infeasible Paths of ProgramsabstractProgram path analysis is an essential component of software defect detection and quality assurance. Accurately identifying infeasible paths can prevent false positives caused by invalid paths, enabling developers to pinpoint actual defects more efficiently and enhancing overall software quality and reliability. This paper proposes an integrated approach for determining infeasible paths based on program path features and constraint-based reinforcement learning. First, a loop-structure path search and reduction algorithm is proposed to systematically simplify path explosion induced by loops. Then, a global subgraph-based path reduction algorithm is introduced to effectively remove redundant and irrelevant paths. Subsequently, we propose a path set generation algorithm guided by control and implication relationships to construct an optimized path set. Path constraints and symbolic path constraints are used to enhance semantic representation. Finally, a reinforcement learning-based model utilizing reachability rewards and exploration rewards to dynamically determine path reachability. Experimental results show that our proposed approach significantly reduces path explosion, accurately identifies infeasible paths and outperforms existing methods in terms of accuracy and computational efficiency. Peng Dai 0007, Tang He, Zebo Peng, Chen Zhao 0015, Yunzhan Gong |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2024 | Multi-type Vulnerability Detection with Staged Feature Fusion and Group Data BalanceabstractWith the progress of software technology, vulnerability detection is more important in software security. The mainstream method lacks discriminability and explainability as it uses graph neural network to fuse features simultaneously. Meanwhile only showing the presence of vulnerabilities limits its usefulness. Therefore, we propose BMAAVD, a Vulnerability Detection framework based on Bilateral Masks Aggregate Attention mechanism, including a staged fusion based on feature level with adaptive structures. It treats detection as a multi-classification task and outputs classification information to improve usefulness. BMAAVD performs well on both synthetic and real datasets, achieving an average increase of 7.61 % and 14.82% in F1 score, and the group bagging training strategy we promoted improves the model's performance in data imbalance. Boyang Zheng, Dongming Zhu, Yunzhan Gong |
ICTAI | 4 |
| 2023 | Incremental Reliability Assessment of Large-Scale Software via Theoretical Structure Reduction
Yunzhan Gong |
ICSOFT | 4 |
| 2023 | Semantic Clone Detection Based on Code Feature Fusion LearningabstractCode clones are duplicated code snippets that significantly threaten software maintenance and the public corpora of code representation learning. Traditionally, code context and its structure information abstract syntax tree (AST), control flow graph (CFG) are typical representations of source code, and context-based models and structure-based models contributed significantly to the development of code clone detection. In this paper, we present a hybrid embedding model for code clone detection (HEM-CCD), a fusion method of token sequential information and graph-based structure information. We insert tokens’ global context information encoded by a bi-directional recurrent neural network into the AST-based graph for comprehensive code semantic representation. Then, feeding the graph into a gated graph neural network we generate code semantic vectors for similarity evaluation. We have implemented our model on two public clone datasets (BigCloneBench and GoogleCodeJam), and the results indicate that HEM-CCD outperforms several state-of-the-art approaches. Dahai Jin, Yunzhan Gong |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2022 | An improving approach to analyzing change impact of C programs
Peng Dai 0007, Dahai Jin, Yunzhan Gong |
Comput. Commun. | 4 |
| 2022 | Eliminating the high false-positive rate in defect prediction through BayesNet with adjustable weightabstractAbstract In defect prediction, a high false‐positive rate (FPR) caused by class imbalance not only increases the workload of testing and development but also consumes unnecessary costs. Many defect models against class imbalance have been proposed to improve the accuracy of defect prediction, but their ability to reduce FPR is unclear. To solve these problems, we first proposed a BayesNet with adjustable weights, called WBN, to reduce the FPR in software defect prediction, which is an algorithm independent of data preprocessing techniques. The mechanism of our WBN is to change the sampling probability of the misclassified instances when training the defect model, making the BayesNet model focus more on false alarm instances. And then, we investigate the FPR of five mainstream defect models for solving class imbalance and select them as comparison models to test the validity of our methods. The experimental result on eight open‐source projects shows that a) our WBN, in in‐version defect prediction (IVDP) and cross‐version defect prediction (CVDP), effectively reduces FPR with means of 0.384 and 0.322, respectively; b) compared with improved subclass discriminant analysis (ISDA) that is the lowest FPR in all control models, our WBN not only reduced the FPR but maintained recall whose mean value was 0.797, whereas ISDA did not, with an average recall of only 0.397; c) our WBN, in CVDP, not only reduces FPR, but also has significant superiority over five control defect models and baseline. Besides, we also found that the class imbalance difference between the test set and the training set has an impact on CVDP performance, recommending that practitioners choose the best dataset for CVDP from the defect data of the historical version through special technology. Yanyang Zhao, Dalin Zhang 0003, Yunzhan Gong |
Expert Syst. J. Knowl. Eng. | 4 |
| 2022 | Improving Large-Gap Clone Detection Recall Using Multiple FeaturesabstractCode clone refers to two or more identical or similar source code fragments. Research on code clone detection has lasted for decades. Investigation and evaluation of existing clone detection techniques indicate that they are resilient to function-level clone detection. Still, there may be room for further research in block-level clone detection. Particularly, type-3 clones that include large gaps, are ongoing challenges. To solve these problems, we propose a clone detection method based on multiple code features. It aims to improve the recall rate of code block clone detection and overcome large-gap and hard-to-detect type-3 clones. This method first splits the source code files based on the program’s structural features and context features to obtain code blocks. The collection of code blocks obtained in this way is complete, and the large gaps in clone pairs will also be removed. In addition, we only need to compute the similarity between code blocks with the same structural features, which can also significantly save time and resources. The similarity is obtained by calculating the proportion of the same tokens between two code blocks. Moreover, since different types of tokens have different weights in similarity calculation, we use supervised learning to obtain a classifier model between token features and code clone. We divide the tokens into 13 types and train the machine learning model with the manually confirmed clone or non-clone pair. Finally, we develop a prototype system and compare our tools with existing tools under the Mutation Framework and in several actual C projects. The experimental results also demonstrate the advancement and practicality of our prototype. Peng Dai 0007, Dahai Jin, Yunzhan Gong |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2022 | ST-TLF: Cross-version defect prediction framework based transfer learningabstractCross-version defect prediction (CVDP) is a practical scenario in which defect prediction models are derived from defect data of historical versions to predict potential defects in the current version. Prior research employed defect data of the latest historical version as the training set using the empirical recommended method, ignoring the concept drift between versions, which undermines the accuracy of CVDP. We customized a Selected Training set and Transfer Learning Framework (ST-TLF) with two objectives: a) to obtain the best training set for the version at hand, proposing an approach to select the training set from the historical data; b) to eliminate the concept drift, designing a transfer strategy for CVDP. To evaluate the performance of ST-TLF, we investigated three research problems, covering the generalization of ST-TLF for multiple classifiers, the accuracy of our training set matching methods, and the performance of ST-TLF in CVDP compared against state-of-the-art approaches. The results reflect that (a) the eight classifiers we examined are all boosted under our ST-TLF, where SVM improves 49.74% considering MCC, as is similar to others; (b) when performing the best training set matching, the accuracy of the method proposed by us is 82.4%, while the experience recommended method is only 41.2%; (c) comparing the 12 control methods, our ST-TLF (with BayesNet), against the best contrast method P15-NB, improves the average MCC by 18.84%. Our framework ST-TLF with various classifiers can work well in CVDP. The training set selection method we proposed can effectively match the best training set for the current version, breaking through the limitation of relying on experience recommendation, which has been ignored in other studies. Also, ST-TLF can efficiently elevate the CVDP performance compared with random forest and 12 control methods. Yanyang Zhao, Yuwei Zhang 0003, Dalin Zhang 0003, Yunzhan Gong, Dahai Jin |
Inf. Softw. Technol. | 5 |
| 2020 | Web service composition on IoT reliability test based on cross entropyabstractAbstract Web service has developed the managed IoT application to let connected devices easily and securely interact with cloud applications and other devices. As an important factor for web service, the reliability of web services refers to the probability of web service running success. For modeling web service composition, we should abstract the process of web service composition. Due to the diversity and complexity of web service composition, it is unlikely to do exhaustive testing. In order to improve the quality of web service composition test cases and find out which path leads to the greatest probability of service combination failure, heuristic test case generation method is adopted to obtain the optimal test path. First, the web service composition test is abstracted into the MDP model. The QoS of the web service composition is taken as the software test optimization goal, and the cross‐entropy strategy is used to optimize the test case. The experimental results show that the test profile given by the cross‐strategy is better than the random test strategy. Detect and exclude the same number of software defects. Cross‐entropy strategy can significantly reduce the number of test cases, reduce test costs, and improve defect detection efficiency. Yunzhan Gong |
Comput. Intell. | 2 |
| 2020 | Automated defect identification via path analysis-based features with transfer learningabstractRecently, artificial intelligence techniques have been widely applied to address various specialized tasks in software engineering, such as code generation, defect identification, and bug repair. Despite the diffuse usage of static analysis tools in automatically detecting potential software defects, developers consider the large number of reported alarms and the expensive cost of manual inspection to be a key barrier to using them in practice. To automate the process of defect identification, researchers utilize machine learning algorithms with a set of hand-engineered features to build classification models for identifying alarms as actionable or unactionable. However, traditional features often fail to represent the deep syntactic structure of alarms. To bridge the gap between programs’ syntactic structure and defect identification features, this paper first extracts a set of novel fine-grained features at variable-level, called path-variable characteristic, by applying path analysis techniques in the feature extraction process. We then raise a two-stage transfer learning approach based on our proposed features, called feature ranking-matching based transfer learning, to increase the performance of cross-project defect identification. Our experimental results for eight open-source projects show that the proposed features at variable-level are promising and can yield significant improvement on both within-project and cross-project defect identification. Yuwei Zhang 0003, Dahai Jin, Yunzhan Gong |
J. Syst. Softw. | 4 |
| 2020 | A variable-level automated defect identification model based on machine learningabstractStatic analysis tools, automatically detecting potential source code defects at an early phase during the software development process, are diffusely applied in safety-critical software fields. However, alarms reported by the tools need to be inspected manually by developers, which is inevitable and costly, whereas a large proportion of them are found to be false positives. Aiming at automatically classifying the reported alarms into true defects and false positives, we propose a defect identification model based on machine learning. We design a set of novel features at variable level, called variable characteristics, for building the classification model, which is more fine-grained than the existing traditional features. We select 13 base classifiers and two ensemble learning methods for model building based on our proposed approach, and the reported alarms classified as unactionable (false positives) are pruned for the purpose of mitigating the effort of manual inspection. In this paper, we firstly evaluate the approach on four open-source C projects, and the classification results show that the proposed model achieves high performance and reliability in practice. Then, we conduct a baseline experiment to evaluate the effectiveness of our proposed model in contrast to traditional features, indicating that features at variable level improve the performance significantly in defect identification. Additionally, we use machine learning techniques to rank the variable characteristics in order to identify the contribution of each feature to our proposed model. Yuwei Zhang 0003, Yunzhan Gong, Dahai Jin |
Soft Comput. | 3 |
| 2019 | Unit Test Data Generation for C Using Rule-Directed Symbolic Execution
Yunzhan Gong, Dahai Jin |
J. Comput. Sci. Technol. | 2 |
| 2017 | Automated String Constraints Solving for Programs Containing String Manipulation Functions
Xuzhou Zhang, Yunzhan Gong |
J. Comput. Sci. Technol. | 2 |
| 2015 | A hybrid static analysis refinement approach within internetware environmentabstractIn this paper, we propose a hybrid refinement approach to improve the accuracy of static analysis. It keeps condition constraints information during forward dataflow analysis and gets the satisfiability of a warning by a constraint solver taking as input such information and path conditions; data regression analysis can remedy the capability of handling loops and library calls of abstract interpretation technique. It has been implemented in our static analysis tool, Defect Testing System (DTS) and deployed on a internetware environment TRUSTIE. Experiment on a large number of C open source projects shows the great improvement this strategy makes. Dalin Zhang 0003, Gang Yin, Dahai Jin, Yunzhan Gong, Tianshuang Wu, Hailong Zhang 0006 |
Internetware | 4 |
| 2015 | The application of iterative interval arithmetic in path-wise test data generation
Yunzhan Gong, Xuzhou Zhang |
Eng. Appl. Artif. Intell. | 2 |
| 2013 | Null Dereference Detection via a Backward AnalysisabstractNull dereferences are commonly occurring bugs in programming languages such as C. In this paper, we present a novel approach that performs a backward dataflow analysis to detect null-dereference bugs. The technical innovation of our approach is that owing to aliasing predicates, it can perform strong updates in the presence of aliasing, thus eliminating false positives. The aliasing predicates are introduced on the premise of a canonical representation for the program being analyzed. Moreover, the other features of our approach also contribute to improve accuracy. We have implemented this approach, and give an evaluation of it on a set of open source benchmarks. The experimental results prove the effectiveness of our approach, and show that it is suitable for exploring large real programs with reasonable accuracy. Dahai Jin, Yunzhan Gong |
APSEC (1) | 3 |
| 2013 | Diagnosis-Oriented Alarm CorrelationsabstractDefect detection generally includes two stages: static analysis and alarm inspection. Helping the user in the alarm inspection task is a major challenge for current static analyzers. A large number of independent alarms are against the understanding and may lead developers and managers to reject the use of static analysis tools due to the overhead of alarm inspection. To help with the inspection tasks, we formally introduce alarm correlations. If the occurrence of one alarm causes another alarm to occur, we say they are correlated. We propose a framework for the investigation of the alarms, so as to help classifying them by their correlations. The underlying algorithms were implemented inside our static analysis tool. We choose one common semantic alarm as case study and proved that our method has the effect of reducing 33.1% of alarm identification. Using correlation information, we are able to automate alarm identification that previously had to be done manually. Dalin Zhang 0003, Dahai Jin, Yunzhan Gong, Hailong Zhang 0006 |
APSEC (1) | 3 |
| 2011 | STVL: Improve the Precision of Static Defect Detection with Symbolic Three-Valued LogicabstractAmong various abstract domains, the interval domain is simple but also less precise. To improve the precision of static defect detection based on the interval domain, we propose a symbolic three-valued logic (STVL) based interval analysis. Our STVL differs from other symbolic techniques in that it is capable of handling the logical relationship between variables, which could help eliminating false positives. In addition, for the pointer related defect detection, we introduce a STVL-based pointer model, which naturally supports the pointer arithmetic operation, alias analysis and point-to memory abstraction. Moreover, we present a unified symbolic procedure summary model, also STVL-based, to extract the call effect of each invocation and achieve context-sensitivity. Experimental results indicate that the technique is able to achieve sizable precision improvements at reasonable costs, compared with the none-symbolic method. Yunshan Zhao, Yunzhan Gong, Honghe Chen, Zhaohong Yang |
APSEC | 3 |
| 2010 | Defect Analysis Respecting Dead Path Elimination in BPEL ProcessabstractThe rise of web services and service composition in recent years makes it necessary to pay special attentions to their robustness and integrity. One of the most ideal solutions is using testing technology, we are interested in checking whether the process satisfies a given temporal safety property, the cost of dynamic testing for distributed and heterogeneous applications is huge, so forward a static analysis technology applied in the BPEL service composition. Defect-based analysis is a suitable technique to measure the adequacy of test results which can detect the defects pulled in by the developers unintentionally. False positive rate and false negative rate are key evaluation criteria of static defect detecting, so in order to improve the accuracy of detection, the algorithm proposed in this paper for defects analysis involving the semantic analysis, i.e., dead path elimination (DPE). By using the abstract domain of variables in the expression of condition on link, we can identify whether it is the start of DPE or not. The whole paper takes the application of uninitialized variable detection to illustrate the effective of this method. Xuehong Yang, Junfei Huang, Yunzhan Gong |
APSCC | 3 |
| 2008 | DTS - A Software Defects Testing SystemabstractThis demo presents DTS (software defects testing system), a tool to catch defects in source code using static testing techniques. In DTS, various defect patterns are defined using defect patterns state machine and tested by a unified testing framework. Since DTS externalizes all the defect patterns it checks, defect patterns can be added, subtracted, or altered without having to modify the tool itself. Moreover, typical interval computation is expanded and applied in DTS to reduce the false positive and compute the state of defect state machine. In order to validate its usefulness, we perform some experiments on a suite of open source software whose results are briefly presented in the last part of the demo. Zhaohong Yang, Yunzhan Gong |
SCAM | 2 |
| 2005 | A State Machine for Detecting C/C++ Memory FaultsabstractMemory faults are major forms of software bugs that severely threaten system availability and security in C/C++ program. Many tools and techniques are available to check memory faults, but few provide systematic full-scale research and quantitative analysis. Furthermore, most of them produce high noise ratio of warning messages that require many human hours to review and eliminate false-positive alarms. And thus, they cannot locate the root causes of memory faults precisely. This paper provides an innovative state machine to check memory faults, which has three main contributions. Firstly, five concise formulas describing memory faults are given to make the mechanism of the state machine simple and flexible. Secondly, the state machine has the ability to locate the cause roots of the memory faults. Finally, a case study applying to an embedded software, which is written in 50 thousand lines of C codes, shows it can provide useful data to evaluate the reliability and quality of software Guangyan Huang, Guangmei Zhang, Xiaowei Li 0001, Yunzhan Gong |
Asian Test Symposium | 4 |
| 2003 | An Expression's Single Fault Model and the Testing MethodsabstractThis paper proposes a single fault model for the faults of the expressions, including operator faults (operator reference fault: an operator is replaced by another, extra or missing operator for single operand), incorrect variable or constant, incorrect parentheses. These types of faults often exist in the software, but some fault classes are hard to detect using traditional testing methods. A general testing method is proposed to detect these types of faults. Furthermore the fault simulation method of the faults is presented which can accelerate the generation of test cases and minimize the testing cost greatly. Our empirical results indicate that our methods require a smaller number of test cases than random testing, while retaining fault-detection capabilities that are as good as, or better than the traditional testing methods. Yunzhan Gong, Wanli Xu, Xiaowei Li 0001 |
Asian Test Symposium | 1 |
| 2003 | An Object-Oriented Program Automatic Execute Model and the Research of AlgorithmabstractThis paper describes an automatic execution model, which can be used in the automatic test of OO programs. By integrating the object transition diagram, state transition diagram, state transition driver and script chooser, this model can choose and execute script automatically, and whenever it meets any exceptional fault, the result inspector indicates the position of it. By comparing and analyzing several script chooser algorithms, an appropriate one to match this model is designed. Dahai Jin, Yunzhan Gong |
Asian Test Symposium | 2 |
| 1993 | Deductive fault simulation algorithm based on fault collapsing
Yunzhan Gong, Daozheng Wei |
J. Comput. Sci. Technol. | 1 |