VLDB 2026 Research / reviewers in the wild / expert
Dahai Jin
dblp:124/2898
· DBLP profile ↗
11ranked-venue papers
1as first author
4since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 3 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Semantic Clone Detection Based on Code Feature Fusion LearningabstractCode clones are duplicated code snippets that significantly threaten software maintenance and the public corpora of code representation learning. Traditionally, code context and its structure information abstract syntax tree (AST), control flow graph (CFG) are typical representations of source code, and context-based models and structure-based models contributed significantly to the development of code clone detection. In this paper, we present a hybrid embedding model for code clone detection (HEM-CCD), a fusion method of token sequential information and graph-based structure information. We insert tokens’ global context information encoded by a bi-directional recurrent neural network into the AST-based graph for comprehensive code semantic representation. Then, feeding the graph into a gated graph neural network we generate code semantic vectors for similarity evaluation. We have implemented our model on two public clone datasets (BigCloneBench and GoogleCodeJam), and the results indicate that HEM-CCD outperforms several state-of-the-art approaches. Dahai Jin, Yunzhan Gong |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2022 | An improving approach to analyzing change impact of C programs
Peng Dai 0007, Dahai Jin, Yunzhan Gong |
Comput. Commun. | 3 |
| 2022 | Improving Large-Gap Clone Detection Recall Using Multiple FeaturesabstractCode clone refers to two or more identical or similar source code fragments. Research on code clone detection has lasted for decades. Investigation and evaluation of existing clone detection techniques indicate that they are resilient to function-level clone detection. Still, there may be room for further research in block-level clone detection. Particularly, type-3 clones that include large gaps, are ongoing challenges. To solve these problems, we propose a clone detection method based on multiple code features. It aims to improve the recall rate of code block clone detection and overcome large-gap and hard-to-detect type-3 clones. This method first splits the source code files based on the program’s structural features and context features to obtain code blocks. The collection of code blocks obtained in this way is complete, and the large gaps in clone pairs will also be removed. In addition, we only need to compute the similarity between code blocks with the same structural features, which can also significantly save time and resources. The similarity is obtained by calculating the proportion of the same tokens between two code blocks. Moreover, since different types of tokens have different weights in similarity calculation, we use supervised learning to obtain a classifier model between token features and code clone. We divide the tokens into 13 types and train the machine learning model with the manually confirmed clone or non-clone pair. Finally, we develop a prototype system and compare our tools with existing tools under the Mutation Framework and in several actual C projects. The experimental results also demonstrate the advancement and practicality of our prototype. Peng Dai 0007, Dahai Jin, Yunzhan Gong |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2022 | ST-TLF: Cross-version defect prediction framework based transfer learningabstractCross-version defect prediction (CVDP) is a practical scenario in which defect prediction models are derived from defect data of historical versions to predict potential defects in the current version. Prior research employed defect data of the latest historical version as the training set using the empirical recommended method, ignoring the concept drift between versions, which undermines the accuracy of CVDP. We customized a Selected Training set and Transfer Learning Framework (ST-TLF) with two objectives: a) to obtain the best training set for the version at hand, proposing an approach to select the training set from the historical data; b) to eliminate the concept drift, designing a transfer strategy for CVDP. To evaluate the performance of ST-TLF, we investigated three research problems, covering the generalization of ST-TLF for multiple classifiers, the accuracy of our training set matching methods, and the performance of ST-TLF in CVDP compared against state-of-the-art approaches. The results reflect that (a) the eight classifiers we examined are all boosted under our ST-TLF, where SVM improves 49.74% considering MCC, as is similar to others; (b) when performing the best training set matching, the accuracy of the method proposed by us is 82.4%, while the experience recommended method is only 41.2%; (c) comparing the 12 control methods, our ST-TLF (with BayesNet), against the best contrast method P15-NB, improves the average MCC by 18.84%. Our framework ST-TLF with various classifiers can work well in CVDP. The training set selection method we proposed can effectively match the best training set for the current version, breaking through the limitation of relying on experience recommendation, which has been ignored in other studies. Also, ST-TLF can efficiently elevate the CVDP performance compared with random forest and 12 control methods. Yanyang Zhao, Yuwei Zhang 0003, Dalin Zhang 0003, Yunzhan Gong, Dahai Jin |
Inf. Softw. Technol. | 6 |
| 2020 | Automated defect identification via path analysis-based features with transfer learningabstractRecently, artificial intelligence techniques have been widely applied to address various specialized tasks in software engineering, such as code generation, defect identification, and bug repair. Despite the diffuse usage of static analysis tools in automatically detecting potential software defects, developers consider the large number of reported alarms and the expensive cost of manual inspection to be a key barrier to using them in practice. To automate the process of defect identification, researchers utilize machine learning algorithms with a set of hand-engineered features to build classification models for identifying alarms as actionable or unactionable. However, traditional features often fail to represent the deep syntactic structure of alarms. To bridge the gap between programs’ syntactic structure and defect identification features, this paper first extracts a set of novel fine-grained features at variable-level, called path-variable characteristic, by applying path analysis techniques in the feature extraction process. We then raise a two-stage transfer learning approach based on our proposed features, called feature ranking-matching based transfer learning, to increase the performance of cross-project defect identification. Our experimental results for eight open-source projects show that the proposed features at variable-level are promising and can yield significant improvement on both within-project and cross-project defect identification. Yuwei Zhang 0003, Dahai Jin, Yunzhan Gong |
J. Syst. Softw. | 2 |
| 2020 | A variable-level automated defect identification model based on machine learningabstractStatic analysis tools, automatically detecting potential source code defects at an early phase during the software development process, are diffusely applied in safety-critical software fields. However, alarms reported by the tools need to be inspected manually by developers, which is inevitable and costly, whereas a large proportion of them are found to be false positives. Aiming at automatically classifying the reported alarms into true defects and false positives, we propose a defect identification model based on machine learning. We design a set of novel features at variable level, called variable characteristics, for building the classification model, which is more fine-grained than the existing traditional features. We select 13 base classifiers and two ensemble learning methods for model building based on our proposed approach, and the reported alarms classified as unactionable (false positives) are pruned for the purpose of mitigating the effort of manual inspection. In this paper, we firstly evaluate the approach on four open-source C projects, and the classification results show that the proposed model achieves high performance and reliability in practice. Then, we conduct a baseline experiment to evaluate the effectiveness of our proposed model in contrast to traditional features, indicating that features at variable level improve the performance significantly in defect identification. Additionally, we use machine learning techniques to rank the variable characteristics in order to identify the contribution of each feature to our proposed model. Yuwei Zhang 0003, Yunzhan Gong, Dahai Jin |
Soft Comput. | 4 |
| 2019 | Unit Test Data Generation for C Using Rule-Directed Symbolic Execution
Yunzhan Gong, Dahai Jin |
J. Comput. Sci. Technol. | 4 |
| 2015 | A hybrid static analysis refinement approach within internetware environmentabstractIn this paper, we propose a hybrid refinement approach to improve the accuracy of static analysis. It keeps condition constraints information during forward dataflow analysis and gets the satisfiability of a warning by a constraint solver taking as input such information and path conditions; data regression analysis can remedy the capability of handling loops and library calls of abstract interpretation technique. It has been implemented in our static analysis tool, Defect Testing System (DTS) and deployed on a internetware environment TRUSTIE. Experiment on a large number of C open source projects shows the great improvement this strategy makes. Dalin Zhang 0003, Gang Yin, Dahai Jin, Yunzhan Gong, Tianshuang Wu, Hailong Zhang 0006 |
Internetware | 3 |
| 2013 | Null Dereference Detection via a Backward AnalysisabstractNull dereferences are commonly occurring bugs in programming languages such as C. In this paper, we present a novel approach that performs a backward dataflow analysis to detect null-dereference bugs. The technical innovation of our approach is that owing to aliasing predicates, it can perform strong updates in the presence of aliasing, thus eliminating false positives. The aliasing predicates are introduced on the premise of a canonical representation for the program being analyzed. Moreover, the other features of our approach also contribute to improve accuracy. We have implemented this approach, and give an evaluation of it on a set of open source benchmarks. The experimental results prove the effectiveness of our approach, and show that it is suitable for exploring large real programs with reasonable accuracy. Dahai Jin, Yunzhan Gong |
APSEC (1) | 2 |
| 2013 | Diagnosis-Oriented Alarm CorrelationsabstractDefect detection generally includes two stages: static analysis and alarm inspection. Helping the user in the alarm inspection task is a major challenge for current static analyzers. A large number of independent alarms are against the understanding and may lead developers and managers to reject the use of static analysis tools due to the overhead of alarm inspection. To help with the inspection tasks, we formally introduce alarm correlations. If the occurrence of one alarm causes another alarm to occur, we say they are correlated. We propose a framework for the investigation of the alarms, so as to help classifying them by their correlations. The underlying algorithms were implemented inside our static analysis tool. We choose one common semantic alarm as case study and proved that our method has the effect of reducing 33.1% of alarm identification. Using correlation information, we are able to automate alarm identification that previously had to be done manually. Dalin Zhang 0003, Dahai Jin, Yunzhan Gong, Hailong Zhang 0006 |
APSEC (1) | 2 |
| 2003 | An Object-Oriented Program Automatic Execute Model and the Research of AlgorithmabstractThis paper describes an automatic execution model, which can be used in the automatic test of OO programs. By integrating the object transition diagram, state transition diagram, state transition driver and script chooser, this model can choose and execute script automatically, and whenever it meets any exceptional fault, the result inspector indicates the position of it. By comparing and analyzing several script chooser algorithms, an appropriate one to match this model is designed. Dahai Jin, Yunzhan Gong |
Asian Test Symposium | 1 |