Kang Yang 0004

dblp:86/8501-4 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0002-5863-1003ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A reinforcement learning malware detection model based on heterogeneous information network path representation
Kang Yang 0004, Lizhi Cai
Appl. Intell.1
2024 Bug report priority prediction using social and technical features
abstract
Summary Software stakeholders report bugs in issue tracking system (ITS) with manually labeled priorities. However, the lack of knowledge and standard for prioritization may cause stakeholders to mislabel the priorities. In response, priority predictors are actively developed to support them. Prior studies trained machine learners based on textual similarity, categorical, and numeric technical features of bug reports. Most models were validated by time‐insensitive approaches, and they were producing suboptimal results for practical usage. While they ignored the social aspects of ITS, the technical aspects were also limited in surface features of bug reports. To better model the bug report, we extract their topic and most similar code structures. Since ITS bridges users and developers as the main contributors, we also integrate their experience, sentiment, and socio‐technical features to construct a new dataset. Then, we perform two‐classed and multiclassed bug priority prediction based on the dataset. We also introduce adversarial training using generated training data with random word swap and random word deletion. We validate our model in within‐project, cross‐project, and time‐wise scenarios, and it outperforms the two baselines by up to 15% in area under curve‐receiver operating characteristics (AUC‐ROC) and 19% in Matthews correlation coefficient (MCC). We reveal involving contributor (i.e., assignee and reporter) features such as sentiment that could boost prediction performance. Finally, we test statistically the mean and distribution of the features that reflect the differences in social and technical aspects (e.g., quality of communication and resource distribution) between high and low priority reports. In conclusion, we suggest that researchers should consider both social and technical aspects of ITS in bug report priority prediction and introduce adversarial training to boost model performance.
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Huiqun Yu, Kang Yang 0004, Ziyi Zhou 0002
J. Softw. Evol. Process.5
2023 Towards Retrieval-Based Neural Code Summarization: A Meta-Learning Approach
abstract
Code summarization aims to generate code summaries automatically, and has attracted a lot of research interest lately. Recent approaches to it commonly adopt neural machine translation techniques, which train a Seq2Seq model on a large corpus and assume it could work on various new code snippets. However, codes are highly varied in practice due to different domains, businesses or programming styles. Therefore, it is challenging to learn such a variety of patterns into a single model. In this paper, we propose a brand-new framework for code summarization based on meta-learning and code retrieval, named MLCS to tackle this issue. In this framework, the summarization of each target code is formalized as a few-shot learning task, where its similar examples are used as training data and the testing example is itself. We retrieve examples similar to the target code in a rank-and-filter manner. Given a neural code summarizer, we optimize it into a meta-learner via Model-Agnostic Meta-Learning (MAML). During inference, the meta-learner first adapts to the retrieved examples and yields an exclusive model for the target code, and then generates its summary. Extensive experiments on real-world datasets show: (1) Utilizing MLCS, a standard Seq2Seq model is able to outperform previous state-of-the-art approaches, including both neural models and retrieval-based neural models; (2) MLCS can flexibly adapt to existing neural code summarizers without modifying their architecture, and could significantly improve their performance with the relative gain of up to 112.7% on BLEU-4, 23.2% on ROUGE-L, and 31.5% on METEOR; (3) Compared to the existing retrieval-based neural approaches, MLCS can better leverage multiple similar examples, and shows better generalization ability on different retrievers, unseen retrieval corpus and low-frequency words.
Ziyi Zhou 0002, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Kang Yang 0004
IEEE Trans. Software Eng.5
2022 Bug Report Priority Prediction Using Developer-Oriented Socio-Technical Features
abstract
Software stakeholders report bugs in Issue Tracking System (ITS) with manually labeled priorities. However, the lack of knowledge and standard for prioritization may cause stakeholders to mislabel the priorities. In response, priority predictors are actively developed to support them. Prior studies trained machine learners based on textual similarity, categorical, and numeric technical features of bug reports. Most models were validated by time-insensitive approaches, and they were producing sub-optimal results for practical usage. Moreover, they tend to ignore the developer and social aspects of ITS. Since ITS bridges users and developers, we integrate their sentiment- and community-oriented socio-technical features to perform 2- and multi-classed bug priority prediction and validate our model in within-project, cross-project, and time-wise scenarios. The proposed model outperforms the 2 baselines by up to 10% in AUC-ROC and 13% in MCC, and the significance of improvement is statistically confirmed. We reveal involving assignee and reporter features from socio-technical perspectives such as sentiment could boost prediction performance. Finally, we test statistically the mean and distribution of the features that reflect the differences in socio-technical aspects (e.g., quality of communication and resource distribution) between high and low priority reports. In conclusion, we suggest researchers should involve contributors’ experience and sentiments in bug report priority prediction.
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Huiqun Yu, Kang Yang 0004, Ziyi Zhou 0002
Internetware5
2022 HQLgen: deep learning based HQL query generation from program context
Ziyi Zhou 0002, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Kang Yang 0004, Jiayin Zhang
Autom. Softw. Eng.5
2022 Code Generation with Hybrid of Structural and Semantic Features Retrieval
abstract
Due to the growing need for faster software delivery, code generation has attracted more and more attention, since it could improve code maintainability by providing suggestions for coding. In the model of generating program source code from natural language (NL), the most effective method is to generate an intermediate architecture (such as Abstract Syntax Tree) combined with a deep learning model. However, these models have the following drawbacks: (1) The data structural information is underutilized and the correlation between samples is not considered. (2) Lack of the ability to memorize large and complex structures, so that complex codes cannot be generated correctly. To address these issues, we propose HRCODE model, a code generation architecture based on Hybrid of structural and semantic features Retrieval CODE model. We transform the NL description into an intermediate structure with structural features. Then, the NL and the intermediate structure are embedded into a vector through weight mixing, and we calculate the similarity score between each vector to retrieve the most relevant samples. Finally, the new input is brought into the PLBART model to generate code. Experiments show that HRCODE is at least 4.7% higher than the state-of-the-art models in the ACC metric and at least 10.3% higher in the BLEU-4 score. We have released our code at https://github.com/jesokang/HRCODE.
Kang Yang 0004, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Ziyi Zhou 0002
Int. J. Softw. Eng. Knowl. Eng.1
2022 Community Smell Occurrence Prediction on Multi-Granularity by Developer-Oriented Features and Process Metrics
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Huiqun Yu, Xingguang Yang, Kang Yang 0004
J. Comput. Sci. Technol.6
2022 HBSniff: A static analysis tool for Java Hibernate object-relational mapping code smell detection
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Huiqun Yu, Kang Yang 0004, Ziyi Zhou 0002
Sci. Comput. Program.5
2022 A graph sequence neural architecture for code completion with semantic structure features
abstract
Abstract Code completion plays an important role in intelligent software development for accelerating coding efficiency. Recently, the prediction models based on deep learning have achieved good performance in code completion task. However, the existing models cannot avoid three drawbacks: (i) In the existing models, the code representation loses the information (parent–child information between nodes) and lacks many effective features (orientation between nodes). (ii) The known code structure information is not fully utilized, which will cause the model to generate completely irrelevant results. (iii) Simple sequence modeling ignores repeated patterns and structural information. Besides, previous works cannot capture the characteristics of correlation and directionality between nodes. In this paper, we propose a Code Completion approach named CC‐GGNN, which is graph model based on Gated Graph Neural Networks (GGNNs) to address the problems. We introduce a new architecture to obtain the effective code features from code representation. In order to utilize the known information, we propose Classification Mechanism, which classifies the representation of the node using the known parent node and constructs training graph in the model. The experimental results show that our model outperforms the state‐of‐the‐art methods MRR@5 at most 9.2% and ACC at most 11.4% in datasets.
Kang Yang 0004, Huiqun Yu, Guisheng Fan, Xingguang Yang, Zijie Huang 0001
J. Softw. Evol. Process.1
2021 An Empirical Study of Model-Agnostic Interpretation Technique for Just-in-Time Software Defect Prediction
Xingguang Yang, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Kang Yang 0004, Ziyi Zhou 0002
CollaborateCom (1)5
2021 Predicting Community Smells' Occurrence on Individual Developers by Sentiments
abstract
Community smells appear in sub-optimal software development community structures, causing unforeseen additional project costs, e.g., lower productivity and more technical debt. Previous studies analyzed and predicted community smells in the granularity of community sub-groups using socio-technical factors. However, refactoring such smells requires the effort of developers individually. To eliminate them, supportive measures for every developer should be constructed according to their motifs and working states. Recent work revealed developers' personalities could influence community smells' variation, and their sentiments could impact productivity. Thus, sentiments could be evaluated to predict community smells' occurrence on them. To this aim, this paper builds a developer-oriented and sentiment-aware community smell prediction model considering 3 smells such as Organizational Silo, Lone Wolf, and Bottleneck. Furthermore, it also predicts if a developer quitted the community after being affected by any smell. The proposed model achieves cross- and within-project prediction F-Measure ranging from 76% to 93%. Research also reveals 6 sentimental features having stronger predictive power compared with activeness metrics. Imperative and indicative expressions, politeness, and several emotions are the most powerful predictors. Finally, we test statistically the mean and distribution of sentimental features. Based on our findings, we suggest developers should communicate in a straightforward and polite way.
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Ziyi Zhou 0002, Kang Yang 0004, Xingguang Yang
ICPC6
2021 DEJIT: A Differential Evolution Algorithm for Effort-Aware Just-in-Time Software Defect Prediction
abstract
Software defect prediction is an effective approach to save testing resources and improve software quality, which is widely studied in the field of software engineering. The effort-aware just-in-time software defect prediction (JIT-SDP) aims to identify defective software changes in limited software testing resources. Although many methods have been proposed to solve the JIT-SDP, the effort-aware prediction performance of the existing models still needs to be further improved. To this end, we propose a differential evolution (DE) based supervised method DEJIT to build JIT-SDP models. Specifically, first we propose a metric called density-percentile-average (DPA), which is used as optimization objective on the training set. Then, we use logistic regression (LR) to build a prediction model. To make the LR obtain the maximum DPA on the training set, we use the DE algorithm to determine the coefficients of the LR. The experiment uses defect data sets from six open source projects. We compare the proposed method with state-of-the-art four supervised models and four unsupervised models in cross-validation, cross-project-validation and timewise-cross-validation scenarios. The empirical results demonstrate that the DEJIT method can significantly improve the effort-aware prediction performance in the three evaluation scenarios. Therefore, the DEJIT method is promising for the effort-aware JIT-SDP.
Xingguang Yang, Huiqun Yu, Guisheng Fan, Kang Yang 0004
Int. J. Softw. Eng. Knowl. Eng.4
2020 Code Prediction Based on Graph Embedding Model
Kang Yang 0004, Huiqun Yu, Guisheng Fan, Xingguang Yang, Liqiong Chen
CollaborateCom (2)1
2019 Deep Semantic Feature Learning with Embedded Static Metrics for Software Defect Prediction
abstract
Software defect prediction, which locates defective code snippets, can assist developers in finding potential bugs and assigning their testing efforts. Traditional defect prediction features are static code metrics, which only contain statistic information of programs and fail to capture semantics in programs, leading to the degradation of defect prediction performance. To take full advantage of the semantics and static metrics of programs, we propose a framework called Defect Prediction via Attention Mechanism (DP-AM) in this paper. Specifically, DPAM first extracts vectors which are then encoded as digital vectors by mapping and word embedding from abstract syntax trees (ASTs) of programs. Then it feeds these numerical vectors into Recurrent Neural Network to automatically learn semantic features of programs. After that, it applies self-attention mechanism to further build relationship among these features. Furthermore, it employs global attention mechanism to generate significant features among them. Finally, we combine these semantic features with traditional static metrics for accurate software defect prediction. We evaluate our method in terms of F1-measure on seven open-source Java projects in Apache. Our experimental results show that DP-AM improves F1-measure by 11% in average, compared with the state-of-the-art methods.
Guisheng Fan, Xuyang Diao, Huiqun Yu, Kang Yang 0004, Liqiong Chen
APSEC4
2019 An Empirical Studies on Optimal Solutions Selection Strategies for Effort-Aware Just-in-Time Software Defect Prediction
abstract
Just-in-time software defect prediction (JIT-SDP) is an active topic in the filed of software engineering, and many methods have been proposed to solve this problem.Stateof-the-art method MULTI applies multi-objective optimization algorithm to the effort-aware JIT-SDP problem, and obtains good average performance.Although the average performance of the MULTI method is high, there are many optimal solutions with poor performance.If an optimal solution is randomly selected, a poor prediction model may be obtained.In order to further improve the performance of the MULTI method, we propose three optimal solutions selection strategies: benefit priority (BP), cost priority (CP), and a compromise between cost and benefit (CCB).In order to compare and validate the effectiveness of the strategies, we conduct a large-scale empirical study on data sets of six open source projects.The experimental results show that, compared with the average performance of MULTI, the optimal solutions selection strategy based on BP has a significant improvement in ACC and Popt indicators.Therefore, we recommend using the BP-based optimal solutions selection strategy to improve the performance of MULTI when using the MULTI method to solve the effort-aware JIT-SDP problem.
Xingguang Yang, Huiqun Yu, Guisheng Fan, Kang Yang 0004
SEKE4