Xingguang Yang

dblp:224/4876 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-7489-2265ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2025 RESEARCH NOTES: Multiclass Classification for Self-Admitted Technical Debt via Large Pre-Trained Language Model
abstract
Technical debt refers to suboptimal solutions adopted for short-term goals. Self-admitted technical debt (SATD) is the debt that is explicitly marked through comments or documentation, making it traceable. Multi-classification of SATD helps developers understand different debt types and improve efficiency. This paper proposes a SATD multi-classification method based on Fine-Tuning the GPT-3.5-turbo model for SATD prediction. This study uses a public dataset containing 10 projects with code comments. We classify design debt, requirement debt, and defect debt and evaluate our method’s performance. The experimental results show that compared to the best baseline model, our method achieves average improvements of 11.41%, 1.72% and 3.72% in MacroF, MacroP and MacroR metrics, respectively, in the MTO scenario. In the OTO scenario, improvements are 2.33%, 3.70% and 2.18%, respectively. These results indicate that our method has a strong generalization ability in SATD multi-classification and offers a new approach to managing technical debt.
Yiyang Du, Xingguang Yang, Zhenyu Shu, Zijie Huang 0001, Gang Wang 0023, Libo Xu
Int. J. Softw. Eng. Knowl. Eng.2
2022 Summarizing source code with hierarchical code representation
Ziyi Zhou 0002, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Xingguang Yang
Inf. Softw. Technol.5
2022 Community Smell Occurrence Prediction on Multi-Granularity by Developer-Oriented Features and Process Metrics
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Huiqun Yu, Xingguang Yang, Kang Yang 0004
J. Comput. Sci. Technol.5
2022 A graph sequence neural architecture for code completion with semantic structure features
abstract
Abstract Code completion plays an important role in intelligent software development for accelerating coding efficiency. Recently, the prediction models based on deep learning have achieved good performance in code completion task. However, the existing models cannot avoid three drawbacks: (i) In the existing models, the code representation loses the information (parent–child information between nodes) and lacks many effective features (orientation between nodes). (ii) The known code structure information is not fully utilized, which will cause the model to generate completely irrelevant results. (iii) Simple sequence modeling ignores repeated patterns and structural information. Besides, previous works cannot capture the characteristics of correlation and directionality between nodes. In this paper, we propose a Code Completion approach named CC‐GGNN, which is graph model based on Gated Graph Neural Networks (GGNNs) to address the problems. We introduce a new architecture to obtain the effective code features from code representation. In order to utilize the known information, we propose Classification Mechanism, which classifies the representation of the node using the known parent node and constructs training graph in the model. The experimental results show that our model outperforms the state‐of‐the‐art methods MRR@5 at most 9.2% and ACC at most 11.4% in datasets.
Kang Yang 0004, Huiqun Yu, Guisheng Fan, Xingguang Yang, Zijie Huang 0001
J. Softw. Evol. Process.4
2021 An Empirical Study of Model-Agnostic Interpretation Technique for Just-in-Time Software Defect Prediction
Xingguang Yang, Huiqun Yu, Guisheng Fan, Zijie Huang 0001, Kang Yang 0004, Ziyi Zhou 0002
CollaborateCom (1)1
2021 Predicting Community Smells' Occurrence on Individual Developers by Sentiments
abstract
Community smells appear in sub-optimal software development community structures, causing unforeseen additional project costs, e.g., lower productivity and more technical debt. Previous studies analyzed and predicted community smells in the granularity of community sub-groups using socio-technical factors. However, refactoring such smells requires the effort of developers individually. To eliminate them, supportive measures for every developer should be constructed according to their motifs and working states. Recent work revealed developers' personalities could influence community smells' variation, and their sentiments could impact productivity. Thus, sentiments could be evaluated to predict community smells' occurrence on them. To this aim, this paper builds a developer-oriented and sentiment-aware community smell prediction model considering 3 smells such as Organizational Silo, Lone Wolf, and Bottleneck. Furthermore, it also predicts if a developer quitted the community after being affected by any smell. The proposed model achieves cross- and within-project prediction F-Measure ranging from 76% to 93%. Research also reveals 6 sentimental features having stronger predictive power compared with activeness metrics. Imperative and indicative expressions, politeness, and several emotions are the most powerful predictors. Finally, we test statistically the mean and distribution of sentimental features. Based on our findings, we suggest developers should communicate in a straightforward and polite way.
Zijie Huang 0001, Zhiqing Shao, Guisheng Fan, Ziyi Zhou 0002, Kang Yang 0004, Xingguang Yang
ICPC7
2021 An Empirical Study on the Impact of Class Overlapin Just-in-Time Software Defect Prediction (S)
abstract
Just-in-time software defect prediction (JIT-SDP) is an active research topic in the field of software engineering, aiming at identifying defect-inducing code changes.Most of the current JIT-SDP work focused on model construction.It is often ignored that the performance of classifiers often depends on high quality data.In this paper, we first investigate the impact of the class overlap problem on the performance of the classifiers in JIT-SDP, and propose a new effective preprocessing method (IKMCCA-TL) combining improved K-Means clustering cleaning approach and Tomek-link method.In order to objectively estimate the impact of class overlap on the classifiers in JIT-SDP, we conduct a large-scale empirical study on the data sets of six open source projects and compare the performance of LR, RF and KNN classifiers by using IKMCCA or KMCCA or NCL and without cleaning data.Experimental results show that after removing overlapping instances, the performance of the classifiers is significantly improved in terms of balance, recall and AUC and our proposed method achieves the best performance.
Minyang Yi, Guisheng Fan, Huiqun Yu, Xingguang Yang
SEKE4
2021 DEJIT: A Differential Evolution Algorithm for Effort-Aware Just-in-Time Software Defect Prediction
abstract
Software defect prediction is an effective approach to save testing resources and improve software quality, which is widely studied in the field of software engineering. The effort-aware just-in-time software defect prediction (JIT-SDP) aims to identify defective software changes in limited software testing resources. Although many methods have been proposed to solve the JIT-SDP, the effort-aware prediction performance of the existing models still needs to be further improved. To this end, we propose a differential evolution (DE) based supervised method DEJIT to build JIT-SDP models. Specifically, first we propose a metric called density-percentile-average (DPA), which is used as optimization objective on the training set. Then, we use logistic regression (LR) to build a prediction model. To make the LR obtain the maximum DPA on the training set, we use the DE algorithm to determine the coefficients of the LR. The experiment uses defect data sets from six open source projects. We compare the proposed method with state-of-the-art four supervised models and four unsupervised models in cross-validation, cross-project-validation and timewise-cross-validation scenarios. The empirical results demonstrate that the DEJIT method can significantly improve the effort-aware prediction performance in the three evaluation scenarios. Therefore, the DEJIT method is promising for the effort-aware JIT-SDP.
Xingguang Yang, Huiqun Yu, Guisheng Fan, Kang Yang 0004
Int. J. Softw. Eng. Knowl. Eng.1
2020 Code Prediction Based on Graph Embedding Model
Kang Yang 0004, Huiqun Yu, Guisheng Fan, Xingguang Yang, Liqiong Chen
CollaborateCom (2)4
2019 An Empirical Studies on Optimal Solutions Selection Strategies for Effort-Aware Just-in-Time Software Defect Prediction
abstract
Just-in-time software defect prediction (JIT-SDP) is an active topic in the filed of software engineering, and many methods have been proposed to solve this problem.Stateof-the-art method MULTI applies multi-objective optimization algorithm to the effort-aware JIT-SDP problem, and obtains good average performance.Although the average performance of the MULTI method is high, there are many optimal solutions with poor performance.If an optimal solution is randomly selected, a poor prediction model may be obtained.In order to further improve the performance of the MULTI method, we propose three optimal solutions selection strategies: benefit priority (BP), cost priority (CP), and a compromise between cost and benefit (CCB).In order to compare and validate the effectiveness of the strategies, we conduct a large-scale empirical study on data sets of six open source projects.The experimental results show that, compared with the average performance of MULTI, the optimal solutions selection strategy based on BP has a significant improvement in ACC and Popt indicators.Therefore, we recommend using the BP-based optimal solutions selection strategy to improve the performance of MULTI when using the MULTI method to solve the effort-aware JIT-SDP problem.
Xingguang Yang, Huiqun Yu, Guisheng Fan, Kang Yang 0004
SEKE1
2019 Mutation with Local Searching and Elite Inheritance Mechanism in Multi-Objective Optimization Algorithm: A Case Study in Software Product Line
abstract
An effective method for addressing the configuration optimization problem (COP) in Software Product Lines (SPLs) is to deploy a multi-objective evolutionary algorithm, for example, the state-of-the-art SATIBEA. In this paper, an improved hybrid algorithm, called SATIBEA-LSSF, is proposed to further improve the algorithm performance of SATIBEA, which is composed of a multi-children generating strategy, an enhanced mutation strategy with local searching and an elite inheritance mechanism. Empirical results on the same case studies demonstrate that our algorithm significantly outperforms the state-of-the-art for four out of five SPLs on a quality Hypervolume indicator and the convergence speed. To verify the effectiveness and robustness of our algorithm, the parameter sensitivity analysis is discussed and three observations are reported in detail.
Kai Shi 0006, Huiqun Yu, Guisheng Fan, Jianmei Guo, Liqiong Chen, Xingguang Yang, Huaiying Sun
Int. J. Softw. Eng. Knowl. Eng.6
2019 A Parallel Framework of Combining Satisfiability Modulo Theory with Indicator-Based Evolutionary Algorithm for Configuring Large and Real Software Product Lines
abstract
Multi-objective evolutionary algorithm (MOEA) has been widely applied to software product lines (SPLs) for addressing the configuration optimization problems. For example, the state-of-the-art SMTIBEA algorithm extends the constraint expressiveness and supports richer constraints to better address these problems. However, it just works better than the competitor for four out of five SPLs in five objectives and the convergence speed is not significantly increased for largest Linux SPL from 5 to 30[Formula: see text]min. To further improve the optimization efficiency, we propose a parallel framework SMTPORT, which combines four corresponding SMTIBEA variants and performs these variants by utilizing parallelization techniques within the limited time budget. For case studies in LVAT repository, we conduct a series of experiments on seven real-world and highly-constrained SPLs. Empirical results demonstrate that our approach significantly outperforms the state-of-the-art for all the seven SPLs in terms of a quality Hypervolume metric and a diversity Pareto Front Size indicator.
Kai Shi 0006, Huiqun Yu, Jianmei Guo, Guisheng Fan, Liqiong Chen, Xingguang Yang
Int. J. Softw. Eng. Knowl. Eng.6
2018 Combining Constraint Solving with Different MOEAs for Configuring Large Software Product Lines: A Case Study
abstract
Multi-objective evolutionary algorithm (MOEA) with the constraint solving has been successfully applied to address the configuration optimization problem in software product line (SPL), for example, the state-of-the-art SATIBEA algorithm. However, each different MOEA with special search operator demonstrates the different strength and weakness in terms of optimality and convergence speed. The SATIBEA just combines the SAT (Boolean satisfiability problem) constraint solving with the Indicator-Based Evolutionary Algorithm (IBEA) for evaluating the algorithm performance. In this paper, we propose six hybrid algorithms which combine the SAT solving with different MOEAs. Case study is based on five large-scale, rich-constrained and real-world SPLs. Empirical results demonstrate that SATMOCell algorithm obtains a competitive optimization performance to the state-of-the-art that outperforms the SATIBEA in terms of quality Hypervolume metric for 2 out of 5 SPLs within the same time budget. Moreover, the convergence speed of SATMOCell and SATssNSGA2 is comparable after 10min terminal times. Particularly, the Hypervolume value of SATssNSGA2 reports the average improvement of 1.33% after 20min terminal times.
Huiqun Yu, Kai Shi 0006, Jianmei Guo, Guisheng Fan, Xingguang Yang, Liqiong Chen
COMPSAC (1)5
2018 imBBO: An Improved Biogeography-Based Optimization Algorithm
Kai Shi 0006, Huiqun Yu, Guisheng Fan, Xingguang Yang
GPC4
2018 A parallel portfolio approach to configuration optimization for large software product lines
abstract
Summary Software product line (SPL) engineering demands for optimal or near‐optimal products that balance multiple often competing and conflicting objectives. A major challenge for large SPLs is to efficiently explore a huge space of various products and satisfy a large number of predefined constraints simultaneously. To improve the optimality and convergence speed, we propose a parallel portfolio approach, called IBEAPORT, which designs three algorithm variants by incorporating constraint solving into the indicator‐based evolutionary algorithm in different ways and performs these variants by utilizing parallelization techniques. Our approach utilizes the exploration capabilities of different algorithms and improves optimality as far as possible within a limited time budget. We evaluate our approach on five large‐scale real‐world SPLs. Empirical results demonstrate that our approach significantly outperforms the state of the art for all five SPLs on a quality indicator and a diversity indicator. Moreover, IBEAPORT quickly converges to a relatively stable hypervolume value even for the largest SPL with 6888 features.
Kai Shi 0006, Huiqun Yu, Jianmei Guo, Guisheng Fan, Xingguang Yang
Softw. Pract. Exp.5