VLDB 2026 Research / reviewers in the wild / expert
Xiaoxu Niu
dblp:171/3105
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0003-0621-8767ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Software defect prediction based on graph code semantics
Hongwei Tao, Zhenhao Geng, Xiaoxu Niu, Qiaoling Cao |
Expert Syst. Appl. | 4 |
| 2025 | Prediction of incompatible bug numbers between versions of java open-source software based on deep fusion features
Xiaoxu Niu, Hongwei Tao, Qiaoling Cao, Jianxun Wang 0010, Zhenhao Geng |
Expert Syst. Appl. | 1 |
| 2025 | Software aging oriented trustworthiness measurement based on weighted Boltzmann entropy
Hongwei Tao, Han Liu 0012, Xiaoxu Niu, Licheng Ding, Yixiang Chen 0001, Qiaoling Cao |
Inf. Softw. Technol. | 3 |
| 2024 | Software Defect Prediction Method Based on Clustering Ensemble LearningabstractThe technique of software defect prediction aims to assess and predict potential defects in software projects and has made significant progress in recent years within software development. In previous studies, this technique largely relied on supervised learning methods, requiring a substantial amount of labeled historical defect data to train the models. However, obtaining these labeled data often demands significant time and resources. In contrast, software defect prediction based on unsupervised learning does not depend on known labeled data, eliminating the need for large‐scale data labeling, thereby saving considerable time and resources while providing a more flexible solution for ensuring software quality. This paper conducts software defect prediction using unsupervised learning methods on data from 16 projects across two public datasets (PROMISE and NASA). During the feature selection step, a chi‐squared sparse feature selection method is proposed. This feature selection strategy combines chi‐squared tests with sparse principal component analysis (SPCA). Specifically, the chi‐squared test is first used to filter out the most statistically significant features, and then the SPCA is applied to reduce the dimensionality of these significant features. In the clustering step, the dot product matrix and Pearson correlation coefficient (PCC) matrix are used to construct weighted adjacency matrices, and a clustering overlap method is proposed. This method integrates spectral clustering, Newman clustering, fluid clustering, and Clauset–Newman–Moore (CNM) clustering through ensemble learning. Experimental results indicate that, in the absence of labeled data, using the chi‐squared sparse method for feature selection demonstrates superior performance, and the proposed clustering overlap method outperforms or is comparable to the effectiveness of the four baseline clustering methods. Hongwei Tao, Qiaoling Cao, Xiaoxu Niu, Zhenhao Geng, Songtao Shang |
IET Softw. | 5 |
| 2024 | Cross-Project Defect Prediction Using Transfer Learning with Long Short-Term Memory NetworksabstractWith the increasing number of software projects, within‐project defect prediction (WPDP) has already been unable to meet the demand, and cross‐project defect prediction (CPDP) is playing an increasingly significant role in the area of software engineering. The classic CPDP methods mainly concentrated on applying metric features to predict defects. However, these approaches failed to consider the rich semantic information, which usually contains the relationship between software defects and context. Since traditional methods are unable to exploit this characteristic, their performance is often unsatisfactory. In this paper, a transfer long short‐term memory (TLSTM) network model is first proposed. Transfer semantic features are extracted by adding a transfer learning algorithm to the long short‐term memory (LSTM) network. Then, the traditional metric features and semantic features are combined for CPDP. First, the abstract syntax trees (AST) are generated based on the source codes. Second, the AST node contents are converted into integer vectors as inputs to the TLSTM model. Then, the semantic features of the program can be extracted by TLSTM. On the other hand, transferable metric features are extracted by transfer component analysis (TCA). Finally, the semantic features and metric features are combined and input into the logical regression (LR) classifier for training. The presented TLSTM model performs better on the f ‐measure indicator than other machine and deep learning models, according to the outcomes of several open‐source projects of the PROMISE repository. The TLSTM model built with a single feature achieves 0.7% and 2.1% improvement on Log4j‐1.2 and Xalan‐2.7, respectively. When using combined features to train the prediction model, we call this model a transfer long short‐term memory for defect prediction (DPTLSTM). DPTLSTM achieves a 2.9% and 5% improvement on Synapse‐1.2 and Xerces‐1.4.4, respectively. Both prove the superiority of the proposed model on the CPDP task. This is because LSTM capture long‐term dependencies in sequence data and extract features that contain source code structure and context information. It can be concluded that: (1) the TLSTM model has the advantage of preserving information, which can better retain the semantic features related to software defects; (2) compared with the CPDP model trained with traditional metric features, the performance of the model can validly enhance by combining semantic features and metric features. Hongwei Tao, Lianyou Fu, Qiaoling Cao, Xiaoxu Niu, Songtao Shang, Yang Xian |
IET Softw. | 4 |
| 2024 | A comparative study of software defect binomial classification prediction models based on machine learning
Hongwei Tao, Xiaoxu Niu, Lianyou Fu, Qiaoling Cao, Songtao Shang, Yang Xian |
Softw. Qual. J. | 2 |
| 2022 | A comprehensive comparison among metaheuristics (MHs) for geohazard modeling using machine learning: Insights from a case study of landslide displacement predictionabstractMachine learning (ML) has been extensively applied to model geohazards, yielding tremendous success. However, researchers and practitioners still face challenges in enhancing the reliability of ML models. In the present study, a systematic framework combining k-fold cross-validation (CV), metaheuristics (MHs), support vector regression (SVR), and Friedman and Nemenyi tests was proposed to improve the reliability and performance of geohazard modeling. The average normalized mean square error (NMSE) from k-fold CV sets was adopted as the fitness metric. Twenty of the most well-established MHs and the most recent MHs were adopted to tune the hyperparameters of SVR and were evaluated through nonparametric Friedman and post hoc Nemenyi tests to identify significant differences. Observations from a typical reservoir landslide were selected as a benchmark dataset, and the accuracy, robustness, computational time, and convergence speed of the MHs were compared. Significant performance differences among the twenty MHs were identified by Friedman and post hoc Nemenyi tests of the mean absolute error (MAE), root mean squared error (RMSE), Kling–Gupta efficiency (KGE), and computational time, with p values lower than 0.05. The comparison of results demonstrated that the multiverse optimizer (MVO) is among the highest-performing, most stable, and computationally efficient algorithms, providing superior performance to other methods, with nearly optimum values of the correlation coefficient (R), a low MAE (23.5086 versus 23.9360), a low mean RMSE (48.6946 versus 50.1882), and a high mean KGE (0.9803 versus 0.9893) in predicting the displacement of the Shuping landslide. This paper considerably enriches the literature regarding hyperparameter optimization algorithms and the enhancement of their reliability. In addition, Friedman and post hoc Nemenyi tests have the potential for evaluating and comparing various ML-based geohazard models. Junwei Ma, Ding Xia, Xiaoxu Niu, Haixiang Guo |
Eng. Appl. Artif. Intell. | 4 |