Shuo Feng 0003

dblp:00/6921-3 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-1575-9891ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 9 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A survey of unit test case generation methods
Ningwei Zhang, Yizhe Luo, Shuo Feng 0003
Inf. Softw. Technol.6
2025 CMS-YOLO for Small-Scale UAV Detection
abstract
In long-range small UAV detection, challenges such as small target size, low contrast, and insufficient information make the detection task highly difficult. Particularly when the algorithm is deployed in edge environments for real-time processing, extensive optimization is required to balance accuracy and speed under limited computational resources. To address these issues, this paper proposes a novel detection method-Coordinated Attention and Multi-Scale Fusion for Small UAV Detection YOLO (CMS-YOLO). CMS-YOLO incorporates two innovative lightweight modules: the Collaborative Attention Module (CAM) and the Multi-Scale Adaptive Context Fusion Module (MACF). CAM significantly enhances the model's ability to perceive key features by integrating channel attention, spatial attention, and pixel attention, while MACF dynamically adjusts the weights of multi-scale features via an adaptive mechanism, effectively improving small target detection accuracy. In addition, we removed the low-resolution detection head and introduced a high-resolution detection head specifically for small target detection, further optimizing the network structure. Experimental validation on the public small target detection dataset AI-TOD and the self-built long-range small UAV dataset S-Drones shows that CMS-YOLO achieves significant performance improvements while reducing the model size by 20.3%. On the AI-TOD dataset, [email protected] increased by 3.1%, and on the S-Drones dataset, [email protected] increased by 7.1%, significantly outperforming multiple benchmark models and state-of-the-art small target detection methods.
Jun Yan Zhu, Bo Wen Yang, Yizhe Luo, Shuo Feng 0003, Mingliang Xu 0001
ICMR5
2025 Proportional clustering-based undersampling for imbalanced data classification
Chengshuo Zhang, Zhanrong Shi, Wangwei Lu, Shuo Feng 0003
Knowl. Inf. Syst.5
2024 Improving the undersampling technique by optimizing the termination condition for software defect prediction
abstract
The class imbalance problem significantly hinders the ability of the software defect prediction (SDP) models to distinguish between defective (minority class) and non-defective (majority class) software instances. Recent studies on the data resampling technique have shown that Random UnderSampling (RUS) is more effective than several complex oversampling techniques at alleviating this problem. However, RUS blindly removes majority class instances, leading to significant information loss. These studies have also pointed out that the conventional termination condition (i.e., terminating the data resampling technique when the number of instances for both the minority and majority classes are the same) of the data resampling technique can result in suboptimal performance. In fact, the undersampling technique can be likened to a recommender system or a web search engine that recommends majority class instances to SDP models. Therefore, we propose the Learning-To-Rank Undersampling technique (LTRUS). Our work is novel in two aspects: (1) We consider the undersampling process as a learning-to-rank task, optimizing a linear model to rank majority class instances and remove them from the bottom of the rank to alleviate the class imbalance problem . (2) We propose two termination conditions for the undersampling technique, which differ from the conventional termination condition. LTRUS significantly outperforms RUS, the clustering-based undersampling technique, the complexity-based oversampling technique, SMOTUNED, and Borderline-SMOTE in terms of F-measure, AUC, and MCC by 8.9%, 7.6%, and 18.0% on average under the conventional termination condition. Furthermore, LTRUS under the two termination conditions we propose yield similar performance, and both outperform LTRUS and all the other baselines under the conventional termination condition. The experimental results demonstrate the effectiveness of LTRUS and indicate that the conventional termination condition for the data resampling technique is improper.
Shuo Feng 0003, Jacky W. Keung, Yan Xiao 0002, Peichang Zhang, Xiao Yu 0008, Xiaochun Cao
Expert Syst. Appl.1
2023 IFCM: An improved Fuzzy C-means clustering method to handle Class Overlap on Aging-related Software Bug Prediction
abstract
Software aging refers to a problem of performance decay in long-running software systems. This phenomenon is primarily attributed to the accumulation of run-time errors, commonly known as aging-related bugs (ARBs). Detecting ARBs through Aging-related Bug Prediction (ARBP) is crucial in ensuring system reliability. The effectiveness of ARBP heavily relies on the quality of datasets. However, ARB datasets often suffer from class overlap, where instances from different classes exhibit similar feature values. Class overlap poses a significant challenge as it compromises the quality of training data and subsequently impacts ARBP accuracy. To address this issue, we propose an improved Fuzzy C-means clustering method named IFCM, designed to mitigate class overlap in ARBP tasks. IFCM can identify whether an instance occurs overlap, and identify the overlap degree of this instance through the predefined parameters. We evaluate our proposed method on two public datasets Linux and MySQL and one self-collected dataset NetBSD using five different classifiers with five performance metrics (AUC, F1, Balance, PD, PF). Comparison with four existing methods (No clean, NCL, IKMCCA, ROCT) demonstrates that IFCM is effective in alleviating class overlap in ARBP. For Instance, IFCM achieves promising results in terms of AUC blue (which are 0.762, 0.757, and 0.642) and Balance (which are 0.709, 0.736, and 0.595) at the dataset level.
Shuo Feng 0003, Wenzhi Xie, Dongdong Zhao 0001, Jianwen Xiang, Roberto Pietrantuono, Roberto Natella, Domenico Cotroneo
ISSRE2
2023 On the relative value of imbalanced learning for code smell detection
abstract
Summary Machine learning‐based code smell detection (CSD) has been demonstrated to be a valuable approach for improving software quality and enabling developers to identify problematic patterns in code. However, previous researches have shown that the code smell datasets commonly used to train these models are heavily imbalanced. While some recent studies have explored the use of imbalanced learning techniques for CSD, they have only evaluated a limited number of techniques and thus their conclusions about the most effective methods may be biased and inconclusive. To thoroughly evaluate the effect of imbalanced learning techniques for machine learning‐based CSD, we examine 31 imbalanced learning techniques with seven classifiers to build CSD models on four code smell data sets. We employ four evaluation metrics to assess the detection performance with the Wilcoxon signed‐rank test and Cliff's . The results show that (1) Not all imbalanced learning techniques significantly improve detection performance, but deep forest significantly outperforms the other techniques on all code smell data sets. (2) SMOTE (Synthetic Minority Over‐sampling TEchnique) is not the most effective technique for resampling code smell data sets. (3) The best‐performing imbalanced learning techniques and the top‐3 data resampling techniques have little time cost for code smell detection. Therefore, we provide some practical guidelines. First, researchers and practitioners should select the appropriate imbalanced learning techniques (e.g., deep forest) to ameliorate the class imbalance problem. In contrast, the blind application of imbalanced learning techniques could be harmful. Then, better data resampling techniques than SMOTE should be selected to preprocess the code smell data sets.
Kuan Zou, Jacky W. Keung, Xiao Yu 0008, Shuo Feng 0003, Yan Xiao 0002
Softw. Pract. Exp.5
2022 The impact of the distance metric and measure on SMOTE-based techniques in software defect prediction
Shuo Feng 0003, Jacky W. Keung, Peichang Zhang, Yan Xiao 0002, Miao Zhang 0025
Inf. Softw. Technol.1
2022 Predicting the precise number of software defects: Are we there yet?
abstract
Context: Defect Number Prediction (DNP) models can offer more benefits than classification-based defect prediction . Recently, many researchers proposed to employ regression algorithms for DNP, and found that the algorithms achieve low Average Absolute Error (AAE) and high Pred(0.3) values. However, since the defect datasets generally contain many non-defective modules, even if a DNP model predicts the number of defects in all modules as zero, the AAE value of the model will be low and Pred(0.3) value will be high. Therefore, the good performance of the regression algorithms in terms of AAE and Pred(0.3) may be questioned due to the imbalanced distribution of the number of defects. Objective: To revisit the impact of regression algorithms for predicting the precise number of defects. Method: We examine the practical effects of 12 widely-used regression algorithms, two data resampling algorithm (SmoteR and ROS), and three ensemble learning algorithms (gradient boosting regression, AdaBoost .R2, and Bagging), one feature selection method (information gain) and one parameter optimization method (grid search) for predicting the precise number of defects on the 18 PROMISE datasets. We propose to evaluate the AAE and Pred(0.3) values for the modules with different numbers of defects separately. Results: The AAE values for defective modules are very high and the Pred(0.3) values are very low, i.e., the regression algorithms are very inaccurate for predicting the precise number of defects in defective modules. Conclusion: The problem of predicting the precise number of defects via regression algorithms is far from being solved. We recommend that software testers use regression algorithms to rank modules for testing resource allocation , rather than predict the precise number of defects to evaluate the software reliability and maintenance effort. In addition, most existing DNP studies employing the whole AAE and Pred(0.3) values of all modules as the evaluation metrics for the proposed DNP algorithms should be revisited.
Xiao Yu 0008, Jacky W. Keung, Yan Xiao 0002, Shuo Feng 0003, Heng Dai
Inf. Softw. Technol.4
2021 ROCT: Radius-based Class Overlap Cleaning Technique to Alleviate the Class Overlap Problem in Software Defect Prediction
abstract
The training data commonly used in software defect prediction (SDP) usually contains some instances that have similar values on features but are in different classes, which significantly degrades the performance of prediction models trained using these instances. This is referred to as the class overlap problem (COP). Previous studies have concluded that COP has a more negative impact on the performance of prediction models than the class imbalance problem (CIP). However, less research has been conducted on COP than CIP. Moreover, the performance of the existing class overlap cleaning techniques heavily relies on the settings of hyperparameters such as the value of K in the K-nearest neighbor algorithm or the K-means algorithm, but how to find those optimal hyperparameters is still a challenge. In this study, we propose a novel technique named the radius-based class overlap cleaning technique (ROCT) to better alleviate COP without tuning hyperparameters in SDP. The basic idea of ROCT is to take each instance as the center of a hypersphere and directly optimize the radius of the hypersphere. Then ROCT identifies those instances with the opposite label of the center instance as the overlapping instance and removes them. To investigate the performance of ROCT, we conduct the empirical experiment across 29 datasets collected from various software repositories on the K-nearest neighbor, random forest, logistic regression, and naive Bayes classifiers measured by AUC, balance, pd, and pf. The experimental results show that ROCT performs the best and significantly improves the performance of prediction models by as much as 15.2% and 29.9% in terms of AUC and balance compared with the existing class overlap cleaning techniques. The superior performance of ROCT indicates that ROCT should be recommended as an efficient alternative to alleviate COP in SDP.
Shuo Feng 0003, Jacky W. Keung, Jie Liu 0016, Yan Xiao 0002, Xiao Yu 0008, Miao Zhang 0025
COMPSAC1
2021 COSTE: Complexity-based OverSampling TEchnique to alleviate the class imbalance problem in software defect prediction
Shuo Feng 0003, Jacky W. Keung, Xiao Yu 0008, Yan Xiao 0002, Kwabena Ebo Bennin, Md. Alamgir Kabir, Miao Zhang 0025
Inf. Softw. Technol.1
2021 Investigation on the stability of SMOTE-based oversampling techniques in software defect prediction
Shuo Feng 0003, Jacky W. Keung, Xiao Yu 0008, Yan Xiao 0002, Miao Zhang 0025
Inf. Softw. Technol.1
2019 A Heuristic Approach to Break Cycles for the Class Integration Test Order Generation
abstract
It is a general objective to minimize overall stubbing cost when performing class integration test order generation. Existing approaches are unable to obtain a cost-optimal class test order, this is largely due to the lack of a comprehensive analysis on the factors that affect overall stubbing cost, i.e., the number of required test stubs and the corresponding stubbing complexity. To address this issue, we propose an approach called HBCITO (Heuristic approach to Break Cycles for the class Integration Test Order generation). Given a set of removed dependencies, a heuristic algorithm is employed to search for a near ideal set of class dependencies. Such dependencies break the same or greater number of cycles as the initialized dependencies but attract less stubbing cost. The experimental results show that HBCITO is capable of generating class test orders with significantly lower stubbing cost compared with other approaches.
Miao Zhang 0025, Jacky W. Keung, Yan Xiao 0002, Md. Alamgir Kabir, Shuo Feng 0003
COMPSAC (1)5