Shihai Wang

dblp:78/1611 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 9 · 6 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 A Novel Image Preprocessing Algorithm for Doppler Wind Spectrum Inversion: Principle, Method, and Performance
abstract
We propose a novel Image Preprocessing Algorithm (IPA) for Coherent Doppler Wind Lidar (CDWL) wind spectrum inversion under low Signal-to-Noise Ratio (SNR) conditions. By utilizing prior knowledge such as the spatial continuity of wind spectra, this new algorithm significantly improves the reliability of the peak frequency estimation and enhances the wind inversion accuracy and precision. The theoretical analysis shows that the IPA can improve the accuracy and precision of the wind spectrum inversion under low SNR, through preselecting and weighting the Doppler wind spectra. Verified by simulation and real CDWL measurements, the IPA effectively enhances the accuracy and reliability of lidar wind spectrum inversion under low SNR conditions. Specifically, the performance enhancement is equivalent to an approximate improvement in the SNR of the data by 2.5-5 dB. The measurements with real CDWL measurements show the maximum effective detection range can be increased by more than 2 km. The IPA is low complexity, demanding low computation resources, and can be applied to high temporal resolution Doppler wind estimation.
Mingguang Zhao, Zhibin Yu 0005, Mengpei Li, Shihai Wang
IEEE Trans. Geosci. Remote. Sens.5
2024 A novel software defect prediction approach via weighted classification based on association rule mining
Shihai Wang, Bin Liu 0032, Yuanxun Shao, Wandong Xie
Eng. Appl. Artif. Intell.2
2024 Joint Instance and Feature Adaptation for Heterogeneous Defect Prediction
abstract
Heterogeneous defect prediction (HDP) predicts defects for the current project (target) using heterogeneous data from other projects (source), which can improve software quality effectively. The ability to reduce the distribution divergence between the source and target project is crucial to the performance of HDP. Existing HDP methods address this issue only by using feature adaptation techniques, i.e., feature matching and feature transformation. However, without considering instance adaptability differences, they cannot fully exploit the potential of instance adaptation and thus cannot make the source and target distribution completely matched. To address this problem, we propose a novel HDP approach called joint instance and feature adaptation (JIFA) combining both instance adaptation and feature adaptation to narrow the gap between projects. Making joint use of both adaptation techniques, JIFA not only further reduces the distribution discrepancy but also enhances its robustness to abnormal outliers. Besides, discriminant information is also preserved in JIFA so that a better classification boundary can be obtained. Particularly, JIFA applies to both heterogeneous and homogeneous cross-project defect prediction (CPDP) tasks. The experimental results from 22 projects indicate that JIFA outperforms a range of advanced heterogeneous and homogeneous CPDP approaches.
Yujie Ren, Bin Liu 0032, Shihai Wang
IEEE Trans. Reliab.3
2023 ARRAY: Adaptive triple feature-weighted transfer Naive Bayes for cross-project defect prediction
Haonan Tong, Wei Lu 0010, Weiwei Xing, Shihai Wang
J. Syst. Softw.4
2022 SHSE: A subspace hybrid sampling ensemble method for software defect number prediction
abstract
Context: Software defect number prediction (SDNP) helps allocate limited testing resources by ranking software modules according to the predicted defect numbers. However, the highly skewed distribution of defects greatly degrades the performance of SDNP models by preventing SDNP models from ranking software modules accurately. Objective: This paper introduces a novel subspace hybrid sampling ensemble (SHSE) method based on feature subspace construction, hybrid sampling , and ensemble learning for building high-performance SDNP models. Method: Specifically, we first construct a series of feature subspace to ensure the diversity of base learners. In each of feature subspace, we then use the proposed hybrid sampling method to balance the training subset without losing too much information and introducing lots of noisy data caused by only using undersampling or oversampling techniques. Finally, we train each base learner and combine them by using the proposed weighted ensemble strategy. Experiments are performed on 27 public defect datasets. We compare SHSE with five state-of-the-art resampling-based models and four zero-inflated/hurdle models in terms of the ranking performance measure fault-percentile-average (FPA). To demonstrate the effectiveness of SHSE, two statistical testing methods including Wilcoxon Signed-rank test and Scott–Knott Effect Size Difference test are utilized. Cliff’s δ is also computed for quantifying the difference when there is significant difference between SHSE and each baseline. Results: The experimental results show that SHSE significantly outperforms the baselines and improves the performance over each baseline with as least medium effect size on most datasets. On average, SHSE improves the performance over the resampling-based methods by 8.7% ∼ 14.4% and the zero-inflate/hurdle models by 10.3% ∼ 15.2%. Conclusion: It can be concluded that SHSE is a more promising alternative for software defect number prediction.
Haonan Tong, Wei Lu 0010, Weiwei Xing, Bin Liu 0032, Shihai Wang
Inf. Softw. Technol.5
2022 WIFLF: An approach independent of the target project for cross-project defect prediction
abstract
Abstract Cross‐project defect prediction (CPDP) is used to build defect prediction models when data from the target project are not enough. There has been several approaches to improve the performance of CPDP, such as feature transformation and instance selection methods. However, existing techniques are strongly dependent on the target data to reduce the distribution discrepancy between source and target projects. That is, the performance of these methods is determined by the effectiveness of feature transformation or the similarity between two projects. Additionally, when there is a large amount of source data that needs to be matched with target data, it will take much time and reduce the efficiency of model construction. Therefore, it is vital to explore a target project‐agnostic approach to build CPDP models. This paper presents a Weighted Isolation Forest with class Label information Filter (WIFLF) to relieve the issues above. Four groups of datasets from AEEEM, Relink and PROMISE Data Repository are used to conduct CPDP models. Besides, WIFLF is compared with 12 approaches. The experimental results indicate that WIFLF significantly outperforms all the baselines. Specifically, WIFLF with random forest significantly improves the performance over the baselines on average by at least 14.64% and 4.90% with respect to Skewed F‐Measure and G‐Measure, respectively.
Bin Liu 0032, Shihai Wang
J. Softw. Evol. Process.3
2021 Efilter: An effective fault localization based on information entropy with unlabelled test cases
Xiaobo Yan, Bin Liu 0032, Shihai Wang, Yelin Yang
Inf. Softw. Technol.3
2021 A Test Restoration Method based on Genetic Algorithm for effective fault localization in multiple-fault programs
Xiaobo Yan, Bin Liu 0032, Shihai Wang
J. Syst. Softw.3
2021 Kernel Spectral Embedding Transfer Ensemble for Heterogeneous Defect Prediction
abstract
Cross-project defect prediction (CPDP) refers to predicting defects in the target project lacking of defect data by using prediction models trained on the historical defect data of other projects (i.e., source data). However, CPDP requires the source and target projects have common metric set (CPDP-CM). Recently, heterogeneous defect prediction (HDP) has drawn the increasing attention, which predicts defects across projects having heterogeneous metric sets. However, building high-performance HDP methods remains a challenge owing to several serious challenges including class imbalance problem, nonlinear, and the distribution differences between source and target datasets. In this paper, we propose a novel kernel spectral embedding transfer ensemble (KSETE) approach for HDP. KSETE first addresses the class-imbalance problem of the source data and then tries to find the latent common feature space for the source and target datasets by combining kernel spectral embedding, transfer learning, and ensemble learning. Experiments are performed on 22 public projects in both HDP and CPDP-CM scenarios in terms of multiple well-known performance measures such as, AUC, G-Measure, and MCC. The experimental results show that (1) KSETE improves the performance over previous HDP methods by at least 22.7, 138.9, and 494.4 percent in terms of AUC, G-Measure, and MCC, respectively. (2) KSETE improves the performance over previous CPDP-CM methods by at least 4.5, 30.2, and 17.9 percent in AUC, G-Measure, and MCC, respectively. It can be concluded that the proposed KSETE is very effective in both the HDP scenario and the CPDP-CM scenario.
Haonan Tong, Bin Liu 0032, Shihai Wang
IEEE Trans. Software Eng.3
2020 Software defect prediction based on correlation weighted class association rule mining
Yuanxun Shao, Bin Liu 0032, Shihai Wang
Knowl. Based Syst.3
2018 A novel software defect prediction based on atomic class-association rule mining
Yuanxun Shao, Bin Liu 0032, Shihai Wang
Expert Syst. Appl.3
2018 Software defect prediction using stacked denoising autoencoders and two-stage ensemble learning
Haonan Tong, Bin Liu 0032, Shihai Wang
Inf. Softw. Technol.3
2018 Feedback-based integrated prediction: Defect prediction based on feedback from software testing process
Peng Xiao 0003, Bin Liu 0032, Shihai Wang
J. Syst. Softw.3
2015 The impact of software process consistency on residual defects
abstract
Abstract Residual defects at the time of delivery are an important concern for safety critical software systems. Suppliers and customers are urged to get evidence for what they can do to reduce residual defects. Thus, it is meaningful to learn from historical data concerning the kinds of defects that have escaped from the existing quality assurance approaches and the factors that lead to the residual defects. A total of 3747 defects from 70 software systems developed by 29 Chinese aviation organizations were collected from acceptance tests during the last 5 years. For all these organizations, 38 domain experts from the industry assessed the process consistency to the standard built in the framework of Capability Maturity Model (CMM). Results demonstrate that the process improvement in the range of high consistency is effective in reducing total defects, as well as the minor and severe defects. The high consistency adoption of the practices in CMM Level 1 to Level 3 is more effective in reducing minor defects than severe defects. Causal analysis was performed to investigate the underlying mechanisms. Results reveal that individual cognitive failures cause 87% of severe defects. More approaches to help software developers manage their interior cognitive process are needed for improving software quality in the future. Copyright © 2015 John Wiley & Sons, Ltd.
Fuqun Huang, Bin Liu 0032, Shihai Wang, Qiuying Li
J. Softw. Evol. Process.3
2014 A new transfer learning Boosting approach based on distribution measure with an application on facial expression recognition
abstract
In the machine learning community, most algorithms proposed, particularly for inductive learning, are based entirely on one crucial assumption: that the training and test data points are drawn or generated from the exact same distribution. If this condition is not fully satisfied, most learning algorithms or models are corrupted. In this paper, we propose a new instance based transductive transfer learning method based on Boosting framework by using a distribution measure approach. There follows a detailed description of this distribution measure approach. Subsequently, we describe our boosting transfer learning method in detail and report its performance in facial expression recognition tasks.
Shihai Wang
IJCNN1
2012 A Robust Boosting by using an Adaptive Weight Scheme
abstract
In the real world, it is extremely difficult to avoid errors; for instance, a doctor may misdiagnose patients. In other words, databases are never free from data entry or other related errors, and many kinds of mistakes are unavoidable in real world data sets. In existing approaches to pattern recognition, handling noisy data in the learning process always produces better generalization performance than if the noise were ignored. In this article, a novel and adaptive weighting mechanism for noise learning tasks is proposed, especially for boosting learning approaches, preventing the algorithm from concentrating on unreasonably noisy learning samples. Several experiments on UC Irvine Machine Learning Repository and a facial expression data set demonstrate the effectiveness of our method.
Shihai Wang, Shin-Jye Lee
Cybern. Syst.1
2011 Semi-Supervised Learning via Regularized Boosting Working on Multiple Semi-Supervised Assumptions
abstract
Semi-supervised learning concerns the problem of learning in the presence of labeled and unlabeled data. Several boosting algorithms have been extended to semi-supervised learning with various strategies. To our knowledge, however, none of them takes all three semi-supervised assumptions, i.e., smoothness, cluster, and manifold assumptions, together into account during boosting learning. In this paper, we propose a novel cost functional consisting of the margin cost on labeled data and the regularization penalty on unlabeled data based on three fundamental semi-supervised assumptions. Thus, minimizing our proposed cost functional with a greedy yet stagewise functional optimization procedure leads to a generic boosting framework for semi-supervised learning. Extensive experiments demonstrate that our algorithm yields favorite results for benchmark and real-world classification tasks in comparison to state-of-the-art semi-supervised learning algorithms, including newly developed boosting algorithms. Finally, we discuss relevant issues and relate our algorithm to the previous work.
Ke Chen 0001, Shihai Wang
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Ensemble Learning with Active Data Selection for Semi-Supervised Pattern Classification
abstract
Unlike traditional pattern classification, semi-supervised learning provides a novel technique to make use of both labeled and unlabeled data for improving the performance of classification. In general, there are two critical issues for semi-supervised learning of discriminative classifiers; i.e., how to create an initial classifier of a good generalization capability with the limited labeled data and the how to make an effective use of unlabeled data without degradation of the established classifier. To tackle two aforementioned problems, we propose an ensemble learning approach based on a recent active data selection strategy, where ensemble learning would yield good generalization and active data selection tends to choose the unlabeled data more likely resulting in an improvement during semi-supervised learning. By using an ensemble of K-NN classifiers, we demonstrate the effectiveness of our approach on a synthetic data classification and a facial expression recognition tasks.
Shihai Wang, Ke Chen 0001
IJCNN1
2007 Regularized Boost for Semi-Supervised Learning
abstract
Semi-supervised inductive learning concerns how to learn a decision rule from a data set containing both labeled and unlabeled data. Several boosting algorithms have been extended to semi-supervised learning with various strategies. To our knowledge, however, none of them takes local smoothness constraints among data into account during ensemble learning. In this paper, we introduce a local smoothness regularizer to semi-supervised boosting algorithms based on the universal optimization framework of margin cost functionals. Our regularizer is applicable to existing semi-supervised boosting algorithms to improve their generalization and speed up their training. Comparative results on synthetic, benchmark and real world tasks demonstrate the effectiveness of our local smoothness regularizer. We discuss relevant issues and relate our regularizer to previous work.
Ke Chen 0001, Shihai Wang
NIPS2