Taghi M. Khoshgoftaar

dblp:k/TMKhoshgoftaar · DBLP profile ↗
← Back
306ranked-venue papers
102as first author
28since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 180 · 36 first-author · 28 since 2021Software engineering, systems software and programming languages · 106 · 61 first-authorApplied, interdisciplinary, general and emerging computing · 22 · 9 first-authorDatabases, data management, data science and information retrieval · 10 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 10Human-computer interaction and ubiquitous computing · 7 · 2 first-authorComputer networks · 2 · 2 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Medical Imaging with Deep Learning: A Comparison of CNN and Transformer Models
abstract
This study presents a comprehensive evaluation of deep learning approaches for medical image classification tasks, focusing on both convolutional neural networks (CNNs) and transformer-based models. Specifically, we assess the performance of three CNN architectures (InceptionNet, DenseNet, and EfficientNet) and three transformer-based models (Vision Transformer, Swin Transformer, and Pyramid Vision Transformer) using three publicly available medical imaging datasets. These datasets represent both balanced multi-class and imbalanced binary classification scenarios, reflecting real-world challenges in medical diagnostics. The experimental results consistently show that transformer-based models outperform conventional CNNs across all datasets, demonstrating superior performance both overall and at the individual class level, particularly in recall for clinically critical classes. This is crucial for minimizing severe Type II errors, which can result in missed diagnoses in practice. While EfficientNet, a CNN architecture, remains a strong baseline with performance often comparable to that of the transformers, attention-based models demonstrate effectiveness in handling complex medical imaging tasks. These findings underscore the potential of transformer-based models for clinical deployment when high recall is essential.
Kehan Gao, Sarah Tasneem, Taghi M. Khoshgoftaar
ICMLA3
2025 Applying Machine Learning to Classify Automobile Repossession Success
abstract
Automobile repossession and impoundment is crucial to enforcing loan repayments and the law, but remains unpredictable and challenging. A wide range of factors can prevent a successful repossession attempt, such as miscommunication on vehicle location or lien status, or debtor hostility. Automated risk estimation models may provide a valuable resource in allowing repossession companies to evaluate different jobs and allocate resources effectively; however, we find no existing work investigating statistical or machine-learning-based models in this field. To this end, we investigate the application of machine learning classification techniques to a private auto repossession dataset containing information on 200,000 distinct repossession attempts, including location, license plate status, and address type. We train a logistic regression classifier and two Gradient Boosted Decision Tree (GBDT) algorithms (CatBoost and XGBoost) to predict whether any given repo attempt will be a success or failure. We find that all three models, when optimized, give moderate classification performance (almost 80% accuracy and >0.50 F1 score), indicating that they provide valuable information for repossession companies and may support more effective decision-making and resource allocation. This is supported by extracted feature importance values from our GBDT models, which may provide additional insights to practitioners when prioritizing assignments. However, the application domain of auto repossession remains largely unexplored, and we highly encourage future work to explore other machine learning approaches or data collection techniques within this field.
Preston Billion-Polak, Andy Sinclair, Taghi M. Khoshgoftaar
ICMLA3
2025 Unsupervised Feature Extraction using Convolutional Autoencoder for Credit Card Fraud Detection
abstract
In many real-world applications, obtaining labeled data is challenging and expensive, making unsupervised learning methods essential for anomaly detection. Fraud detection is especially difficult because fraudulent transactions are often rare and hidden within a complex, high-dimensional feature space. Standard algorithms tend to favor non-fraudulent cases (majority class) and miss true fraud cases (minority class). To address these challenges of class imbalance and high-dimensionality, we propose CAE-IF, a unique hybrid approach that combines a Convolutional Autoencoder (CAE) for unsupervised feature extraction with an Isolation Forest (IF) for anomaly detection. We evaluate our method on the severely imbalanced Credit Card Fraud Detection dataset. First, the CAE learns compact, nonlinear embeddings of the transaction data. Next, these learned features are fed into the IF model, which identifies potentially fraudulent transactions. Our hybrid approach, CAE-IF, outperforms a baseline IF (without CAE) across Precision, Recall, F1-score, and Area Under the Precision-Recall Curve (AUPRC). Our findings indicate that CAE-IF is highly effective for unsupervised anomaly detection tasks even when the data is inherently imbalanced in nature.
Zahra Salekshahrezaee, Mary Anne Walauskis, Taghi M. Khoshgoftaar
ICMLA3
2025 Return-to-Work Classification of Occupational Injury Claims via Rank-Ordering
abstract
When an employee is injured on the job, the business may have costs related to lost wages and productivity, largely dependent on the length of the employee’s injury-related leave. While regression models can estimate absence duration, their predictions alone are often not precise enough to support decision-making, prompting the need for an advanced approach. In contrast, transforming the problem into a classification task allows for simpler determination of whether an absence is short or long, but depending on practical constraints, allocating resources to all cases classified as long absences may be inefficient. To address this, we introduce a rank-ordering framework that prioritizes the most critical cases using classification, enabling businesses to concentrate efforts on the top percentage of predicted absences. Using an extensive dataset of workers’ compensation claims, this study leverages CatBoost regression to predict injury absence in days and applies binary classification labels based on absence duration. Instead of predictive performance alone, we focus on the model’s ability to prioritize cases at risk of prolonged injury leave. Our rank-based evaluation shows that the models are able to flag a substantial portion of long-absence cases. To the best of our knowledge, our paper proposes the first application of a novel order-classification framework for absence length after a workplace injury, which can support return-to-work and resource planning.
Gonzalo A. Vivian, Chelsea M. Zuvieta, Taghi M. Khoshgoftaar
ICMLA3
2025 Examining the Impact of Feature Selection for Contrastive Learning in Fraud Detection
abstract
Contrastive learning approaches - which forgo labeling individual instances into different classes, and instead classify pairs of instances as either similar or dissimilar - have seen great success in difficult tasks such as few-shot image classification, due to their ability to quickly generalize to rare classes. However, these techniques are under-explored in the context of more traditional machine learning tasks with tabular, numeric data and only two classes. Our previous research has shown a Siamese Recurrent Neural Network (Siamese-RNN) to outperform traditional state-of-the-art techniques such as CatBoost on a highly-imbalanced credit card fraud detection task when combined with data sampling; we follow up on this work by examining the impact of applying supervised feature selection on the performance of this approach. We find that feature selection allows Siamese-RNN to produce streamlined feature representations, resulting in a statistically significant increase in AUPRC scores while also providing a 6-fold reduction in total data size.
Preston Billion-Polak, Taghi M. Khoshgoftaar
ICTAI2
2025 A Novel Technique to Rank-Order Occupational Injury Claims for Return-to-Work Prediction
abstract
Accurately forecasting the length of an employee's absence following workplace injuries remains a critical yet ongoing challenge. Such incidents place a significant strain on businesses through potential costs such as lost wages and productivity. Organizations may lack reliable tools to predict and prioritize the injury-related absences with the most lost workdays. This study leverages two machine learning models, CatBoost and a decision tree, to predict the number of days from injury to return-to-work date using a comprehensive dataset of workers' compensation claims that consists of employee demographics and injury-related features, such as injury severity, nature, and cause. Instead of regression performance alone, we focus on the models' ability to prioritize cases at risk of extended injury leave. We evaluate this using a rank-ordering framework to examine how well each model identifies cases with the longest absence durations. Our rank-based evaluation shows that both models are able to flag a substantial portion of high-risk cases. This paper proposes a novel predictive framework to support return-to-work planning and resource management. To the best of our knowledge, this is the first study to predict absence length after a workplace injury and assess regression model performance using ranking analysis in occupational health and safety.
Gonzalo A. Vivian, Chelsea M. Zuvieta, Taghi M. Khoshgoftaar
ICTAI3
2024 An Evaluation of Low-Shot Learning Techniques for the Detection of Credit Card Fraud
abstract
The task of credit card fraud detection presents many obstacles to traditional machine learning techniques, most notably severe class imbalance. Meanwhile, Low-Shot Learning (LSL) techniques have shown great effectiveness in multi-class problems involving significant class rarity, but have seen little examination in binary tasks such as fraud detection. In this paper, we select two low-shot learning papers from the literature (representing two major approaches to LSL), and replicate their methods on a highly imbalanced credit-card fraud dataset, to compare their performances to that of six contemporary state-of-the-art (SOTA) models. In the process, we improve on the experimental methodologies presented in the two selected papers by introducing multiple runs of cross-fold validation, measurement of the imbalance-robust performance metric Area Under the Precision-Recall Curve (AUPRC), and statistical analyses of our results. To the best of our knowledge, our work is the first to compare the two different LSL approaches to non-LSL SOTA methods, as well as the first peer-reviewed paper to evaluate any LSL method on a fraud detection dataset. We find that, while our chosen optimization-based method, Meta-Balance, underperforms the SOTA baseline in terms of AUPRC, our similarity-based method, Siamese-RNN, significantly outperforms the SOTA, and yields the highest AUPRC score recorded in the literature on the Kaggle credit card fraud dataset.
Preston Billion-Polak, Taghi M. Khoshgoftaar
ICMLA2
2024 New Class Labeling and Evaluation Methodology for Balanced and Highly Imbalanced Data
abstract
Acquiring labeled data in machine learning is challenging due to the high costs and need for domain expertise and human annotation, leaving much of the data unlabeled. In the healthcare domain, the prompt acquisition of labeled data is important for the early diagnosis of diseases such as dementia. Our research introduces an unsupervised method for generating binary class labels on two publicly available cognitive datasets, one balanced and one highly imbalanced (3% minority), derived from the Health and Retirement Study (HRS). Additionally, we employ an innovative evaluation approach that directly assesses the accuracy of the newly generated class labels, bypassing the need to train a supervised model to evaluate traditional performance metrics. This allows for an accurate quantification of the efficacy and reliability of the generated labels. Through this evaluation framework, we demonstrate the potential of our method to produce high-quality labels. Upon evaluation, our newly generated class labels significantly outperform our baseline method, Isolation Forest.
Mary Anne Walauskis, Taghi M. Khoshgoftaar
ICMLA2
2024 Enhancing Medicare Fraud Detection: Random Undersampling Followed by SHAP-Driven Feature Selection with Big Data
abstract
SHapley Additive exPlanations (SHAP) is a method used to explain the output of machine learning models. SHAP provides a unified measure of feature importance through SHAP values and serves as a feature selection tool for handling Big Data. This paper presents a study on optimizing feature selection using SHAP for Medicare fraud detection by applying the Random Undersampling (RUS) technique. Our approach aims to mitigate the big dataset's complexity, stemming from class imbalance, size, and high dimensionality, by employing RUS, followed by the integration of the SHAP model within a feature selection framework. The SHAP model is integrated using algorithms including LightGBM, XGBoost, CatBoost, and Decision Tree. To evaluate the effectiveness of our approach, we use the Area Under the Precision-Recall Curve (AUPRC) as the primary evaluation metric to measure the performance of the classification model, utilizing a Random Forest algorithm. Our experiments utilize the Medicare Part D dataset from the Centers for Medicare and Medicaid Services (CMS). Our primary objective is to investigate whether applying RUS with SHAP-based feature selection leads to measurable improvements in binary classification performance. Our findings indicate that the feature subset generated by applying RUS before selecting features using SHAP outperforms those created without the RUS enhancement. This approach not only enhances model performance, but also improves efficiency by reducing computational demands.
Qianxin Liang, Richard A. Bauder, Taghi M. Khoshgoftaar
ICTAI3
2024 A Comparison of Low-Shot Learning Methods for Imbalanced Binary Classification
abstract
The modern tasks of few-shot, one-shot, and zero-shot learning - or collectively Low-Shot Learning (LSL) -at first glance are quite similar to the long-standing task of class-imbalanced learning; specifically, they both aim to learn classes for which there is little labeled data available. Despite this similarity, and the shown effectiveness of LSL techniques in complex many-class image datasets, a previous literature review of the recent work in this overlap found that LSL methods are rarely applied to “classical” imbalanced tasks, i.e., those that are binary and tabular. To fill this gap, we select two of the papers in this survey that provide open-source models (each representing one of the two major approaches to LSL, optimization-based and similarity-based), train and thoroughly evaluate them on a traditional credit card fraud dataset, and compare their performances to each other. Our evaluation methodologies improve on those of the original works, by implementing an Area Under the Precision-Recall Curve (AUPRC) measurement, utilizing ten runs of five-fold cross-validation to ensure fairness, and conducting statistical analyses to confirm the statistical significance of our results. We find that, while our chosen optimization-based model outperforms somewhat in terms of Area Under the Receiver Operating Characteristic Curve (AUC-ROC), the similarity-based model vastly outperforms in terms of our primary metric, AUPRC, and shows promise for future research on similar datasets. To the best of our knowledge, we are the first to directly compare the performance of these two LSL approaches, and the first to examine a similarity-based model on this credit card fraud dataset.
Preston Billion-Polak, Taghi M. Khoshgoftaar
ICTAI2
2024 Confident Labels: A Novel Approach to New Class Labeling and Evaluation on Highly Imbalanced Data
abstract
A common challenge in machine learning is obtaining readily available labeled data, as the majority of data remains unlabeled and thereby necessitates human annotation from domain experts, e.g. doctors. Thus, substantial cost is associated with labeling the data, e.g. healthcare diagnostics. The efficient and effective acquisition of labeled data is crucial for early detection, specifically, in relation to cognitive decline. In our work, we employ a novel unsupervised method to generate binary class labels on a publicly available, highly imbalanced cognition dataset derived from the Health and Retirement Study (HRS). To evaluate the efficacy of our newly generated class labels, we also employ a novel approach to evaluate the labels by directly comparing them to ground-truth labels, rather than the traditional approach which measures the performance of a supervised model trained on the generated labels. Our results demonstrate that the newly generated class labels significantly outperform the baseline method, IF, in terms of Balanced Accuracy (BA), Geometric Mean (GM), F-Measure ($F$1), and Matthew's Correlation Coefficient (MCC).
Mary Anne Walauskis, Taghi M. Khoshgoftaar
ICTAI2
2023 Data Reduction to Improve the Performance of One-Class Classifiers on Highly Imbalanced Big Data
abstract
We propose and demonstrate a data reduction technique to improve the performance of one-class classifiers for highly imbalanced Big Data. We focus on insurance fraud detection in Medicare Part D prescription drug claims data. An initial comparison of the performance One-Class Support Vector Machines (SVMs) and One-Class Gaussian Mixture Models (GMMs) showed One-Class GMMs significantly out-performing One-Class SVMs. Therefore, we utilize One-Class Gaussian Mixture Models (GMMs). To enable faster training, we experiment with different levels of reduction in the size of the training data. Our novel application of a data reduction technique to one-class classification maintains model performance in terms of Area Under the Receiver Operating Characteristic Curve (AUC) and Area Under the Precision Recall Curve (AUPRC) scores, while simultaneously lowering model training time and computational resource consumption. Statistical tests confirm models trained on 20 percent of the training data perform equally well in terms of AUC and AUPRC. Furthermore, models trained on five percent of the training data yield AUC scores equivalent to models trained with all available training data. To the best of our knowledge, we are the first to demonstrate a viable data reduction technique for one-class classifiers applied to severely imbalanced Big Data.
John T. Hancock, Taghi M. Khoshgoftaar
ICMLA2
2023 A Model-Agnostic Feature Selection Technique to Improve the Performance of One-Class Classifiers
abstract
One-class classifiers hold promise for applications like fraudulent credit card transaction identification. However, interpreting these models to understand which features drive predictions is challenging. Such an understanding is necessary to avoid brute-force approaches to feature selection. This paper explores SHAP (SHapley Additive exPlanations) for feature selection with one-class classffiers on a credit card fraud dataset. We apply SHAP to select key features and evaluate One-Class Gaussian mixture models and One-Class Support Vector Machines. Statistical analysis tests show Gaussian mixture models built with SHAP-selected features perform significantly better than Gaussian mixture models built without feature selection. To the best of our knowledge, we are the first to show the benefit of SHAP-based feature selection to One-Class Gaussian mixture models. Moreover, we show that robust performance with features of the full dataset may be a prerequisite in order for SHAP feature selection to impart further gains. Our results provide novel evidence that SHAP can identify informative features for one-class classifiers.
John T. Hancock, Richard A. Bauder, Taghi M. Khoshgoftaar
ICTAI3
2023 One-Class Classifier Performance: Comparing Majority versus Minority Class Training
abstract
Our study evaluates the impact of training OneClass Classification (OCC) algorithms on the majority class compared to training them on the minority class, using large and big datasets that are highly imbalanced. It is important to note that class availability can present a significant challenge in model training. In reality, there may be situations where one class is readily obtainable within a reasonable time frame while another class is not. Our task involves detecting instances of fraud in the Credit Fraud Detection Dataset and our Medicare dataset derived from Medicare Part D data and List of Excluded Individuals and Entities (LEIE) data. The Credit Card Fraud Detection Dataset has real-world transaction content as well as a significant class imbalance, making it suitable for use as a benchmark for credit card fraud detection. In addition, it is the only publicly available large data for credit card fraud analysis. Part D is big data, allowing researchers to analyze national trends and patterns in prescription drug usage and expenditures. The algorithms used in the study are One-Class Gaussian Mixture Model (GMM), OneClass Adversarial Nets (OCAN), and One-Class Support Vector Machine (SVM). Their performance is measured with the Area Under the Precision-Recall Curve (AUPRC) and Area Under the Receiver Operating Characteristic Curve (AUC). Our results indicate that OCC produces better results when models are trained on the majority class.
Joffrey L. Leevy, John T. Hancock, Taghi M. Khoshgoftaar, Azadeh Abdollah Zadeh
ICTAI3
2022 Informative Evaluation Metrics for Highly Imbalanced Big Data Classification
abstract
We conduct experiments that show the Area Under the Precision Recall Curve (AUPRC) metric provides a more meaningful insight into the impact of Random Undersampling than Area Under the Receiver Operating Characteristic Curve (AUC). Evaluating experiments with multiple metrics is a robust method for overcoming challenges in Machine Learning, such as class imbalance. Random Undersampling is a technique to deal with class imbalance. We find Random Undersampling may provide an improvement to AUC scores. However, at the same time, Random Undersampling may be detrimental to AUPRC scores. AUPRC is a metric that involves precision, whereas AUC does not. In the classification of imbalanced Big Data, an increase in false positive counts has a more noticeable drop in precision scores. Therefore, in application domains where false positives are undesirable, optimizing models for AUPRC is a wise choice. Our contribution is to compare the performance of models in terms of AUPRC and AUC to show the impact of Random Undersampling on the classification of imbalanced Big Data. We compare the performance via experiments in the classification of highly imbalanced Big Data. Models are built with data in its original class ratio, and with data undersampled into 5 distinct class ratios. We report the results of 600 experiments where we apply Random Undersampling to a dataset with about 175 million instances. To the best of our knowledge we are the first to utilize Medicare Part D data which became available in 2021.
John T. Hancock, Taghi M. Khoshgoftaar, Justin M. Johnson
ICMLA2
2022 Cost-Sensitive Ensemble Learning for Highly Imbalanced Classification
abstract
There are a variety of data-level and algorithm-level methods available for treating class imbalance. Data-level methods include data sampling strategies that pre-process training data to reduce levels of class imbalance. Algorithm-level methods modify the learning and inference processes to reduce bias towards the majority class. This study evaluates both data-level and algorithm-level methods for class imbalance using a highly imbalanced healthcare fraud data set. We approach the problem from a cost-sensitive learning perspective, and demonstrate how these direct and indirect cost-sensitive methods can be implemented using a common cost matrix. For each method, a wide range of costs are evaluated using three popular ensemble learning algorithms. Initial results show that random undersampling (RUS) and class weighting are both effective ways to improve classification when the default classification threshold is used. Further analysis using the area under the precision-recall curve, however, shows that both RUS and class weighting actually decrease the discriminative power of these learners. Through multiple complementary performance metrics and confidence interval analysis, we find that the best model performance is consistently obtained when RUS and class weighting are not applied, but when output thresholding is used to maximize the confusion matrix instead. Our contributions include various recommendations related to implementing cost-sensitive ensemble learning and effective model evaluation, as well as empirical evidence that contradicts popular beliefs about learning from imbalanced data.
Justin M. Johnson, Taghi M. Khoshgoftaar
ICMLA2
2022 Evaluating Performance Metrics for Credit Card Fraud Classification
abstract
Practitioners and researchers of machine learning should have a deep understanding about the selection of the right performance metrics for classifier evaluation. Using a credit card fraud dataset, we demonstrate that the Area Under the Precision-Recall Curve (AUPRC) metric is a more reliable measurement, for the classification of highly imbalanced data, than the Area Under the Receiver Operating Characteristic Curve (AUC) metric. Furthermore, we establish that AUC is minimally impacted by the use of Random Undersampling (RUS). The classifiers used in this study are ensemble learners: LightGBM, CatBoost, Extremely Randomized Trees (ET), XGBoost, and Random Forest. Our results are governed by the fact that in a highly imbalanced dataset, the comparatively large number of true negative instances has an influence on AUC but not on AUPRC. Hence, AUPRC is able to accurately detect changes in the number of false positives because it ignores the true negatives.
Joffrey L. Leevy, Taghi M. Khoshgoftaar, John T. Hancock
ICTAI2
2022 GANs for Class-Imbalanced Data: A Meta-Analysis of GitHub Projects
abstract
Generative Adversarial Networks (GANs) have increasingly been the subject of intense research interest for their ability to augment datasets to correct for class imbalance. The collaborative and complex code bases for these GANs are often written in a high-level, general-purpose programming language such as Python and housed on the GitHub platform. The goal of this work is to summarize the research aims of 18 GitHub repositories of projects that implement GANs in regimes of class imbalance, in both tabular and non-tabular settings, as well to summarize and analyze the patterns and characteristics of these code bases. With respect to the latter task, we conduct our analysis from a perspective of library reliance and from the structural properties of the code base. The insights discovered herein are meant to serve as a gentle introduction to the various tools available and recommended best practices for the enterprising researcher seeking to apply GANs as data augmenters in imbalanced settings.
Rick Sauber-Cole, Taghi M. Khoshgoftaar, Justin M. Johnson
ICTAI2
2022 Exploring Language-Interfaced Fine-Tuning for COVID-19 Patient Survival Classification
abstract
We present Language-Interfaced Fine-Tuning (LIFT) in application to COVID-19 patient survival classification. LIFT describes translating tabular Electronic Health Records (EHRs) into text inputs for transformer neural networks. We study LIFT with a dataset of 5,371 COVID-19 patients. We focus on the predictive task of survival classification utilizing demographic and medical history features. We begin by presenting information about our dataset. We preface our investigation in text-based transformers by reporting the performances of conventional machine learning models such as Logistic Regression and Random Forest classifiers. We also present the results of a few configurations of tabular input-based Deep Multilayer Perceptron (MLP) networks. 86% of the patients in our database survived in the measured time window. Thus, predictive models are heavily biased to predict that a patient will survive. We emphasize that this problem of Class Imbalance was a major challenge in developing these models. Our balanced sampling strategy from examples in the majority and minority classes is crucial to achieving even reasonable predictive performance. For this reason, we also report performance based on Precision, Recall, and F-score metrics, in addition to Accuracy. Having established baselines with tabular inputs, we then shift our focus to the prompts for translating from tabular to text inputs. We report the performance of 5 prompts. The LIFT model achieves an F-score on the held-out test set of 0.21, slightly behind the Deep MLP with Tabular Features score of 0.23. Both models outperform the Random Forest with Tabular Features at 0.15. We believe that LIFT is a very exciting direction for machine learning in healthcare applications because text-based inputs enables us to take advantage of recent advances in Transfer Learning and Retrieval-Augmented Learning. This study illustrates the effectiveness of converting tabular EHRs to text inputs and utilizing transformer neural networks for prediction.
Connor Shorten, Erika Cardenas, Taghi M. Khoshgoftaar, Javad Hashemi, Safiya George Dalmida, Debarshi Datta, Laurie Martinez, Candice Sareli, Paula Eckardt
ICTAI3
2021 Detecting SSH and FTP Brute Force Attacks in Big Data
abstract
We present a simple approach for detecting brute force attacks in the CSE-CIC-IDS2018 Big Data dataset. We show our approach is preferable to more complex approaches since it is simpler, and yields stronger classification performance. Our contribution is to show that it is possible to train and test simple Decision Tree models with two independent variables to classify CSE-CIC-IDS2018 data with better results than reported in previous research, where more complex Deep Learning models are employed. Moreover, we show that Decision Tree models trained on data with two independent variables perform similarly to Decision Tree models trained on a larger number independent variables. Our experiments reveal that simple models, with AUC and AUPRC scores greater than 0.99, are capable of detecting brute force attacks in CSE-CIC-IDS2018. To the best of our knowledge, these are the strongest performance metrics published for the machine learning task of detecting these types of attacks. Furthermore, the simplicity of our approach, combined with its strong performance, makes it an appealing technique.
John T. Hancock, Taghi M. Khoshgoftaar, Joffrey L. Leevy
ICMLA2
2021 Robust Thresholding Strategies for Highly Imbalanced and Noisy Data
abstract
Many studies have shown that non-default decision thresholds are required to maximize classification performance on highly imbalanced data sets. Thresholding strategies include using a threshold equal to the prior probability of the positive class or identifying an optimal threshold on training data. It is not clear, however, how these thresholding strategies will generalize to imbalanced data sets that contain class label noise. When class noise is present, the positive class prior is influenced by the class label noise, and a threshold that is optimized on noisy training data may not generalize to test data. We employ four thresholding strategies: two thresholds that are optimized on training data and two thresholds that depend on the positive class prior. Threshold strategies are evaluated on a range of noise levels and noise distributions using the Random Forest, Multilayer Perceptron, and XGBoost learners. While all four thresholding strategies significantly outperform the default threshold with respect to the Geometric Mean (G-Mean), three of the four thresholds yield unstable true positive rates (TPR) and true negative rates (TNR) in the presence of class noise. Results show that setting the threshold equal to the prior probability of the noisy positive class consistently performs best according to G-Mean, TPR, and TNR. This is the first evaluation of thresholding strategies for imbalanced and noisy data, to the best of our knowledge, and our results contradict related works that have suggested optimizing thresholds on training data as the best approach.
Justin M. Johnson, Taghi M. Khoshgoftaar
ICMLA2
2021 Detecting Information Theft Attacks in the Bot-IoT Dataset
abstract
There are growing security risks tied to the recent proliferation of Internet of Things (IoT) devices. Due to this fact, datasets such as Bot-IoT were designed to train machine learning classifiers on network intrusion detection in IoT networks. In this research, we use Bot-IoT to build a predictive model for detecting information theft attacks. Our contribution is defined by the unique approach of using eight classifiers and two performance metrics to detect information theft traffic. Also, to the best of our knowledge, the Bot-IoT Information Theft category of attacks has never been the focus of a research paper. Our group of classifiers is a diverse range of four ensembles (CatBoost, Light-GBM, XGBoost, and Random Forest) and four non-ensembles (Decision Tree, Logistic Regression, Naive Bayes, and a Multilayer Perceptron (MLP)). The metrics used to evaluate the classifiers are Area Under the Receiver Operating Characteristic Curve (AUC) and Area Under the Precision-Recall Curve (AUPRC). Through cross-validation, we train and test Bot-IoT instances (only normal and information theft traffic) to evaluate the best classifier(s). According to our results, the ensemble classifiers, particularly CatBoost, LightGBM, and XGBoost, are the top-performing models.
Joffrey L. Leevy, John T. Hancock, Taghi M. Khoshgoftaar, Jared M. Peterson
ICMLA3
2021 KerasBERT: Modeling the Keras Language
abstract
We introduce a new application domain to evaluate the knowledge retention of language models. Our model, Keras-BERT, is trained on the Keras code documentation. This is a unique challenge of language modeling small datasets, as well as combining natural language and code data. We evaluate how well KerasBERT learns the Keras Deep Learning framework through cloze test evaluation. We present miscellaneous properties of these cloze tests such as mask positioning and prompt paraphrasing. KerasBERT is an 80 million parameter RoBERTa model, which we compare to the Zero-Shot learning capability of the 6 billion parameter GPT-Neo model. We present a suite of cloze tests crafted from the Keras documentation to evaluate these models. We find some exciting completions that show KerasBERT is a promising direction for question answering and schema-free database querying. We conclude our work by discussing some future directions for KerasBERT and the development of language models for code documentation support.
Connor Shorten, Taghi M. Khoshgoftaar
ICMLA2
2021 Feature Popularity Between Different Web Attacks with Supervised Feature Selection Rankers
abstract
We introduce the novel concept of feature popularity with three different web attacks and big data from the CSE-CIC-IDS2018 dataset: Brute Force, SQL Injection, and XSS web attacks. Feature popularity is based upon ensemble Feature Selection Techniques (FSTs) and allows us to more easily understand common important features between different cyberattacks, for two main reasons. First, feature popularity lists can be generated to provide an easy comprehension of important features across different attacks. Second, the Jaccard similarity metric can provide a quantitative score for how similar feature subsets are between different attacks. Both of these approaches not only provide more explainable and easier-to-understand models, but they can also reduce the complexity of implementing models in real-world systems. Four supervised learning-based FSTs are used to generate feature subsets for each of our three different web attack datasets, and then our feature popularity frameworks are applied. For these three web attacks, the XSS and SQL Injection feature subsets are the most similar per the Jaccard similarity. The most popular features across all three web attacks are: Flow_Bytes_s, Flow_IAT_Max, and Flow_Packets_s. While this introductory study is only a simple example using only three web attacks, this feature popularity concept can be easily extended, allowing an automated framework to more easily determine the most popular features across a very large number of attacks and features.
Richard Zuech, John T. Hancock, Taghi M. Khoshgoftaar
ICMLA3
2021 Output Thresholding for Ensemble Learners and Imbalanced Big Data
abstract
Class imbalance is a common problem in many real-world machine learning applications that has been shown to significantly degrade classification performance. This is especially true in the context of big data, where large volumes of data from the majority class dominate training processes and bias learning algorithms. Of the various methods for treating class imbalance, output thresholding is one technique that improves classification performance by tuning the decision threshold that is used to assign class labels to class probabilities. While thresholding techniques have been successful, systematic studies within big and imbalanced data applications are limited. In this study, we compare four popular thresholding strategies using two big and imbalanced fraud classification data sets. We focus specifically on tree-based ensemble learners, and employ four popular bagging and boosting ensemble learners that are well known for achieving state-of-the-art performance. Overall classification is measured using the Geometric Mean (G-Mean) and F-Measure metrics and class-wise performance tradeoffs are compared using the true positive rate (TPR) and true negative rate (TNR). The average threshold values of each strategy are compared, and statistical tests are provided to illustrate the importance of careful threshold tuning. Results show that the G-Mean and F-Measure metrics provide misleading results, and careful validation of TPR and TNR is necessary for selecting optimal thresholds. Furthermore, we show how small changes to decision thresholds yield significant changes to classification performance. Our comparison of popular thresholding techniques on both big and highly-imbalanced data makes this a unique contribution in the area of output thresholding with ensemble learners and big data.
Justin M. Johnson, Taghi M. Khoshgoftaar
ICTAI2
2021 The Effects of Class Label Noise on Highly-Imbalanced Big Data
abstract
This study explores the effects of class label noise on a highly-imbalanced big data set by injecting varying levels of class noise into a Medicare Part B fraud detection data set. Noise parameters are used to vary the total level of class noise and the proportion of class noise between the majority and minority classes. This allows us to better understand how class noise affects imbalanced data sets and where data cleaning efforts should be focused. Four popular machine learning algorithms are evaluated using six rounds of five-fold cross-validation to determine which learners are most robust to class noise. Area under the precision-recall curve (AUPRC) results shows that negative class noise, i.e. when positive instances are incorrectly labeled as negative, has the greatest adverse effect on classification performance. Statistical results show that the XGBoost learner performs significantly better than Random Forest, Multilayer Perceptron, and Logistic Regression Learners. True positive rates and true negative rates show that there is a trade-off that occurs as the noise proportion switches between the majority and negative classes. Finally, we show that the effects of class label noise can be combatted by regularizing the XGBoost learner through shallower decision trees.
Robert K. L. Kennedy, Justin M. Johnson, Taghi M. Khoshgoftaar
ICTAI3
2021 Feature Extraction for Class Imbalance Using a Convolutional Autoencoder and Data Sampling
abstract
Training a machine learning algorithm from a class-imbalanced dataset is an inherently challenging task. The task becomes more challenging when compounded by high dimensionality (a high number of features). Feature extraction is a data reduction process that transforms features into linear or non-linear combinations of the original features, resulting in a smaller and richer set of attributes. Data sampling is a popular approach for addressing class imbalance. In this paper, our proposed method requires the implementation of feature extraction before data sampling, based on the idea that richer high-level features facilitate more efficient sampling and hence, produce better classification results. We use principal component analysis (PCA) and convolutional autoencoder (CAE) as the feature extraction techniques and synthetic minority oversampling technique (SMOTE) as the data sampling technique. In evaluating the performance of the random forest classifier on a credit card fraud dataset, our results show that CAE is a better feature extraction technique than PCA. The combination of CAE followed by SMOTE yields the best F1-score of 90.5%.
Zahra Salekshahrezaee, Joffrey L. Leevy, Taghi M. Khoshgoftaar
ICTAI3
2021 Investigating the Generalization of Image Classifiers with Augmented Test Sets
abstract
Adding prior knowledge about the task or domain being learned can greatly facilitate learning. Two of the most common examples of injecting prior knowledge into Deep Learning systems are architecture design and data augmentation. For example, the convolutional architecture biases the model to learn local features. This emphasis on local features has been a useful prior for image processing. Data augmentation is full of examples that utilize prior knowledge. For example, cropping an image and preserving the original label gives the inductive bias of local feature importance. In this study, we aim to see how the priors in architecture interplay with augmentations. We begin by showing the overall benefit of training with data augmentation, improving the Vision Transformer’s test accuracy from 74.4% to 84.3% and improving the ResNet’s test accuracy from 79.3% to 86.7%. We focus on the distinction between global and local priors such as the difference between the convolution and attention layers and cropping versus noise addition augmentations. These tests are not yet able to find complementing flaws in architectures and augmentations. We find that neither the ResNet or the Vision Transformer is robust to distribution shifts controlled with data augmentation. The performance of both models degrades heavily even with moderate augmentation strengths. Although remedied by explicitly training with the augmentation used to construct the test set, we still see a notable decrease in performance. This study illustrates the utility of generalization testing with data augmentation and the challenge of measuring the impact of global and local priors in architecture.
Connor Shorten, Taghi M. Khoshgoftaar
ICTAI2
2020 Evaluating The Number of Trainable Parameters on Deep Maxout and LReLU Networks for Visual Recognition
abstract
Object recognition research has made notable steps since the appearance of convolutional neural networks, and many activation functions have been proposed to enhance the classification performance of these networks. Maxout networks have achieved great success in many computer vision tasks, but there is limited information on whether an increase in the number of trainable parameters can increase the performance in Leaky Rectified Linear Unit (LReLU) networks compared to maxout networks. Our experiments compare LReLU, rectified linear unit, scaled exponential linear unit, and hyperbolic tangent to four maxout variants. We evaluate ReLU and LReLU with 2x, 3x and 6x the number of filters in each convolutional layer. We also evaluate ReLU, LReLU and maxout networks with approximately the same number of trainable parameters. Under equal conditions, we found that on average, across all datasets, LReLU performs better than any of the evaluated activation functions.
Gabriel Castaneda Oscos, Paul Morris, Taghi M. Khoshgoftaar
ICMLA3
2020 Performance of CatBoost and XGBoost in Medicare Fraud Detection
abstract
Due to the size of the data involved, performance is an important consideration in the task of detecting fraudulent Medicare insurance claims. We evaluate CatBoost and XGBoost on the task of Medicare fraud detection, and report performance in terms of running time and Area Under the Receiver Operating Characteristic Curve (AUC). We show that adding a categorical feature for XGBoost and CatBoost improves performance in terms of AUC, and that CatBoost's performance is higher in a statistically significant sense. Moreover, we conduct experiments to find the optimal number of decision trees to use for XGBoost and CatBoost in the task of Medicare fraud detection. This is an important contribution because the number of trees in the ensemble governs overall resource consumption of a Gradient Boosted Decision Tree implementation. We find that with a purely numerical dataset, CatBoost and XGBoost yield nearly equivalent performance in terms of AUC, and XGBoost has a shorter training time. With respect to Medicare fraud detection, to the best of our knowledge, this is the first study to evaluate the performance of CatBoost and XGBoost in terms of running time and AUC on highly imbalanced, Big Data. Our contribution of evaluating running time performance on a large imbalanced dataset benefits researchers looking for more efficient utilization of valuable resources.
John T. Hancock, Taghi M. Khoshgoftaar
ICMLA2
2020 Accelerated Deep Learning on HPCC Systems
abstract
Modern deep learning architectures are computationally expensive to train due to their size, complexity, and the data they typically train on. There are different approaches to accelerate the training process. Using graphics processing units (GPUs) to accelerate deep learning is an increasingly effective and accessible way for researchers to train large neural networks in a fraction of the time. In this paper, we trained neural networks on High-Performance Computing Cluster (HPCC) Systems Platform using GPU acceleration with a novel implementation that combines widely used neural network libraries and the HPCC Systems Platform. This is the first work that uses GPU acceleration on HPCC Systems and is the first to demonstrate the effectiveness of using single and multiple GPUs with HPCC Systems. We provide experiments that measure the performance increase in training time as compared to only using central processing units (CPUs). The experiments trained a convolutional neural network on image data and a multilayer perceptron trained on a Medicare fraud detection dataset, using both single and multiple GPUs to accelerate the training process. Our results show that although GPU usage does not always guarantee a significant performance increase, using one or more GPUs does generally decrease the required training time.
Robert K. L. Kennedy, Taghi M. Khoshgoftaar
ICMLA2
2020 A study on rare fraud predictions with big Medicare claims fraud data
abstract
Access to copious amounts of information has reached unprecedented levels, and can generate very large data sources. These big data sources often contain a plethora of useful information but, in some cases, finding what is actually useful can be quite problematic. For binary classification problems , such as fraud detection, a major concern therein is one of class imbalance. This is when a dataset has more of one label versus another, such as a large number of non-fraud observations with comparatively few observations of fraud (which we consider the class of interest). Class rarity further delineates class imbalance with significantly smaller numbers in the class of interest. In this study, we assess the impacts of class rarity in big data, and apply data sampling to mitigate some of the performance degradation caused by rarity. Real-world Medicare claims datasets with known excluded providers are used as fraud labels for a fraud detection scenario, incorporating three machine learning models. We discuss the necessary data processing and engineering steps in order to understand, integrate, and use the Medicare data. From these already imbalanced datasets, we generate three additional datasets representing varying levels of class rarity. We show that, as expected, rarity significantly decreases model performance, but data sampling, specifically random undersampling, can help significantly with rare class detection in identifying Medicare claims fraud cases.
Richard A. Bauder, Taghi M. Khoshgoftaar
Intell. Data Anal.2
2019 Deep Learning and Thresholding with Class-Imbalanced Big Data
abstract
Class imbalance is a regularly occurring problem in machine learning that has been studied extensively over the last two decades. Various methods for addressing class imbalance have been introduced, including algorithm-level methods, datalevel methods, and hybrid methods. While these methods are well studied using traditional machine learning algorithms, there are relatively few studies that explore their application to deep neural networks. Thresholding, in particular, is rarely discussed in the deep learning with class imbalance literature. This paper addresses this gap by conducting a systematic study on the application of thresholding with deep neural networks using a Big Data Medicare fraud data set. We use random oversampling (ROS), random under-sampling (RUS), and a hybrid ROS-RUS to create 15 training distributions with varying levels of class imbalance. With the fraudulent class size ranging from 0.03%-60%, we identify optimal classification thresholds for each distribution on random validation sets and then score the thresholds on a 20% holdout test set. Through repetition and statistical analysis, confidence intervals show that the default threshold is never optimal when training data is imbalanced. Results also show that the optimal threshold outperforms the default threshold in nearly all cases, and linear models indicate a strong linear relationship between the minority class size and the optimal decision threshold. To the best of our knowledge, this is the first study to provide statistical results that describe optimal classification thresholds for deep neural networks over a range of class distributions.
Justin M. Johnson, Taghi M. Khoshgoftaar
ICMLA2
2019 The Effect of Time on the Maintenance of a Predictive Model
abstract
Periodic updating of a machine learning model may become necessary because new data could have a distribution that has drifted significantly over time from the original data distribution, thus impacting the model's usefulness. The primary objective of this paper is to evaluate temporal influence on the maintenance of a predictive model. We investigate the impact of using training data from various year-groupings on a model designed to detect Medicare Part B billing fraud. Training datasets are obtained from year-groupings of 2015, 2014-2015, 2013-2015, and 2012-2015. The test dataset is represented by 2016 data. Our study utilizes five popular learners and five class ratios obtained by Random Undersampling. Using the Area Under the Receiver Operating Characteristic (ROC) Curve as the performance metric, our case study indicates that the Logistic Regression learner yields the highest overall value for the yeargrouping of 2013-2015, with a majority-to-minority ratio of 90:10. For the problem of maintaining predictive models for Medicare fraud, we conclude that a sampled dataset should be chosen over the full dataset and that the largest training dataset (i.e., 2012- 2015) does not always produce the best results. To the best of our knowledge, this is the first big data study that examines the influence of time on the maintenance of machine learning models.
Joffrey L. Leevy, Taghi M. Khoshgoftaar, Richard A. Bauder, Naeem Seliya
ICMLA2
2019 Learning Curve Estimation with Large Imbalanced Datasets
abstract
Datasets for machine learning are constantly increasing in size, along with computational requirements for processing the data. A useful exercise for machine learning experiments is to approximate model performance as dataset size increases. This can inform application building and data collection efforts as well as improve computational efficiency by using subsets of the data. In this paper, we evaluate a learning curve estimation method on three large imbalanced datasets. Estimation is performed by fitting an inverse power law model to a learning curve created on a small amount of data. We then explore how well this estimated curve fits to the full learning curve of each dataset. The method has been previously evaluated for small datasets (hundreds or thousands of instances), and in this study we show that the method is indeed effective for larger datasets with millions of instances. This is beneficial because only a few thousand instances are required to accurately estimate the performance of models using millions of instances. To the best of our knowledge, this is the first study to systematically explore the use of an inverse power law curve fitting method for big data.
Aaron N. Richter, Taghi M. Khoshgoftaar
ICMLA2
2019 A Study on Software Metric Selection for Software Fault Prediction
abstract
For most software systems, superfluous software metrics are often collected. Sometimes, metrics that are collected may be redundant or irrelevant to fault prediction results. Feature (software metric) selection helps separating relevant software metrics from irrelevant or redundant ones, thereby identifying the small set of software metrics that are best predictors of fault proneness for new components, modules, or releases. In this study, we compare three forms of feature selection techniques (filter-and wrapper-based subset evaluators along with two search techniques (Best First (BF) and Greedy Stepwise (GS)), and feature ranking on four datasets from a real world software project. Five learners are used to build fault prediction models with the selected software metrics. Each model is assessed using the Area Under the Receiver Operating Characteristic Curve (AUC). We find that wrapper-based subset evaluators performed best and feature ranking performed worst. In addition, the model built with the logistic regression (LR) learner performs best in terms of the AUC performance metric. This leads us to recommend the use of the wrapper-based subset evaluators to select software metric subsets and the LR learner for building software fault prediction models.
Huanjing Wang, Taghi M. Khoshgoftaar
ICMLA2
2019 Threshold Based Optimization of Performance Metrics with Severely Imbalanced Big Security Data
abstract
Proper evaluation of classifier predictive models requires the selection of appropriate metrics to gauge the effectiveness of a model's performance. The Area Under the Receiver Operating Characteristic Curve (AUC) has become the de facto standard metric for evaluating this classifier performance. However, recent studies have suggested that AUC is not necessarily the best metric for all types of datasets, especially those in which there exists a high or severe level of class imbalance. There is a need to assess which specific metrics are most beneficial to evaluate the performance of highly imbalanced big data. In this work, we evaluate the performance of eight machine learning techniques on a severely imbalanced big dataset pertaining to the cyber security domain. We analyze the behavior of six different metrics to determine which provides the best representation of a model's predictive performance. We also evaluate the impact that adjusting the classification threshold has on our metrics. Our results find that the C4.5N decision tree is the optimal learner when evaluating all presented metrics for severely imbalanced Slow HTTP DoS attack data. Based on our results, we propose that the use of AUC alone as a primary metric for evaluating highly imbalanced big data may be ineffective, and the evaluation of metrics such as F-measure and Geometric mean can offer substantial insight into the true performance of a given model.
Chad Calvert, Taghi M. Khoshgoftaar
ICTAI2
2019 Approximating Learning Curves for Imbalanced Big Data with Limited Labels
abstract
Labeling data for supervised learning can be an expensive task, especially when large amounts of data are required to build an adequate classifier. For most problems, there exists a point of diminishing returns on a learning curve where adding more data only marginally increases model performance. It would be beneficial to approximate this point for scenarios where there is a large amount of data available but only a small amount of labeled data. Then, time and resources can be spent wisely to label the sample that is required for acceptable model performance. In this study, we explore learning curve approximation methods on a big imbalanced dataset from the bioinformatics domain. We evaluate a curve fitting method developed on small data using an inverse power law model, and propose a new semi-supervised method to take advantage of the large amount of unlabeled data. We find that the traditional curve fitting method is not effective for large sample sizes, while the semi-supervised method more accurately identifies the point of diminishing returns.
Aaron N. Richter, Taghi M. Khoshgoftaar
ICTAI2
2018 An Empirical Study on Class Rarity in Big Data
abstract
The problem of class imbalance, especially the classification of rare cases, is an important area in machine learning. These rare cases are typically the ones of interest, thus accurate classification of these instances is required. Class imbalance is a well-studied area with relatively small datasets, but there is limited research focusing on both rarity and class imbalance with Big Data. In this study, we focus on the impact of rare class classification in the area of fraud detection using publicly available real-world Big Data from Medicare data sources. We demonstrate that rarity significantly degrades fraud detection performance over three machine learning models and nine datasets, with varying numbers of positive class instances. From these experiments, we show clear groupings indicating different levels of class imbalance and rarity. Furthermore, our results, showing decreasing performance with increasing rarity, are corroborated using three additional Medicare Big Data sources.
Richard A. Bauder, Taghi M. Khoshgoftaar, Tawfiq Hasanin
ICMLA2
2018 Data Sampling Approaches with Severely Imbalanced Big Data for Medicare Fraud Detection
abstract
Class imbalance is an important problem in machine learning. With increases in available information and the growing use of Big Data sources to extract meaning from data, the challenges associated with class imbalance continue to influence research and shape business value. In this paper, we focus on using highly imbalanced Big Data from Medicare to detect provider claims fraud. We combine three Medicare parts and generate fraud labels using real-world excluded providers. The number of known fraudulent providers is very small, with 0.062% of the combined dataset being labeled as fraud, indicating severe class imbalance. To address class imbalance concerns, we provide experimental results incorporating six different data sampling methods (undersampling and oversampling) to create datasets for five class ratios (imbalanced to balanced), as well as using the full dataset (with no sampling). Three state-of-the-art machine learning models with Apache Spark are used to assess Medicare fraud detection performance across data sampling methods and class ratios. We demonstrate that data sampling, in particular random undersampling, presents good results across all learners, whereas oversampling provides no benefit versus models built using the full dataset.
Richard A. Bauder, Taghi M. Khoshgoftaar, Tawfiq Hasanin
ICTAI2
2018 Building and Interpreting Risk Models from Imbalanced Clinical Data
abstract
As more clinical data becomes available for research, it is important to be able to build effective models and understand the predictions made from them. In this paper, we present a case study modeling melanoma risk using structured clinical records. Advanced modeling techniques are required as the data set is large, sparse, and imbalanced. We explore the use of logistic regression, decision tree, and random forest classifiers with various feature selection and random undersampling techniques. For clinical models to be used in practice, both providers and patients should have insight into why a certain prediction is made. Therefore, interpretability must be a key factor when choosing a model for a clinical prediction task, and we explore the level of interpretation given by the models compared to their predictive performance.
Aaron N. Richter, Taghi M. Khoshgoftaar
ICTAI2
2018 A review of statistical and machine learning methods for modeling cancer risk using structured clinical data
abstract
Advancements are constantly being made in oncology, improving prevention and treatment of cancers. To help reduce the impact and deadliness of cancers, they must be detected early. Additionally, there is a risk of cancers recurring after potentially curative treatments are performed. Predictive models can be built using historical patient data to model the characteristics of patients that developed cancer or relapsed. These models can then be deployed into clinical settings to determine if new patients are at high risk for cancer development or recurrence. For large-scale predictive models to be built, structured data must be captured for a wide range of diverse patients. This paper explores current methods for building cancer risk models using structured clinical patient data. Trends in statistical and machine learning techniques are explored, and gaps are identified for future research. The field of cancer risk prediction is a high-impact one, and research must continue for these models to be embraced for clinical decision support of both practitioners and patients.
Aaron N. Richter, Taghi M. Khoshgoftaar
Artif. Intell. Medicine2
2017 Predicting sentinel node status in melanoma from a real-world EHR dataset
abstract
Melanoma is the fastest growing cancer worldwide, and 1 in 50 Americans will develop it in their lifetime. Sentinel lymph node (SLN) metastasis is one of the most important prognostic indicators for melanoma survival. We present several machine learning models for predicting SLN metastasis using data from a real-world dermatology electronic health record (EHR) system. The class label is the result of a sentinel lymph node biopsy, an elective procedure that can be performed for newly-diagnosed melanoma patients to determine if there is metastasis in the nearest lymph node. We show that a simple model, using solely Breslow thickness, can achieve predictive performance (AUC=0.769) comparable to a logistic regression model using 5 features (AUC=0.772, p=0.518). Current clinical recommendations are to perform a biopsy for patients with melanomas thicker than 1mm, however, when applying this 1mm threshold to the simple thickness model, it achieves 0% sensitivity for melanomas2=0.934), and that machine learning models can effectively detect thin melanomas that warrant an SLN biopsy.
Aaron N. Richter, Taghi M. Khoshgoftaar
BIBM2
2017 Medicare Fraud Detection Using Machine Learning Methods
abstract
Healthcare is an integral component in people's lives, especially for the rising elderly population, and must be affordable. Medicare is one such healthcare program. Claims fraud is a major contributor to increased healthcare costs, but its impact can be lessened through fraud detection. In this paper, we compare several machine learning methods to detect Medicare fraud. We perform a comparative study with supervised, unsupervised, and hybrid machine learning approaches using four performance metrics and class imbalance reduction via oversampling and an 80-20 undersampling method. We group the 2015 Medicare data into provider types, with fraud labels from the List of Excluded Individuals/Entities database. Our results show that the successful detection of fraudulent providers is possible, with the 80-20 sampling method demonstrating the best performance across the learners. Furthermore, supervised methods performed better than unsupervised or hybrid methods, but these results varied based on the class imbalance sampling technique and provider type.
Richard A. Bauder, Taghi M. Khoshgoftaar
ICMLA2
2017 Comparing Transfer Learning and Traditional Learning Under Domain Class Imbalance
abstract
Transfer learning is a subclass of machine learning, which uses training data (source) drawn from a different domain than that of the testing data (target). A transfer learning environment is characterized by the unavailability of labeled data from the target domain, due to data being rare or too expensive to obtain. However, there exists abundant labeled data from a different, but similar domain. These two domains are likely to have different distribution characteristics. Transfer learning algorithms attempt to align the distribution characteristics of the source and target domains to create high-performance classifiers. This paper provides comparative performance analysis between stateof- the-art transfer learning algorithms and traditional machine learning algorithms under the domain class imbalance condition. The domain class imbalance condition is characterized by the source and target domains having different class probabilities, which can create marginal distribution differences between the source and target data. Statistical analysis is provided to show the significance of the results.
Karl R. Weiss, Taghi M. Khoshgoftaar
ICMLA2
2017 Training Convolutional Networks on Truncated Text
abstract
Classifiers trained using deep neural networks have been shown to be effective for a wide variety of classification tasks including text sentiment. One such approach is to use convolutional neural networks to learn from character-level representations of documents. This approach is appealing as all feature engineering, extraction or selection is performed by the neural network, automatically generating high level abstract representations of the data. With character-level learning, network topology is dependent on document length as this determines the input shape for the network. In this paper, we investigate how limiting the number of characters used from a document impacts performance and compare neural network performance against a Multinomial Naive Bayes baseline. Our results show the required number of characters is linked to the document domain and neural network performance exceeds that of Multinomial Naive Bayes; however, too low of a number results in performance degradation.
Joseph D. Prusa, Taghi M. Khoshgoftaar
ICTAI2
2017 Evaluation of Transfer Learning Algorithms Using Different Base Learners
abstract
In the field of supervised machine learning, a transfer learning environment is defined as the training data having different distribution characteristics than the testing data. This is due to the lack of available labeled data for the domain of interest, which prompts an alternate domain to be used as the training data. Because there is insufficient labeled data from the domain of interest, validation techniques cannot be reliably used for the algorithm selection process in a transfer learning environment. A transfer learning algorithm is typically comprised of a domain adaptation step followed by a learning step. The learning step is usually implemented using a traditional machine learning algorithm. In this paper, we examine and analyze the impact that the traditional machine learning algorithm (the learning step) has on the overall performance of a transfer learning algorithm. Using the transfer learning test framework, we test five state-of-the-art transfer learning algorithms coupled with seven different traditional learning algorithms for a total of 35 unique transfer learning algorithms. For our experiment, no labeled data from the domain of interest is available for the training process. Since validation techniques cannot be reliably used for the algorithm selection process in a transfer learning environment, it is important for machine learning researchers and practitioners to understand the impact of a traditional machine learner on the overall performance of a transfer learning algorithm.
Karl R. Weiss, Taghi M. Khoshgoftaar
ICTAI2
2016 A Probabilistic Programming Approach for Outlier Detection in Healthcare Claims
abstract
Healthcare is an integral component in people's lives, especially for the rising elderly population. Medicare is one such healthcare program that provides for the needs of the elderly. It is imperative that these healthcare programs are affordable, but this is not always the case. Out of the many possible factors for the rising cost of healthcare, claims fraud is a major contributor, but its impact can be lessened through effective fraud detection. We propose a general outlier detection model, based on Bayesian inference, using probabilistic programming. Our model provides probability distributions rather than just point values, as with most common outlier detection methods. Credible intervals are also generated to further enhance confidence that the detected outliers should in fact be considered outliers. Two case studies are presented demonstrating our model's effectiveness in detecting outliers. The first case study uses temperature data in order to provide a clear comparison of several outlier detection techniques. The second case study uses a Medicare dataset to showcase our proposed outlier detection model. Our results show that the successful detection of outliers, which indicate possible fraudulent activities, can provide effective and meaningful results for further investigation within medical specialties or by using real-world, medical provider fraud investigation cases.
Richard A. Bauder, Taghi M. Khoshgoftaar
ICMLA2
2016 An Investigation of Ensemble Techniques for Detection of Spam Reviews
abstract
Whether purchasing a product or searching for a new doctor, consumers often turn to online reviews for recommendations. Determining whether reviews are truthful is imperative to the consumer, as to not get misled by false recommendations. Unfortunately, it is often difficult, or impossible, for humans to ascertain the validity of a review through reading the text, however, studies have shown machine learning methods perform well for detecting untruthful reviews. Previously, no studies have examined the effects of ensemble learners on the detection of untruthful reviews, despite these techniques being effective in related text classification domains. We seek to inform other researchers of the effects of ensemble techniques on the detection of spam reviews. To this aim, we evaluate four classifiers and three ensemble techniques using those four classifiers as base learners. We compare the results of Multinomial Naïve Bayes, C4.5, Logistic Regression, Support Vector Machine, Random Forest with 100, 250, and 500 trees, and Boosting and Bagging using the base learners. We found that none of the ensemble techniques tested were able to significantly improve review spam detection over standard Multinomial Naïve Bayes and thus, are not worth the computational expense they inflict.
Brian Heredia, Taghi M. Khoshgoftaar, Joseph D. Prusa, Michael Crawford
ICMLA2
2016 Investigating Transfer Learners for Robustness to Domain Class Imbalance
abstract
A transfer learning environment is characterized by a machine learning algorithm being trained with data from one domain (the source domain) and being tested on data from a different domain (the target domain). In a transfer learning scenario, the class probability of the source domain may be different from the class probability of the target domain, which is referred to as "domain class imbalance". Domain class imbalance is different from "class imbalance". Class imbalance refers to the condition of a single domain having unequal class probabilities. In traditional machine learning, the training and testing data are drawn from a single domain. The effects of class imbalance in traditional machine learning are well studied, however, the issue of domain class imbalance in the field of transfer learning has received little research attention. This paper provides a comparative performance test of state-of-the-art transfer learning algorithms, using a wide-range of domain class imbalance combinations. A detailed discussion on the relative performances of the different algorithms, with statistical validation, is presented for the different domain class imbalance scenarios.
Karl R. Weiss, Taghi M. Khoshgoftaar
ICMLA2
2016 Predicting Medical Provider Specialties to Detect Anomalous Insurance Claims
abstract
The healthcare industry is a complex system with many moving parts. One issue in this field is the misuse of medical insurance systems, such as Medicare. In this paper, we build a machine learning model to detect when physicians exhibit anomalous behavior in their medical insurance claims. This new research has the potential to give some insight in determining if, and when, physicians are acting outside the norm of their respective specialty, which could indicate misuse, fraud, or lack of knowledge around billing procedures. We use a publicly available procedure billing dataset, released by the U. S. Medicare system. Due to the large size of the dataset, we sampled the dataset to include all physicians practicing within one state only. The model uses the multinomial Naïve Bayes algorithm and is evaluated by calculating precision, recall, and Fscore with 5-fold cross-validation. The model is able to successfully predict several classes of physicians with an F-score over 0.9. These results show that it is possible to effectively use machine learning in a novel way to classify physicians into their respective fields solely using the procedures they bill for. This research provides a model that can identify physicians who are potentially misusing insurance systems for further investigation.
Richard A. Bauder, Taghi M. Khoshgoftaar, Aaron N. Richter, Matthew Herland
ICTAI2
2016 An Investigation of Transfer Learning and Traditional Machine Learning Algorithms
abstract
Previous research focusing on the evaluation of transfer learning algorithms has predominantly used real-world datasets to measure an algorithm's performance. A test with a real-world dataset exposes an algorithm to a single instance of distribution difference between the training (source) and test (target) datasets. These previous works have not measured performance over a wide-range of source and target distribution differences. We propose to use a test framework that creates many source and target datasets from a single base dataset, representing a diverse-range of distribution differences. These datasets will be used as a stress test to measure an algorithm's performance. The stress test process will measure and compare different transfer learning algorithms and traditional learning algorithms. The unique contributions of this paper, with respect to transfer learning, are defining a test framework, defining multiple distortion profiles, defining a stress test suite, and the evaluation and comparison of different transfer learning and traditional machine learning algorithms over a wide-range of distributions.
Karl R. Weiss, Taghi M. Khoshgoftaar
ICTAI2
2016 The improved grey model based on particle swarm optimization algorithm for time series prediction
Kewen Li 0002, Jiannan Zhai, Taghi M. Khoshgoftaar, Timing Li
Eng. Appl. Artif. Intell.4
2015 Does the Inclusion of Data Sampling Improve the Performance of Boosting Algorithms on Imbalanced Bioinformatics Data?
abstract
Bioinformatics datasets contain many challenging characteristics, such as class imbalance, which adversely impacts the performance of supervised classification models built on these datasets. Techniques such as ensemble learning and data sampling from the domain of data mining can be deployed to alleviate the problem and to improve the classification performance. In this study, we sought to seek whether inclusion of data sampling within the ensemble framework can further improve the performance of classification models. To this end, we performed an experimental study using two newly hybrid ensemble techniques, one integrates feature selection within the boosting process and the other incorporates random under-sampling followed by feature selection within the boosting framework, two learners, three forms of feature rankers, and four feature subset sizes on 15 highly imbalanced bioinformatics datasets. Our results and statistical analysis demonstrate that the difference between the two boosting methods is statistically insignificant. Therefore, as the inclusion of data sampling has no significant positive effect on the performance of ensemble classifiers, it is not required to achieve maximum classification performance. To our knowledge, this is the first empirical study that examined the effects of data sampling, random under-sampling, to enhance classification performance of boosting algorithm for highly imbalanced bioinformatics data.
Alireza Fazelpour, Taghi M. Khoshgoftaar, David J. Dittman, Amri Napolitano
ICMLA2
2015 Investigating New Bootstrapping Approaches of Bagging Classifiers to Account for Class Imbalance in Bioinformatics Datasets
abstract
One major challenge posed by bioinformatics datasets is class imbalance which occurs when one class has many more instances than the other class(es). Its undesirable effect on the classification performance is compounded with the fact that, in general, the class with fewer instances is the class of interest. Bagging has been utilized by practitioners in the field to overcome the challenge of class imbalance and to improve the classification performance. Our motivation for this study is to investigate whether changes to the bootstrapping step of bagging classifiers can further improve their performance. Specifically, these modifications to the bootstrapping process take into account the membership of the classes. We performed an extensive empirical study utilizing four bootstrap approaches within bagging framework using three feature rankers along with four feature subset sizes and two base classifiers across 15 imbalanced bioinformatics datasets. Three of these bootstrap approaches were proposed and implemented by our research team for this study. Our results show that all new approaches improve performance over the classic bootstrap approach, with balanced bagging having the highest performance, however, observed increases in performance are not statistically significant. We recommend the balanced bootstrap approach because it shows the most improvement, in terms of frequency of having the highest performance, and it generates fully balanced bootstrap datasets that can account for the class imbalance problem. The uniqueness of this paper is proposing and implementing the three innovative bootstrapping approaches to examine the effects of these bootstrapping processes against the classic one on the performance of bagging classifiers in the domain of bioinformatics.
Alireza Fazelpour, Taghi M. Khoshgoftaar, David J. Dittman, Amri Napolitano
ICMLA2
2015 Detection of SSH Brute Force Attacks Using Aggregated Netflow Data
abstract
The SSH Brute force attack is one of the most prevalent attacks in computer networks. These attacks aim to gain ineligible access to users' accounts by trying plenty of different password combinations. The detection of this type of attack at the network level can overcome the scalability issue of host-based detection methods. In this paper, we provide a machine learning approach for the detection of SSH brute force attacks at the network level. Since extracting discriminative features for any machine learning task is a fundamental step, we explain the process of extracting discriminative features for the detection of brute force attacks. We incorporate domain knowledge about SSH brute force attacks as well as the analysis of a representative collection of the data to define the features. We collected real SSH traffic from a campus network. We also generated some failed login data that a legitimate user who has forgotten his/her password can produce as normal traffic that can be similar to the SSH brute force attack traffic. Our inspection on the collected brute force Netflow data and the manually produced SSH failed login data showed that the Netflow features are not discriminative enough to discern brute force traffic from the failed login traffic produced by a legitimate user. We introduced an aggregation of Netflows to extract the proper features for building machine learning models. Our results show that the models built upon these features provide excellent performances for the detection of brute force attacks.
Maryam M. Najafabadi, Taghi M. Khoshgoftaar, Chad Calvert, Clifford Kemp
ICMLA2
2015 Utilizing Ensemble, Data Sampling and Feature Selection Techniques for Improving Classification Performance on Tweet Sentiment Data
abstract
Sentiment analysis of tweets is a popular method of opinion mining social media. Many machine learning techniques exist that can improve the performance of classifiers trained to determine the sentiment or emotional polarity of a tweet, however, they are designed with different objectives and it is unclear which techniques are most beneficial. Additionally, these techniques may behave differently depending on quality of data issues, such as class imbalance, a common problem when using real world data. In an effort to determine which techniques are more important, we tested 12 techniques consisting of: eight feature selection techniques, bagging, boosting and data sampling with two post sampling class ratios. Using five base learners, we compare these techniques against each other and each base learners with no additional technique. We train and test each classifier on a balanced dataset and two imbalanced datasets with different class ratios. Additionally, we conduct statistical tests to determine if the differences observed between techniques are significant. Our results show that bagging and seven of the eight feature selection techniques significantly improve performance (compared to using no technique) on all three datasets, while boosting and data sampling are less beneficial for imbalanced tweet sentiment data. To the best of our knowledge, this is the first study comparing these three types of techniques on tweet sentiment data and the first to show that feature selection and ensemble techniques perform better than data sampling on tweet sentiment data.
Joseph D. Prusa, Taghi M. Khoshgoftaar, Amri Napolitano
ICMLA2
2015 The Effect of Dataset Size on Training Tweet Sentiment Classifiers
abstract
Using automated methods of labeling tweet sentiment, large volumes of tweets can be labeled and used to train classifiers. Millions of tweets could be used to train a classifier, however, doing so is computationally expensive. Thus, it is valuable to establish how many tweets should be utilized to train a classifier, since using additional instances with no gain in performance is a waste of resources. In this study, we seek to find out how many tweets are needed before no significant improvements are observed for sentiment analysis when adding additional instances. We train and evaluate classifiers using C4.5 decision tree, Naïve Bayes, 5 Nearest Neighbor and Radial Basis Function Network, with seven datasets varying from 1000 to 243,000 instances. Models are trained using four runs of 5-fold cross validation. Additionally, we conduct statistical tests to verify our observations and examine the impact of limiting features using frequency. All learners were found to improve with dataset size, with Naïve Bayes being the best performing learner. We found that Naïve Bayes did not significantly benefit from using more than 81,000 instances. To the best of our knowledge, this is the first study to investigate how learners scale in respect to dataset size with results verified using statistical tests and multiple models trained for each learner and dataset size. Additionally, we investigated using feature frequency to greatly reduce data grid size with either a small increase or decrease in classifier performance depending on choice of learner.
Joseph D. Prusa, Taghi M. Khoshgoftaar, Naeem Seliya
ICMLA2
2015 Ensemble vs. Data Sampling: Which Option Is Best Suited to Improve Classification Performance of Imbalanced Bioinformatics Data?
abstract
Bioinformatics datasets contain challenging characteristics, such as class imbalance that occurs when one class has many more instances than the other class(es). These challenges make the task of classification much more subtle for practitioners and researchers in the field. Fortunately, there are tools, such as ensemble learning and data sampling methods that can be applied to overcome these problems and improve the performance of supervised classification models. Our motivation for this study is to investigate which option is well suited to tackle this significant challenge for bioinformatics data. Our literature survey shows that no previous work has conducted such an extensive study to examine whether ensemble learning or data sampling is best suited for imbalanced gene expression data. To this end, we carried out an extensive experimental study using five ensemble classification methods, four other classification methods with random under-sampling, three feature rankers along with four feature subset sizes across 15 highly imbalanced bioinformatics datasets. Our results along with statistical analysis confirm that ensemble learning methods in general outperform data sampling techniques in improving classification results. Furthermore, Select-Bagging with Naïve Bayes (NB) followed by Random Forest are the top two performing ensemble techniques. Based on these results, we recommend either Select-Bagging with NB or Random Forest with 100 trees (RF100) for imbalanced datasets. However, RF100, unlike Select-Bagging, does not rely on choice of the base learner.
Taghi M. Khoshgoftaar, Alireza Fazelpour, David J. Dittman, Amri Napolitano
ICTAI1
2015 Using Feature Selection in Combination with Ensemble Learning Techniques to Improve Tweet Sentiment Classification Performance
abstract
Performing sentiment analysis of tweets by training a classifier is a challenging and complex task, requiring that the classifier can correctly and reliably identify the emotional polarity of a tweet. Poor data quality, due to class imbalance or mislabeled instances, may negatively impact classification performance. Ensemble learning techniques combine multiple models in an attempt to improve classification performance, especially on poor quality or imbalanced data, however, these techniques do not address the concern of high dimensionality present in tweets sentiment data and may require a prohibitive amount of resources to train on high dimensional data. This work addresses these issues by studying bagging and boosting combined with feature selection. These two techniques are denoted as Select-Bagging and Select-Boost, and seek to address both poor data quality and high dimensionality. We compare the performance of Select-Bagging and Select-Boost against feature selection alone. These techniques are tested with four base learners, two datasets and ten feature subset sizes. Our results show that Select-Boost offers the highest performance, is significantly better than using no ensemble technique, and is significantly better than Select-Bagging for most learners on both datasets. To the best of our knowledge, this is the first study to focus on the effects of using ensemble learning in combination with feature selection for the purpose of tweet sentiment classification.
Joseph D. Prusa, Taghi M. Khoshgoftaar, Amri Napolitano
ICTAI2
2015 Efficient Modeling of User-Entity Preference in Big Social Networks
abstract
Data generated by social media are frequently leveraged to build machine learning models that can accurately profile human behavior and sentiment. Twitter is a readily available source of population data that can be collected and used by any organization. Therefore, accurate machine learning models must be created to learn from this user-generated content. In this paper, we explore the task of classifying a user's preference towards a specific entity. Particularly, we study the accuracy of classification models as an increasing number of tweets (status posts) per user is provided to the models. New users and tweets are constantly being created, warranting the use of techniques to reduce the size of data needed for machine learning algorithms. We find that there is a diminishing return on model performance as the number of tweets per user is increased, and identify a threshold where adding more tweets per user does not result in statistically better performance. Utilizing this threshold, as opposed to the maximum amount of tweets per user, data collection time is reduced by 80% while dataset size is reduced by 75%.
Aaron N. Richter, Michael Crawford, Brian Heredia, Taghi M. Khoshgoftaar
ICTAI4
2015 Combining Feature Subset Selection and Data Sampling for Coping with Highly Imbalanced Software Data
abstract
In the software quality modeling process, many practitioners often ignore problems such as high dimensionality and class imbalance that exist in data repositories.They directly use the available set of software metrics to build classification models without regard to the condition of the underlying software measurement data, leading to a decline in prediction performance and extension of training time.In this study, we propose an approach, in which feature selection is combined with data sampling, to overcome these problems.Feature selection is a process of choosing a subset of relevant features so that the quality of prediction models can be maintained or improved.Data sampling seeks a more balanced dataset through the addition or removal of instances.Three different approaches would be produced when combing these two techniques: 1-sampling performed prior to feature selection, but retaining the unsampled data instances; 2-sampling performed prior to feature selection, retaining the sampled data instances; 3-sampling performed after feature selection.The empirical study was carried out on six datasets from a real-world software system.We employed one filter-based (no learning algorithm involved in the selection process) feature subset selection technique called correlationbased feature selection combined with the random undersampling method.The results demonstrate that sampling performed prior to feature selection, but retaining the unsampled data instances (Approach 1) performs better than the other two approaches.
Kehan Gao, Taghi M. Khoshgoftaar, Amri Napolitano
SEKE2
2015 Stability of Three Forms of Feature Selection Methods on Software Engineering Data
abstract
One of the major challenges when working with software metrics datasets is that some metrics may be redundant or irrelevant to software defect prediction.This may be addressed using feature (metric) selection, which chooses an appropriate subset of features for use in downstream computation.There are three major forms of feature selection: filter-based feature rankers, which uses statistical measures to assign a score to each feature and present the user with a ranked list; filter-based subset evaluation, which uses statistical measures on feature subsets to find the best choice; and wrapper-based subset selection, which builds classification models using different subsets to find the one which maximizes performance.Software practitioners are interested in which feature selection methods are best at providing the most stable feature subset in the face of changes to the data (here, the addition or removal of instances).In this study we select feature subsets using fifteen feature selection methods and then use our newly proposed Average Pairwise Tanimoto Index (APTI) to evaluate the stability of feature selection methods.We evaluate the stability of feature selection methods on a pair of subsamples generated by fixed-overlap partitions algorithm.Four different levels of overlap are considered in this study.Four software metric datasets from a real-world software project are used in this study.Results demonstrate that ReliefF (RF) is the most stable feature selection method and wrapper based feature subset selection shows least stability.In addition, as the overlap of partitions increased, the stability of the feature selection strategies increased.
Huanjing Wang, Taghi M. Khoshgoftaar, Amri Napolitano
SEKE2
2015 Investigating Two Approaches for Adding Feature Ranking to Sampled Ensemble Learning for Software Quality Estimation
abstract
Defect prediction is very challenging in software development practice. Classification models are useful tools that can help for such prediction. Classification models can classify program modules into quality-based classes, e.g. fault-prone (fp) or not-fault-prone (nfp). This facilitates the allocation of limited project resources. For example, more resources are assigned to program modules that are of poor quality or likely to have a high number of faults based on the classification. However, two main problems, high dimensionality and class imbalance, affect the quality of training datasets and therefore classification models. Feature selection and data sampling are often used to overcome these problems. Feature selection is a process of choosing the most important attributes from the original dataset. Data sampling alters the dataset to change its balance level. Another technique, called boosting (building multiple models, with each model tuned to work better on instances misclassified by previous models), is found to also be effective for resolving the class imbalance problem. In this study, we investigate an approach for combining feature selection with this ensemble learning (boosting) process. We focus on two different scenarios: feature selection performed prior to the boosting process and feature selection performed inside the boosting process. Ten individual base feature ranking techniques, as well as an ensemble ranker based on the ten, are examined and compared over the two scenarios. We also employ the boosting algorithm to construct classification models without performing feature selection and use the results as the baseline for further comparison. The experimental results demonstrate that feature selection is important and needed prior to the learning process. In addition, the ensemble feature ranking method generally has better or similar performance than the average of the base ranking techniques, and more importantly, the ensemble method exhibits better robustness than most base ranking techniques. As for the two scenarios, the results show that applying feature selection inside boosting performs better than using feature selection prior to boosting.
Kehan Gao, Taghi M. Khoshgoftaar, Amri Napolitano
Int. J. Softw. Eng. Knowl. Eng.2
2015 Aggregating Data Sampling with Feature Subset Selection to Address Skewed Software Defect Data
abstract
Defect prediction is an important process activity frequently used for improving the quality and reliability of software products. Defect prediction results provide a list of fault-prone modules which are necessary in helping project managers better utilize valuable project resources. In the software quality modeling process, high dimensionality and class imbalance are the two potential problems that may exist in data repositories. In this study, we investigate three data preprocessing approaches, in which feature selection is combined with data sampling, to overcome these problems in the context of software quality estimation. These three approaches are: Approach 1 — sampling performed prior to feature selection, but retaining the unsampled data instances; Approach 2 — sampling performed prior to feature selection, retaining the sampled data instances; and Approach 3 — sampling performed after feature selection. A comparative investigation is presented for evaluating the three approaches. In the experiments, we employed three sampling methods (random undersampling, random oversampling, and synthetic minority oversampling), each combined with a filter-based feature subset selection technique called correlation-based feature selection. We built the defect prediction models using five common classification algorithms. The case study was based on software metrics and defect data collected from multiple releases of a real-world software system. The results demonstrated that the type of sampling methods used in data preprocessing significantly affected the performance of the combination approaches. It was found that when the random undersampling technique was used, Approach 1 performed better than the other two approaches. However, when the feature selection technique was used in conjunction with an oversampling method (random oversampling or synthetic minority oversampling), we strongly recommended Approach 3.
Kehan Gao, Taghi M. Khoshgoftaar, Amri Napolitano
Int. J. Softw. Eng. Knowl. Eng.2
2015 An Empirical Investigation on Wrapper-Based Feature Selection for Predicting Software Quality
abstract
The basic measurements for software quality control and management are the various project and software metrics collected at various states of a software development life cycle. The software metrics may not all be relevant for predicting the fault proneness of software components, modules, or releases. Thus creating the need for the use of feature (software metric) selection. The goal of feature selection is to find a minimum subset of attributes that can characterize the underlying data with results as well as, or even better than the original data when all available features are considered. As an example of inter-disciplinary research (between data science and software engineering), this study is unique in presenting a large comparative study of wrapper-based feature (or attribute) selection techniques for building defect predictors. In this paper, we investigated thirty wrapper-based feature selection methods to remove irrelevant and redundant software metrics used for building defect predictors. In this study, these thirty wrappers vary based on the choice of search method (Best First or Greedy Stepwise), leaner (Naïve Bayes, Support Vector Machine, and Logistic Regression), and performance metric (Overall Accuracy, Area Under ROC (Receiver Operating Characteristic) Curve, Area Under the Precision-Recall Curve, Best Geometric Mean, and Best Arithmetic Mean) used in the defect prediction model evaluation process. The models are trained using the three learners and evaluated using the five performance metrics. The case study is based on software metrics and defect data collected from a real world software project. The results demonstrate that Best Arithmetic Mean is the best performance metric used within the wrapper. Naïve Bayes performed significantly better than Logistic Regression and Support Vector Machine as a wrapper learner on slightly and less imbalanced datasets. We also recommend Greedy Stepwise as a search method for wrappers. Moreover, comparing to models built with full datasets, the performances of defect prediction models can be improved when metric subsets are selected through a wrapper subset selector.
Huanjing Wang, Taghi M. Khoshgoftaar, Amri Napolitano
Int. J. Softw. Eng. Knowl. Eng.2
2015 On the Stability of Feature Selection Methods in Software Quality Prediction: An Empirical Investigation
abstract
Software quality modeling is the process of using software metrics from previous iterations of development to locate potentially faulty modules in current under-development code. This has become an important part of the software development process, allowing practitioners to focus development efforts where they are most needed. One difficulty encountered in software quality modeling is the problem of high dimensionality, where the number of available software metrics is too large for a classifier to work well. In this case, many of the metrics may be redundant or irrelevant to defect prediction results, thereby selecting a subset of software metrics that are the best predictors becomes important. This process is called feature (metric) selection. There are three major forms of feature selection: filter-based feature rankers, which uses statistical measures to assign a score to each feature and present the user with a ranked list; filter-based feature subset evaluation, which uses statistical measures on feature subsets to find the best feature subset; and wrapper-based subset selection, which builds classification models using different subsets to find the one which maximizes performance. Software practitioners are interested in which feature selection methods are best at providing the most stable feature subset in the face of changes to the data (here, the addition or removal of instances). In this study we select feature subsets using fifteen feature selection methods and then use our newly proposed Average Pairwise Tanimoto Index (APTI) to evaluate the stability of the feature selection methods. We evaluate the stability of feature selection methods on a pair of subsamples generated by our fixed-overlap partitions algorithm. Four different levels of overlap are considered in this study. 13 software metric datasets from two real-world software projects are used in this study. Results demonstrate that ReliefF (RF) is the most stable feature selection method and wrapper based feature subset selection shows least stability. In addition, as the overlap of partitions increased, the stability of the feature selection strategies increased.
Huanjing Wang, Taghi M. Khoshgoftaar, Naeem Seliya
Int. J. Softw. Eng. Knowl. Eng.2
2014 Selecting the Appropriate Data Sampling Approach for Imbalanced and High-Dimensional Bioinformatics Datasets
abstract
One of the more prevalent problems when working with bioinformatics datasets is class imbalance, when there are more instances in one class compared to the other class (es). This problem is made worse because frequently, the class of interest is also the minority class. A possible solution is data sampling, a powerful tool for combating class imbalance by adding or removing instances to make the dataset more balanced. In addition to the choice of including data sampling, one of the most important decisions when applying data sampling is what the final class ratio should be. Commonly, the final class ratio when data sampling is applied is 50:50, however it is an open question whether other ratios are more appropriate for certain imbalanced datasets (all datasets in this paper have 25.16% minority instances or less) where a 50:50 ratio requires extreme modification to the dataset. In this work we compare six different data sampling approaches (feature selection with the pair wise combinations of three data sampling techniques and two final class ratios) with feature selection without data sampling with the goal of determining if the inclusion of data sampling is beneficial and if so, what should be the final class ratio. In order to test the six data sampling approaches and feature selection alone thoroughly, we utilize seven imbalanced and high-dimensional datasets, three feature selection techniques, and six classifiers. Our results show that for a majority of scenarios, random under sampling along with either 35:65 or 50:50 is the best data sampling approach. Statistical analysis shows that there is no significant difference between the data sampling approaches. However, despite this, we still recommend using random under sampling along with 35:65 as the final class ratio. This is because of the frequency of both random under sampling and 35:65 being the most frequent top performing data sampling technique and class ratio respectively. Additionally, 35:65 will have fewer negative impacts than 50:50 (less data loss or over fitting, which makes it a better choice if all other factors are equal) and random under sampling is more computationally efficient than any other form of sampling, including "no sampling" (both by not requiring any internal calculations and by producing a reduced, easier-to-work-with dataset). To our knowledge, this is the most comprehensive work which focuses on the choice of the inclusion and implementation of data sampling with different final class ratios on bioinformatics datasets which exhibit such large levels of class imbalance.
David J. Dittman, Taghi M. Khoshgoftaar, Amri Napolitano
BIBE2
2014 Select-Bagging: Effectively Combining Gene Selection and Bagging for Balanced Bioinformatics Data
abstract
Bioinformatics datasets have historically been difficult to work with. However, within machine learning, there is a potentially effective tool to combat such problems: ensemble learning. Ensemble learning generates a series of models and combines their results to make a single decision. This process has the benefit of utilizing the power of multiple models but the overhead of having to compute the multiple models. Thus, we must ask whether the benefits outweigh the detriments. In this study, we seek to determine if the ensemble learning technique Select-Bagging improves classification results over feature selection on the training dataset followed by classification (denoted as FS-Classifier in this work) on a series of balanced bioinformatics datasets. We test the two approaches with two filter-based feature rankers, four feature subset sizes and the Naïve Bayes classifier. Our results show that Select-Bagging clearly outperforms FS-Classifier for nearly all scenarios. Subsequent statistical analysis shows that the increase in performance generated by Select-Bagging is statistically significantly better than FS-Classifier. Therefore, we can state that the inclusion of Select-Bagging is beneficial to the classification performance of models built on high-dimensional and balanced bioinformatics datasets and should be implemented. To our knowledge this is the first study which looks at the effectiveness of bagging in conjunction with internal feature selection for balanced bioinformatics datasets.
David J. Dittman, Taghi M. Khoshgoftaar, Amri Napolitano, Alireza Fazelpour
BIBE2
2014 Effects of the Use of Boosting on Classification Performance of Imbalanced Bioinformatics Datasets
abstract
In the domain of bioinformatics, two common problems encountered when analyzing real-world datasets are class imbalance and high dimensionality. Boosting is a technique that can be used to improve classification performance, even in the presence of class imbalance. In addition, data sampling and feature selection are two important preprocessing techniques used to counter the adverse effects of both challenges collectively. In this study, we examine whether the inclusion of boosting along with joint deployment of feature selection and data sampling techniques affect the classification performance of inductive models. To this end, we used two approaches: filter-based feature selection followed by either data sampling (denoted as FS-DS) or a hybrid data sampling and boosting technique entitled RUSBoost (denoted as FRB) which integrates random under sampling within the boosting process. We conducted an extensive experimental study using six high dimensional and imbalanced bioinformatics datasets along with three learners and four feature subset sizes. Our results show that the improvement of classification performance due to boosting depends on the choice of learner used to build the model. We recommend FRB because it outperforms FS-DS for nearly all scenarios. Additionally, our ANOVA analysis shows that the FRB is statistically distinguishable from the FS-DS when using the LR learner. To our knowledge, this is the first study to investigate the effects of boosting along with combined feature selection and data sampling on classification performance of inductive models in the domain of bioinformatics.
Taghi M. Khoshgoftaar, Alireza Fazelpour, David J. Dittman, Amri Napolitano
BIBE1
2014 Machine Learning for Detecting Brute Force Attacks at the Network Level
abstract
The tremendous growth in computer network and Internet usage, combined with the growing number of attacks makes network security a topic of serious concern. One of the most prevalent network attacks that can threaten computers connected to the network is brute force attack. In this work we investigate the use of machine learners for detecting brute force attacks (on the SSH protocol) at the network level. We base our approach on applying machine learning algorithms on a newly generated dataset based upon network flow data collected at the network level. Applying detection at the network level makes the detection approach more scalable. It also provides protection for the hosts who do not have their own protection. The new dataset consists of real-world network data collected from a production network. We use four different classifiers to build brute force attack detection models. The use of different classifiers facilitates a relatively comprehensive study on the effectiveness of machine learners in the detection of brute force attack on the SSH protocol at the network level. Empirical results show that the machine learners were quite successful in detecting the brute force attacks with a high detection rate and low false alarms. We also investigate the effectiveness of using ports as features during the learning process. We provide a detailed analysis of how the models built can change as a result of including or excluding port features.
Maryam M. Najafabadi, Taghi M. Khoshgoftaar, Clifford Kemp, Naeem Seliya, Richard Zuech
BIBE2
2014 Evaluation of Wrapper-Based Feature Selection Using Hard, Moderate, and Easy Bioinformatics Data
abstract
One of the most challenging problems encountered when analyzing real-world gene expression datasets is high dimensionality (overabundance of features/attributes). This large number of features can lead to suboptimal classification performance and increased computation time. Feature selection, whereby only a subset of the original features are used for building a classification model, is the most commonly used technique to counter high dimensionality. One category of feature selection called wrapper-based techniques employ a classifier to directly find the subset of features which performs best. Unfortunately, noise can negatively impact the effectiveness of data mining techniques and subsequently lead to suboptimal results. Class noise in particular has a detrimental effect on the classification performance, making datasets perform poorly across a wide range of classifiers (i.e. Having a high """"difficulty-of-learning.""""). No previous work has examined the effectiveness of wrapper-based feature selection when learning from real world high dimensional gene expression datasets in the context of difficulty-of-learning due to noise. To study this effectiveness, we perform experiments using ten gene expression datasets which was first determined to be easy-to-learn-from then had artificial class noise injected in a controlled fashion creating three levels of difficulty-of-learning (Easy, Moderate, and Hard). Using the Naïve Bayes learner, we perform wrapper feature selection followed by classification, using four classifiers (Naïve Bayes, Multilayer Perceptron, 5-Nearest Neighbor, and Support Vector Machines), and we compare these results to the classification performance without feature selection. The results show that wrapper-based feature selection effectiveness depends on the choice of learner: for Multilayer Perceptron, wrapper selection improved performance compared to not using feature selection, while for Naïve Bayes it slightly reduced performance and for the remaining learners it further reduced performance. Because its performance relative to no feature selection varied depending on the choice of learner, we recommend that wrapper selection be at least considered in future bioinformatics experiments, especially if the goal is gene discovery not classification. Also, as dimensionality reduction techniques are not only useful but necessary for high-dimensional bioinformatics datasets, the no-feature-selection case may not be feasible in practice.
Ahmad Abu Shanab, Taghi M. Khoshgoftaar, Randall Wald
BIBE2
2014 Using Correlation-Based Feature Selection for a Diverse Collection of Bioinformatics Datasets
abstract
The large number of genes found in most gene micro array datasets demands the use of feature selection techniques to alleviate this problem of high-dimensionality. However, the computational cost of filter-based subset evaluation techniques such as Correlation-Based Feature Selection (CFS) has generally limited the use of these techniques to smaller datasets, or at least smaller collections of gene micro array datasets. No previous work has applied CFS to a large and diverse range of bioinformatics datasets. To address this deficit, we employ nine different micro array datasets exhibiting a wide range of characteristics in terms of dataset balance (fraction of instances found in the minority class) and dataset difficulty of learning (overall difficulty of building effective classification models on raw, pre-feature-selection datasets). We also use five classification learners to discover how these perform in conjunction with CFS, along with five performance metrics to give a broad perspective on our results. The results find that CFS can be used to help build effective models, in particular when used with the 5-Nearest Neighbors learner on data that is Easy or Moderate (in terms of difficulty-of-learning) or Balanced (in terms of class distribution). For other types of data, the optimal learner varies, although in most cases the Logistic Regression learner works worst in conjunction with CFS.
Randall Wald, Taghi M. Khoshgoftaar, Amri Napolitano
BIBE2
2014 Network Traffic Prediction Models for Near- and Long-Term Predictions
abstract
The large quantity of data flowing through network equipment demands that effective and efficient models be built to identify whether sessions are healthy or malicious. These models can be complex to build, and may rely on manually-labeled data. As a result, it is desirable to update or rebuild these models as rarely as possible without impairing classification performance. In this work, we consider the Kyoto dataset, training models on a single day's worth of data and testing these models under two circumstances: using 12 datasets gathered between six and twelve months after the training date, and using 9 datasets gathered between 18 and 19 months after the training date. In all cases, we apply three feature rankers (in addition to no feature ranking) and consider four classification models. We find that the results for the "near-term" 12 datasets are similar to those from the "long-term" 9 datasets, demonstrating that once a model has been built, it can potentially be used for over a year afterwards.
Randall Wald, Taghi M. Khoshgoftaar, Richard Zuech, Amri Napolitano
BIBE2
2014 A Session Based Approach for Aggregating Network Traffic Data - The SANTA Dataset
abstract
This paper compares and contrasts the most widely used network security datasets, evaluating their efficacy in providing a benchmark for intrusion and anomaly detection systems. The antiquated nature of some of the most widely used datasets along with their inadequacies is examined and used as a basis for discussion of a new approach to analyzing network traffic data. Live network traffic is collected that consists of real normal traffic and both real and penetration testing attack data. Attack data is then inspected and labeled by means of manual analysis. While network attacks and anomaly features vary widely, they share some commonalities that are examined here. Among these are: self-similarity convergence, periodicity, and repetition. Further, the knowledge inherent in the definition of network boundaries and advertised services can provide crucial context that allows the network analyst to consider self-aware attributes when examining network traffic sessions. To these ends the Session Aggregation for Network Traffic Analysis (SANTA) dataset is proposed. The motivation and the methodology of collection, aggregation and evaluation of the raw data are presented, as well as the conceptualization of the SANTA attributes and advantages provided by this approach.
Charles Wheelus, Taghi M. Khoshgoftaar, Richard Zuech, Maryam M. Najafabadi
BIBE2
2014 Comparing Two Approaches for Adding Feature Ranking to Sampled Ensemble Learning for Software Quality Estimation
Kehan Gao, Taghi M. Khoshgoftaar, Amri Napolitano
SEKE2
2014 Choosing the Best Classification Performance Metric for Wrapper-based Software Metric Selection for Defect Prediction
Huanjing Wang, Taghi M. Khoshgoftaar, Amri Napolitano
SEKE2
2014 The Use of Ensemble-Based Data Preprocessing Techniques for Software Defect Prediction
abstract
Software defect prediction models that use software metrics such as code-level measurements and defect data to build classification models are useful tools for identifying potentially-problematic program modules. Effectiveness of detecting such modules is affected by the software measurements used, making data preprocessing an important step during software quality prediction. Generally, there are two problems affecting software measurement data: high dimensionality (where a training dataset has an extremely large number of independent attributes, or features) and class imbalance (where a training dataset has one class with relatively many more members than the other class). In this paper, we present a novel form of ensemble learning based on boosting that incorporates data sampling to alleviate class imbalance and feature (software metric) selection to address high dimensionality. As we adopt two different sampling methods (Random Undersampling (RUS) and Synthetic Minority Oversampling (SMOTE)) in the technique, we have two forms of our new ensemble-based approach: selectRUSBoost and selectSMOTEBoost. To evaluate the effectiveness of these new techniques, we apply them to two groups of datasets from two real-world software systems. In the experiments, four learners and nine feature selection techniques are employed to build our models. We also consider versions of the technique which do not incorporate feature selection, and compare all four techniques (the two different ensemble-based approaches which utilize feature selection and the two versions which use sampling only). The experimental results demonstrate that selectRUSBoost is generally more effective in improving defect prediction performance than selectSMOTEBoost, and that the techniques with feature selection do help for getting better prediction than the techniques without feature selection.
Kehan Gao, Taghi M. Khoshgoftaar, Amri Napolitano
Int. J. Softw. Eng. Knowl. Eng.2
2014 Incomplete-case nearest neighbor imputation in software measurement data
Jason Van Hulse, Taghi M. Khoshgoftaar
Inf. Sci.2
2014 Software quality assessment using a multi-strategy classifier
Taghi M. Khoshgoftaar, Yudong Xiao, Kehan Gao
Inf. Sci.1
2014 An empirical study of the classification performance of learners on imbalanced and noisy software quality data
Chris Seiffert, Taghi M. Khoshgoftaar, Jason Van Hulse, Andres Folleco
Inf. Sci.2
2014 System regression test planning with a fuzzy expert system
Zhiwei Xu 0001, Kehan Gao, Taghi M. Khoshgoftaar, Naeem Seliya
Inf. Sci.3
2013 Simplifying the Utilization of Machine Learning Techniques for Bioinformatics
abstract
The domain of bioinformatics has a number of challenges such as handling datasets which exhibit extreme levels of high dimensionality (large number of features per sample) and datasets which are particularly difficult to work with. These datasets contain many pieces of data (features) which are irrelevant and redundant to the problem being studied, which makes analysis quite difficult. However, techniques from the domain of machine learning and data mining are well suited to combating these difficulties. Techniques like feature selection (choosing an optimal subset of features for subsequent analysis by removing irrelevant or redundant features) and classifiers (used to build inductive models in order to classify unknown instances) can assist researchers in working with such difficult datasets. Unfortunately, many practitioners of bioinformatics do not have the machine learning knowledge to choose the correct techniques in order to achieve good classification results. If the choices could be simplified or predetermined then it would be easier to apply the techniques. This study is a comprehensive analysis of machine learning techniques on twenty-five bioinformatics datasets using six classifiers, and twenty-four feature rankers. We analyzed the factors at each of four feature subset sizes chosen for being large enough to be effective in creating inductive models but small enough to be of use for further research. Our results shows that Random Forest with 100 trees is the top performing classifier and that the choice of feature ranker is of little importance as long as feature selection occurs. Statistical analysis confirms our results. By choosing these parameters, machine learning techniques are more accessible to bioinformatics.
David J. Dittman, Taghi M. Khoshgoftaar, Randall Wald, Amri Napolitano
ICMLA (2)2
2013 Improving Software Quality Estimation by Combining Boosting and Feature Selection
abstract
The predictive accuracy of a classification modelis often affected by the quality of training data. However, there are two problems which may affect the quality of the training data: high dimensionality (too many independent attributes in a dataset) and class imbalance (many more instances of one class than the other class in a binary-classification problem). In this study, we present an iterative feature selection approach working with an ensemble learning method to solve both of these problems. The iterative feature selection approach samples the dataset k times and applies feature ranking to each sampled dataset, the k different rankings are then aggregated to create a single feature ranking. The ensemble learning method used is RUSBoost, in which random under sampling(RUS) is integrated into a boosting algorithm. The main purpose of this paper is to investigate the impact of feature selection as well as the RUSBoost approach on the classification performance in the context of software quality prediction. In the experiment, we explore six rankers, each used along with RUS in the iterative feature selection process. Following feature selection, models are built either using a plain learner or byusing the RUSBoost algorithm. We also examine the case of no feature selection and use this as the baseline for comparisons. The experimental results demonstrate that with the exception of one learner, feature selection combined with boosting provides better classification performance than when either is applied alone or when neither are applied.
Kehan Gao, Taghi M. Khoshgoftaar, Amri Napolitano
ICMLA (1)2
2013 Survey of Clinical Data Mining Applications on Big Data in Health Informatics
abstract
Every piece of information learned in human health has the potential to improve the length and quality of life for patients. However, the search for more knowledge in the domain of Health Informatics has led to vast quantities of data, far more than can be easily processed by researchers. Fortunately, Justas this problem of Big Data has begun to challenge advances in Health Informatics, so too can the data mining and machine learning techniques used for the general study of Big Data be brought to bear on the these problems. This paper will discuss recent studies done on Big Data Analytics in the field of Health Informatics which sought to answer various clinical questions, using data acquired from the molecular, tissue, and patient levels of Health Informatics. We also consider what work remains to be done in this area.
Matthew Herland, Taghi M. Khoshgoftaar, Randall Wald
ICMLA (2)2
2013 Contrasting Undersampled Boosting with Internal and External Feature Selection for Patient Response Datasets
abstract
Class imbalance (where one class has many more instances than the other class(es)) and high dimensionality (large number of features per instance) are two prevalent problems that are frequently present in patient response datasets. In addition to these problems, these datasets are notoriously difficult to build effective models from. This paper introduces a new hybrid boosting algorithm named SelectRUSBoost which combines data sampling and feature selection with every iteration of boosting. We test SelectRUSBoost along with RUSBoost combined with external feature selection on a set of five patient response datasets. In addition to the datasets we also utilize two classifiers, three filter-based feature selection techniques, and four feature subset sizes. Our results show that SelectRUSBoost will, with few exceptions, outperform RUSBoost combined with external feature selection. Also, the feature selection technique information gain outperformed the other techniques for all combinations of boosting approach, classifier, and feature subset size, and in addition for this feature selection technique SelectRUSBoost always (without exception) outperformed RUSBoost combined with external selection. Statistical analysis confirmed that SelectRUSBoost gives better performance than RUSBoost combined with external selection. This is the first work which utilizes SelectRUSBoost in a bioinformatics study.
Taghi M. Khoshgoftaar, David J. Dittman, Randall Wald, Amri Napolitano
ICMLA (2)1
2013 Survey of Data Cleansing and Monitoring for Large-Scale Battery Backup Installations
abstract
A continuous supply of electrical power is necessary for many areas of modern life, including industry, healthcare, and telecommunications. Therefore, battery backup systems, which can provide power in the event of emergencies, have become extremely important for many types of industries. Due to the importance of these systems, they are often installed and maintained by large firms dedicated to this task, but such firms then must monitor a huge number of such systems. Handling this Big Data problem requires facing two challenges: dealing with potentially noisy and erroneous data in a fashion which preserves the important information and may even help point towards repairable failures in the monitoring systems, and using the cleansed data to build models of the battery systems which will allow for prediction of their state. In this work, we survey the scope of progress in these two areas, presenting papers which have looked at the data cleansing and battery monitoring problems in the context of battery backup installations. We also consider the work which has yet been performed, areas which retain the potential for future research.
Liz Aranguren Pachano, Taghi M. Khoshgoftaar, Randall Wald
ICMLA (2)2
2013 Random Forest with 200 Selected Features: An Optimal Model for Bioinformatics Research
abstract
Many problems in bioinformatics involve high-dimensional, difficult-to-process collections of data. For example, gene micro arrays can record the expression levels of thousands of genes, many of which have no relevance to the underlying medical or biological question. Building classification model son such datasets can thus take excessive computational time and still give poor results. Many strategies exist to combat these problems, including feature selection (which chooses only the most relevant genes for building models) and ensemble learners (which combine multiple weak classification learners into one collection which should give a broader view of the data). However, these techniques present a new challenge: choosing which combination of strategies is most appropriate for a given collection of data. This is especially difficult for health informatics and bioinformatics practitioners who do not have an extensive machine learning background. An ideal model should be easy to use and apply, helping the practitioner by either making these choices in advance or by being insensitive to these choices. In this work we demonstrate that the Random Forest learner, when using 100 trees and 200 features (selected by any reasonable feature ranking technique, as the specific choice does not matter), is such a model. To show this, we use 25 bioinformatics datasets from a number of different cancer diagnosis and identification problems, and we compare Random Forest with 5 other learners. We also tested 25 feature ranking techniques and 12 feature subset sizes, to optimize the feature selection step. Our results show that Random Forest with 100 trees and 200 selected features is statistically significantly better than any of the alternatives (orin the case of using 200 features, is statistically equivalent with the top choices), and that the specific choice of ranking technique is statistically insignificant.
Randall Wald, Taghi M. Khoshgoftaar, David J. Dittman, Amri Napolitano
ICMLA (1)2
2013 Comparison of Stability for Different Families of Filter-Based and Wrapper-Based Feature Selection
abstract
Due to the prevalence of high dimensionality(having a large number of independent attributes), feature selection techniques (which reduce the feature subset to amore manageable size) have become quite popular. These reduced feature subsets can help improve the performance of classification models and can also inform researchers about which features are most relevant for the problem at hand. For this latter problem, it is often most important that the features chosen are consistent even in the face of changes(perturbations) to the dataset. While previous studies have considered the problem of finding so-called "stable" feature selection techniques, none has examined stability across all three major categories of feature selection technique: filter-based feature rankers (which use statistical measures to assign scores to each feature), filter-based subset evaluators (which also employ statistical approaches, but consider whole feature subsets at a time), and wrapper-based subset evaluation (which also considers whole subsets, but which builds classification models to evaluate these subsets). In the present study, we use two datasets from the domain of Twitter profile mining to compare the stability of five filter-based rankers, two filter-based subset evaluators, and five wrapper-based subset evaluators. We find that the rankers are most stable, followed by the filter-based subset evaluators, with the wrappers being the least stable. We also show that the relative performance among the techniques within each group is consistent across dataset and perturbation level. However, the relative stability of the two datasets does vary between the groups, showing that the effects are more complex than simply "one group is always more stable than another group".
Randall Wald, Taghi M. Khoshgoftaar, Amri Napolitano
ICMLA (2)2
2013 Comparative Analysis on the Stability of Feature Selection Techniques Using Three Frameworks on Biological Datasets
abstract
Feature (gene) selection is a common preprocessing technique used to counter the problem of high dimensionality(too many independent features) found in many bioinformaticsdatasets, addressing this problem by creating a smaller feature subset including only the most important features. Although feature selection techniques are often evaluated based on how they can help improve classification performance, it is also important to find stable feature selection techniques which will give consistent results even in the face of dataset perturbations(such as class noise or sampling used to alleviate the problem of imbalanced data). This is especially important in bioinformatics, where the prime concern may be gene discovery rather than classification. In this study we use three frameworks to evaluate the stability of gene selection techniques: "sampledcleanvs. sampled-clean, " "sampled-noisy vs. sampled-noisy, " and" sampled-clean vs. sampled-noisy." All frameworks involve pairwisecomparisons among the results from the perturbed datasets(due to sampling or class noise injection followed by sampling). They differ in terms of whether they observe how sampling can create variation within the feature subsets (sampled-clean vs. sampled-clean), how noisy datasets (which were then sampled)can create a wide spread of selected features (sampled-noisyvs. sampled-noisy), or how features selected on clean and noisy datasets differ, after both datasets have been sampled (sampledcleanvs. sampled-noisy). Along with these three frameworks, our comparison of seven feature ranking techniques uses four cancer gene datasets, applies three sampling techniques, and generates artificial class noise to better simulate real-world datasets. The results from the frameworks are generally similar, with Signal-To-Noise and ReliefF showing the best stability and Gain Ratio showing the worst across all three frameworks, although Relief-W is notable for showing moderate to above-average stability when the clean datasets are used, but giving the second worst performance when noise was present.
Randall Wald, Taghi M. Khoshgoftaar, Ahmad Abu Shanab, Amri Napolitano
ICMLA (1)2
2013 An Empirical Study on Wrapper-Based Feature Selection for Software Engineering Data
abstract
Software metrics give valuable information for understanding and predicting the quality of software modules, and thus it is important to select the right software metrics for building software quality classification models. In this paper we focus on wrapper-based feature (metric) selection techniques, which evaluate the merit of feature subsets based on the performance of classification models. We seek to understand the relationship between the internal learner used inside wrappers and the external learner for building the final classification model. We perform experiments using four consecutive releases of a very large telecommunications system, which include 42 software metrics (and with defect data collected for every program module). Our results demonstrate that (1) the best performance is never found when the internal and external learner match, (2)the best performance is usually found by using NB (Naïve Bayes) inside the wrapper unless SVM (Support Vector Machine) is external learner, (3) LR (Logistic Regression) is often the best learner to use for building classification models regardless of which learner was used inside the wrapper.
Huanjing Wang, Taghi M. Khoshgoftaar, Amri Napolitano
ICMLA (2)2
2013 Maximizing Classification Performance for Patient Response Datasets
abstract
The ability to predict a patient's response to a treatment has long been a goal in the fields of medicine andpharmacology. This is especially true for cancer treatments, as many of these incur extreme side effects as a consequenceof destroying healthy cells along with cancerous ones. Geneprofiles such as DNA microarrays could potentially containinformation on which treatments are most likely to work withminimal side effects. However, DNA microarray datasets canbe challenging due to the large number of features (genes) per sample, many of which are irrelevant or redundant. Techniques from the domain of data mining may help both identifythe most important features and build classification modelsusing those features. This paper is a comprehensive study onthe relative performance of many different feature selectionapproaches and classification models when applied to fifteenpatient response datasets. We use six classifiers along withtwelve feature subset sizes and twenty-five feature selection techniques. Our results show that the Random Forest classifieris the top performing classifier in terms of both average resultsacross all feature selection techniques and when using thebest-performing feature selection technique, and also had thesmallest range between the best and worst performing featureselection techniques. Additionally, we found that for the averageand best feature selection technique performance, as the featuresubset size increases, the classification performance increases. Finally, we found that different feature selection techniquesdominated performance for different feature subset sizes, andlikewise the worst performers also depended on the chosenfeature subset size. Statistical analysis was conducted to furthervalidate our results. Overall, based on our results we wouldrecommend the use of Random Forest along with a featureselection technique (the choice not being statistically significant)that reduces the feature set to around 1000 features, in orderto both maximize classification performance and remove onestep (choosing an appropriate feature ranking technique) from the process.
David J. Dittman, Taghi M. Khoshgoftaar, Randall Wald, Amri Napolitano
ICTAI2
2013 A Review of Ensemble Classification for DNA Microarrays Data
abstract
Ensemble classification has been a frequent topic of research in recent years, especially in bioinformatics. The benefits of ensemble classification (less prone to overfitting, increased classification performance, and reduced bias) are a perfect match for a number of issues that plague bioinformatics experiments. This is especially true for DNA microarray data experiments, due to the large amount of data (results from potentially tens of thousands of gene probes per sample) and large levels of noise inherent in the data. This work is a review of the current state of research regarding the applications of ensemble classification for DNA microarrays. We discuss what research thus far has demonstrated, as well as identify the areas where more research is required.
Taghi M. Khoshgoftaar, David J. Dittman, Randall Wald, Wael Awada
ICTAI1
2013 Stability of Filter- and Wrapper-Based Feature Subset Selection
abstract
High dimensionality (too many features) is found across many data science domains. Feature selection techniques address this problem by choosing a subset of features whichare more relevant to the problem at hand. These technique scan simply rank the features, but this risks including multiple features which are individually useful but which contain redundant information, subset evaluation techniques, on the other hand, consider the usefulness of whole subsets, and therefore avoid selecting redundant features. Subset-based techniques can either be filters, which apply some statistical test to thesubsets to measure their worth, or wrappers, which judgefeatures based on how effective they are when building a model. One known problem with subset-based techniques is stability: because redundant features are not included, slight changes to the input data can have a significant effect on which features are chosen. In this study, we explore the stability of feature subset selection, including two filter-based techniques and five choices for both the wrapper learner and the wrapper performance metric. We also introduce a new stability metric, the modified Kuncheva's consistency index, which is able tocompare two feature subsets of different size. We also considerboth the stability of the feature selection technique and the average/standard deviation of feature subset size. Our results show that the Consistency feature subset evaluator has thegreatest stability overall, but CFS (Correlation-Based Feature Selection) shows moderate stability with a much smaller standard deviation of feature subset size. All of the wrapper-basedtechniques are less stable than the filter-based techniques, although the Naïve Bayes learner using the AUC performancemetric is the most stable wrapper-based approach.
Randall Wald, Taghi M. Khoshgoftaar, Amri Napolitano
ICTAI2
2013 How the Choice of Wrapper Learner and Performance Metric Affects Subset Evaluation
abstract
Due to the widespread problem of high dimensionality(datasets with many features/independent attributes), feature selection has become an important research topic in many areas of machine learning. One form of feature selection, wrapper-based subset evaluation, has been the focus of a moderate amount of research, because its use of classification learners to find optimal feature subsets has the potential to remove redundant features and find feature subsets which directly achieve the goal of improving classification performance. However, while the choice of learner to use within the wrapper framework has previously been studied, no paper has thoroughly investigated the role of the performance metric used within the wrapper process. Especially with imbalanced data (data where one class predominates over other classes), traditional metrics such as accuracy can give a misleading view of how many instances from each class are mislabeled. While it seems intuitive that metrics which take balance into account will affect the chosen features, no previous study has investigated this effect directly. In the present work, we test five different learners and five different performance metrics within the wrapper framework and use a newly-proposed variant of the Tanimoto Index to evaluate the similarity among the different choices of learner and metric as all other factors are held constant, using two datasets from the domain of social network profile mining. We find that while the Best Arithmetic Mean and Best Geometric Mean metrics (both of which find the stated means of True Positive Rate and True Negative Rate) are somewhat similar, they still are quite distinct, and no other metrics are particularly similar to one another. The five learners were also found to produce extremely dissimilar feature subsets. Thus, we show that the choice of both learner and metric has a major effect on which features are selected by through wrapper-based feature selection.
Randall Wald, Taghi M. Khoshgoftaar, Amri Napolitano
ICTAI2
2013 Should the Same Learners Be Used Both within Wrapper Feature Selection and for Building Classification Models?
abstract
Due to the problem of high-dimensionality(datasets which contain many independent attributes or features), feature selection has become an important part of data mining research. One popular form of feature selection, wrapper selection, chooses the best features by directly addressing the question of which features build the best models. Various feature subsets are used to build classification models, and the performance of these models is the score of each feature subset. The feature subset with the best score is then used to build the final classification model. As wrappers use a classification algorithm (learner) both to select the features and to build a predictive model, it has been traditional to use the same learner for both, such that the features chosen will be those which optimize that model's performance. However, no research has considered whether having different learners operate inside and outside the wrapper (that is, for selectingthe features and for building the final model) might actually result in improved classification performance. In this work, weconsider five learners both inside the wrapper and for building the classification model, along with two datasets drawn fromthe domain of Twitter profile mining. By considering both the raw performance values and a statistical analysis, we find that contrary to intuition, usually the best performance for a given choice of external learner is not found by using the same learner within the wrapper. Instead, the Naïve Bayes learner is usually the best choice for selecting features, regardless of which learner is used for the external model. We also find that Multi-Layer Perceptron is able to build consistent classification models for many different choices of internal learner. Finally, the 5-Nearest Neighbor learner gave poor results both inside and outside the wrapper.
Randall Wald, Taghi M. Khoshgoftaar, Amri Napolitano
ICTAI2
2013 Which Users Reply to and Interact with Twitter Social Bots?
abstract
Bots, autonomous programs which attempt to imitate human behavior, are becoming more of an issue on Twitter as the popular social network becomes an important target for spammers and others who wish to guide publicsentiment. While much research has attempted to discover andfilter out bots, less work has focused on the properties of those most susceptible to bots. In this study, we examine data from 610 users who were contacted by a bot and who could choose to either reply to or follow the bot, in order to discover which traits lead to replies alone or to interaction (replies or follows). In particular, we use feature ranking algorithms to order thedifferent traits (features) in terms of their relevance to thetwo classes ("interacted with bot" and "replied to bot"), andthen both consider those features which are frequently selectedby many different algorithms as well as use these features tobuild classification models. This is the first study to considerthe features which can influence both of these classes, and inparticular note how the chosen features differ between the twoclasses. We found that while these two classes have much incommon (for example, users with high Klout.com scores aremore likely to both reply to and interact with bots), therewere notable differences between them both in terms of whichfeatures are most important and which algorithms can be usedto build the best models. We also found that overall, the bestclassification models built with feature selection outperformedthe best models without feature selection, but the optimalchoices of feature ranker, learner, and feature subset size werenecessary to achieve this performance.
Randall Wald, Taghi M. Khoshgoftaar, Amri Napolitano, Chris Sumner
ICTAI2
2013 Comparison of Two Frameworks for Measuring the Stability of Gene-Selection Techniques on Noisy Class-Imbalanced Data
abstract
A common challenge encountered with feature (gene) selection is the instability of selected genes, which is defined as the degree of agreement between its outputs to differently-perturbed versions of the same input data. Very little work considers the impact of noise and sampling (e.g., preprocessing techniques used to cope with class imbalance, such as undersampling and oversampling) on the stability of gene selection techniques. In this study we compare two frameworks for evaluating this stability: "sampled-noisy vs. clean" and "sampled-noisy vs. sampled-noisy." Both frameworks involve noise injection followed by sampling, they differ in that the first compares the features selected from the perturbed (due to noise injection followed by sampling) datasets with the features selected from the original (clean) dataset, while the second performs a pairwise comparisons among the results from the perturbed datasets. Intuitively, the first framework should have more consistent results, since it is only randomizing one half of the comparison rather than both halves. The primary goal of this paper is to discover whether despite this, these two frameworks show similar patterns and conclusions. This is tested using four groups of cancer gene datasets. We employ ten feature rankers from three different families, apply three sampling techniques, and generate artificial class noise to better simulate real-world datasets. The results show that Mutual Information, Signal-To-Noise, and Deviance show the best stability across the two frameworks, while Gain Ratio shows the worst stability on average. The results also show that two frameworks have the same stability pattern, i.e., the rankers that perform well (or poorly) in the first framework perform as well (or as poorly) in the second. This means that the second framework, which is less computationally intensive (due to not performing feature selection on the clean data), can be used without requiring the first framework.
Randall Wald, Taghi M. Khoshgoftaar, Ahmad Abu Shanab
ICTAI2
2013 Exploring Ensemble-Based Data Preprocessing Techniques for Software Quality Estimation
Kehan Gao, Taghi M. Khoshgoftaar, Amri Napolitano
SEKE2
2013 Overcoming Big Data Challenges
Taghi M. Khoshgoftaar
SEKE1
2013 A Study on First Order Statistics-Based Feature Selection Techniques on Software Metric Data
Huanjing Wang, Taghi M. Khoshgoftaar, Randall Wald, Amri Napolitano
SEKE2
2012 Similarity analysis of feature ranking techniques on imbalanced DNA microarray datasets
abstract
DNA microarrays are a modern advancement in the analysis of genetic data. This technology allows a researcher to test samples for thousands of genes simultaneously. However, once the samples in the DNA microarrays have been tested, the researcher must then search through the data collected and identify genes important to their problem. A possible solution to this issue is the data mining pre-processing technique called feature selection. Feature (gene) selection takes the original set of features (in the case of DNA microarrays, gene probes) and chooses an optimal subset to perform analysis from. Ideally, the reduced subset only contains the most important features as determined by the feature selection technique (or set of feature selection techniques), which allows for further research in the discovered genes. However in the case of using multiple feature selection techniques, the set of techniques must be diverse in order to reduce redundancy among the chosen features. Another benefit of increasing diversity is that any features chosen across a diverse set of feature selection techniques will have more importance than those chosen by a single technique or a set of related ones. Therefore, it would be useful to know how similar the feature selection techniques are to each other. In this study we perform an analysis of eighteen feature selection techniques across nine imbalanced DNA microarray datasets and using four feature subset sizes. Our results found that one should not use Gini Index and Probability Ratio together or the Kolmogorov-Smirnov statistic and Geometric Mean together at any feature subset size in order to minimize redundancy, and that the members of the first of these pairs (along with the pair of ReliefF and ReliefF-W) are very dissimilar to all rankers outside their own cluster. We also found that Chi-Squared, Information Gain, and Symmetric Uncertainty form a cluster of similarity, as do Chi-Squared, Deviance, F-Measure, and Mutual Information.
David J. Dittman, Taghi M. Khoshgoftaar, Randall Wald, Amri Napolitano
BIBM2
2012 The effect of measurement approach and noise level on gene selection stability
abstract
Many biological datasets exhibit high dimensionality, a large abundance of attributes (genes) per instance (sample). This problem is often solved using feature selection, which works by selecting the most relevant attributes and removing irrelevant and redundant attributes. Although feature selection techniques are often evaluated based on the performance of classification models (e.g., algorithms designed to distinguish between multiple classes of instances, such as cancerous vs. noncancerous) built using the selected features, another important criterion which is often neglected is stability, the degree of agreement among a feature selection technique's outputs when there are changes to the dataset. More stable feature selection techniques will give the same features even if aspects of the data change. In this study we consider two different approaches for evaluating the stability of feature selection techniques, with each approach consisting of noise injection followed by feature ranking. The two approaches differ in that the first approach compares the features selected from the noisy datasets with the features selected from the original (clean) dataset, while the second approach performs pairwise comparisons among the results from the noisy datasets. To evaluate these two approaches, we use four biological datasets and employ six commonly-used feature rankers. We draw two primary conclusions from our experiments: First, the rankers show different levels of stability in the face of noise. In particular, the ReliefF ranker has significantly greater stability than the other rankers. Also, we found that both approaches gave the same results in terms of stability patterns, although the first approach had greater stability overall. Additionally, because the first approach is significantly less computationally expensive, future studies may employ a faster technique to gain the same results.
Randall Wald, Taghi M. Khoshgoftaar, Ahmad Abu Shanab
BIBM2
2012 The Effect of Number of Iterations on Ensemble Gene Selection
abstract
Dimensionality-reducing techniques such as gene selection have become commonplace in order to reduce the high dimensionality found within bioinformatics datasets such as DNA microarray datasets. The degree of dimensionality is reduced by identifying and removing redundant and irrelevant features or genes and leaving only an optimum subset of features for subsequent analysis. However, a number of feature selection techniques show poor stability (resistance to change in the underlying data). One approach for increasing the stability of feature subsets is ensemble feature selection. This is performed first by generating multiple ranked gene lists and then aggregating the results using an aggregation function. While research has been performed on ensemble feature selection and its effect on gene list stability, there has been little research on an important choice made in the process of ensemble feature selection: the number of iterations (or repetitions) of feature selection. The computation time of ensemble feature selection is greatly affected by the number of ranked lists generated: the higher the number of iterations, the more computation time is required. To study this, we evaluate the similarity among feature subsets generated from two different approaches to ensemble feature selection (data diversity and hybrid approach). We calculate the similarity between the final ranked lists generated using 10, 20 and 50 iterations, using the mean aggregation function. Our results show that the similarity between 20 and 50 iterations is high enough for us to recommend using 20 iterations instead of 50 and thus saving the large amount of computation time required for 50 iterations.
Wael Awada, Taghi M. Khoshgoftaar, David J. Dittman, Randall Wald
ICMLA (2)2
2012 Determining the Number of Iterations Appropriate for Ensemble Gene Selection on Microarray Data
abstract
Ensemble gene (feature) selection is a promising new strategy with many benefits including more stable gene lists and improved classification results. The ensemble portion is achieved through multiple runs of feature selection which are then aggregated into a single result. The critical question is how many iterations of feature selection are appropriate. Too few iterations can make classification performance suffer. However, too many iterations will cause issues regarding computational costs. The goal is to choose the correct number of iterations to maximize classification performance without expending too much computational power. Our paper is an in-depth study on the effect of the number of iterations of feature selection on classification performance. Our work employs eleven DNA microarray datasets, on which we apply various ensemble methods, feature selection techniques, classifiers, and feature subset sizes. The results show that using 10 iterations of feature ranking during ensemble feature selection is not sufficient to optimize classification results and that a larger number of iterations is required (20 or 50). However, there is very little distinction between 20 iterations and 50 iterations, as both produce very similar classification results. Our recommendation is to use 20 iterations because while 20 iterations and 50 iterations perform similarly, 20 iterations has a much smaller computation time. To our knowledge there has not been a previous study as expansive as this one on the effects of the number of iterations of feature selection on ensemble feature selection.
David J. Dittman, Taghi M. Khoshgoftaar, Randall Wald, Amri Napolitano
ICMLA (1)2
2012 Comparing Two New Gene Selection Ensemble Approaches with the Commonly-Used Approach
abstract
Ensemble feature selection has recently become a topic of interest for researchers, especially in the area of bioinformatics. The benefits of ensemble feature selection include increased feature (gene) subset stability and usefulness as well as comparable (or better) classification performance compared to using a single feature selection method. However, existing work on ensemble feature selection has concentrated on data diversity (using a single feature selection method on multiple datasets or sampled data from a single dataset), neglecting two other potential sources of diversity. We present these two new approaches for gene selection, functional diversity (using multiple feature selection technique on a single dataset) and hybrid (a combination of data and functional diversity). To demonstrate the value of these new approaches, we measure the similarity between the feature subsets created by each of the three approaches across twenty-six datasets and ten feature selection techniques (or an ensemble of these techniques as appropriate). We also compare the classification performance of models built using each of the three ensembles. Our results show that the similarity between the functional diversity and hybrid approaches is much higher than the similarity between either of those and data diversity, with the distinction between data diversity and our new approaches being particularly strong for hard-to-learn datasets. In addition to having the highest similarity, functional and hybrid diversity generally show greater classification performance than data diversity, especially when selecting small feature subsets. These results demonstrate that these new approaches can both provide a different feature subset than the existing approach and that the resulting novel feature subset is potentially of interest to researchers. To our knowledge there has been no study which explores these new approaches to ensemble feature selection within the domain of bioinformatics.
David J. Dittman, Taghi M. Khoshgoftaar, Randall Wald, Amri Napolitano
ICMLA (2)2
2012 Decision Level Fusion of Wavelet Features for Ocean Turbine State Detection
abstract
Data fusion is the process of combining data from multiple sources, allowing for a more complete and accurate assessment of a system or an environment than could have been otherwise provided by a single source. This paper considers and empirically evaluates a decision-level data fusion method for enabling reliable ocean turbine state detection based on data from multiple sensors. This method involves first generating a classification model from the data from individual sensor channels. For each new incoming instance, the probability that this new instance belongs to each of the possible system states is then computed individually based on observations made by each source. These probabilities are averaged and the system state with the highest probability is selected as the fused output. In a case study presented in this paper, six accelerometers mounted at different positions along a dynamometer test bed for an ocean turbine measure the vibration of various components while the machine is in operation. Each sensor gives unique information about the dynamometer thus ignoring data from one or more sources (or sensors) means possibly discarding useful information. We apply the decision-level fusion method to combine the decisions made from the individual channels and allow for more informed state detection which considers all sources together. Five popular machine learning algorithms are used to generate the classification models from the individual sensor data to see how the performances of each algorithm is affected by fusion and to determine whether this decision level fusion approach can lead to optimal behavior for an ocean turbine state detection module. All five learners are found to benefit from such an approach. Of all the learners, the k-Nearest Neighbors algorithm produces the best results after fusion.
Janell Duhaney, Taghi M. Khoshgoftaar
ICMLA (2)2
2012 Studying the Effect of Class Imbalance in Ocean Turbine Fault Data on Reliable State Detection
abstract
Class imbalance is prevalent in many real world datasets. It occurs when there are significantly fewer examples in one or more classes in a dataset compared to the number of instances in the remaining classes. When trained on highly imbalanced datasets, traditional machine learning techniques can often simply ignore the minority class(es) and label all instances as being of the majority class to maximize accuracy. This problem has been studied in many domains but there is little or no research related to the effect of class imbalance in fault data for condition monitoring of an ocean turbine. This study makes the first efforts in bridging that gap by providing insight into how class imbalance in vibration data can impact a learner's ability to reliably identify changes in the ocean turbine's operational state. To do so, we empirically evaluate the performances of three popular, but very different, machine learning algorithms when trained on four datasets with varying class distributions (one balanced and three imbalanced) to distinguish between a normal and an abnormal state. All data used in this study were collected from the testbed for an ocean turbine and were under sampled to simulate the different levels of imbalance. We find here, as in other domains, that the three learners seemed to suffer overall when trained on data with a highly skewed class distribution (with 0.1% examples in a faulty/abnormal state while the remaining 99.9% were captured in a normal operational state). It was noted, however, that the Logistic Regression and Decision Tree classifiers performed better when only 5% of the total number of examples were representative of an abnormal state (the remaining 95% therefore indicating normal operation) than they did when there was no imbalance present.
Janell Duhaney, Taghi M. Khoshgoftaar, Amri Napolitano
ICMLA (1)2
2012 Applying Feature Selection to Short Time Wavelet Transformed Vibration Data for Reliability Analysis of an Ocean Turbine
abstract
This paper considers the use of feature selection within the state detection module for an ocean turbine condition monitoring system. The goal is to reduce the quantity of data to be processed while maintaining or improving state detection capabilities. Five feature selection techniques (Chi-squared, Information Gain, Signal-To-Noise, AUC and PRC) are evaluated based on their effects on four widely used machine learning algorithms, namely Naive Bayes, k-Nearest Neighbors, Decision Tree and Logistic Regression, when each machine learner is trained on the top n features selected by each feature selection technique. Six values of n (2, 4, 6, 8, 10 and 15) were considered. Features were extracted from the raw vibration signals using a Short Time Wavelet Transform with Baselining (STWTB) technique designed to allow for reliable state detection regardless of the turbine's operating conditions, which are often reflected within its vibration readings. The condition-independent features extracted by the STWTB are then fused to combine all the data observed by all sensor sources. Models were built on data gathered at one operating condition and tested against data from a different operating condition to simulate the problem of building models which work regardless of operating condition. Results show that k-Nearest Neighbors, Naive Bayes and Logistic Regression have improved classification performance when using less than 11% of the 78 available features, with Logistic Regression needing just 2 features selected by the Signal-To-Noise technique to generate a perfect classification model. The Decision Tree performed best without feature selection.
Janell Duhaney, Taghi M. Khoshgoftaar, Randall Wald
ICMLA (1)2
2012 A Hybrid Approach to Coping with High Dimensionality and Class Imbalance for Software Defect Prediction
abstract
High dimensionality and class imbalance are the two main problems affecting many software defect prediction. In this paper, we propose a new technique, named SelectRUSBoost, which is a form of ensemble learning that in-corporates data sampling to alleviate class imbalance and feature selection to resolve high dimensionality. To evaluate the effectiveness of the new technique, we apply it to a group of datasets in the context of software defect prediction. We employ two classification learners and six feature selection techniques. We compare the technique to the approach where feature selection and data sampling are used together, as well as the case where feature selection is used alone (no sampling used at all). The experimental results demonstrate that the SelectRUSBoost technique is more effective in improving classification performance compared to the other approaches.
Kehan Gao, Taghi M. Khoshgoftaar, Amri Napolitano
ICMLA (2)2
2012 A Novel Noise-Resistant Boosting Algorithm for Class-Skewed Data
abstract
Boosting methods have been successfully applied in a wide variety of machine learning applications. In the context of data quality issues, a number of variants of the standard boosting method have been proposed and evaluated. To address the problem of mislabeled examples, ORBoost was developed to prevent over fitting to noisy examples. Our research group has recently proposed RUSBoost as an enhancement to the AdaBoost algorithm for dealing with skewed class distributions. This work proposes a modification to the RUSBoost algorithm, incorporating the noise-handling ability of ORBoost, to improve its handling of noisy data. The new method is compared with both ORBoost and RUSBoost in an extensive set of experiments using five real-world datasets with various levels of simulated noise.
Jason Van Hulse, Taghi M. Khoshgoftaar, Amri Napolitano
ICMLA (2)2
2012 First Order Statistics Based Feature Selection: A Diverse and Powerful Family of Feature Seleciton Techniques
abstract
Dimensionality reduction techniques have become a required step when working with bioinformatics datasets. Techniques such as feature selection have been known to not only improve computation time, but to improve the results of experiments by removing the redundant and irrelevant features or genes from consideration in subsequent analysis. Univariate feature selection techniques in particular are well suited for the large levels of high dimensionality that are inherent in bioinformatics datasets (for example: DNA microarray datasets) due to their intuitive output (a ranked lists of features or genes) and their relatively small computational time compared to other techniques. This paper presents seven univariate feature selection techniques and collects them into a single family entitled First Order Statistics (FOS) based feature selection. These seven all share the trait of using first order statistical measures such as mean and standard deviation, although this is the first work to relate them to one another and consider their performance compared with one another. In order to examine the properties of these seven techniques we performed a series of similarity and classification experiments on eleven DNA microarray datasets. Our results show that in general, each feature selection technique will create diverse feature subsets when compared to the other members of the family. However when we look at classification we find that, with one exception, the techniques will produce good classification results and that the techniques will have similar performances to each other. Our recommendation, is to use the rankers Signal-to-Noise and SAM for the best classification results and to avoid Fold Change Ratio as it is consistently the worst performer of the seven rankers.
Taghi M. Khoshgoftaar, David J. Dittman, Randall Wald, Alireza Fazelpour
ICMLA (2)1
2012 Mean Aggregation versus Robust Rank Aggregation for Ensemble Gene Selection
abstract
Feature (gene) selection is an important preprocessing step for performing data mining on large-scale bioinformatics datasets. However, one known concern is that feature selection can sometimes give very different results when applied to very similar data sets. Ensemble gene selection is a promising new approach which may help resolve this concern, producing more stable gene lists and better classification results. Ensemble selection consists of multiple runs of feature ranking which are then combined into a single ranking for each feature. However, one of the most critical decisions when performing ensemble gene selection is deciding on which aggregation technique to use for combining the resulting ranked feature lists from the multiple runs of feature ranking into a single decision for each gene. This paper is an in-depth comparison between two aggregation techniques: Mean Aggregation (a simple and commonly-used technique) and Robust Rank Aggregation (a recently proposed aggregation technique designed specifically for bioinformatics). Our results show that in general Mean Aggregation will outperform (or at least match) Robust Rank Aggregation in terms of classification performance, while being significantly simpler to implement and perform. These results allows us to recommend with reasonable confidence the use of Mean Aggregation over Robust Rank Aggregation.
Randall Wald, Taghi M. Khoshgoftaar, David J. Dittman
ICMLA (1)2
2012 A New Fixed-Overlap Partitioning Algorithm for Determining Stability of Bioinformatics Gene Rankers
abstract
Feature (gene) selection has become an important and necessary step for combating high dimensionality, a problem found in bioinformatics datasets. Many studies have focused on gene selection, examining both the design of these techniques and the classification performance of prediction models built using these techniques. However, it is only recently that any work has focused on the robustness or stability of these gene selection techniques. Robustness is important because techniques which do not give reliable gene lists cannot be trusted to give useful genes. Previous papers studying stability typically generate multiple random sub samples of the original dataset and compare the genes chosen from these with one another, or compare the genes from the sub samples directly with the genes from the original data. These methods both have known problems, either with comparing two randomly-generated datasets with an unknown level of overlap or with comparing two datasets of different sizes. This paper introduces a new algorithm for generating sub sample datasets called fixed-overlap partitions. This will generate sub samples which have exactly the desired level of overlap and number of instances. Using this method we evaluate nineteen feature selection techniques using twenty-six real world DNA microarray datasets. Our results show that there are three rankers (Deviance, Receiver Operating Characteristic curve, and Precision-Recall Curve) which are consistently the most stable. However, the level of overlap, the quality of the data, and the number of genes selected have an effect on which ranker will be the most stable in a given situation. The fixed-overlap partitions algorithm in particular is able to find how varying levels of overlap can cause different levels of difficulty to sometimes resemble one another (for example, moderate-difficulty datasets behave like easy-difficulty datasets at low levels of overlap, but diverge as the overlap increases).
Randall Wald, Taghi M. Khoshgoftaar, David J. Dittman
ICMLA (2)2
2012 Using Twitter Content to Predict Psychopathy
abstract
An ever-growing number of users share their thoughts and experiences using the Twitter micro logging service. Although sometimes dismissed as containing too little content to convey significant information, these messages can be combined to build a larger picture of the user posting them. One particularly notable personality trait which can be discovered this way is psychopathy: the tendency for disregarding others and the rule of society. In this paper, we explore techniques to apply data mining towards the goal of identifying those who score in the top 1.4% of a well-known psychopathy metric using information available from their Twitter accounts. We apply a newly-proposed form of ensemble learning, Select RUSBoost (which adds feature selection to our earlier imbalance-aware ensemble in order to resolve high dimensionality), employ four classification learners, and use four feature selection techniques. The results show that when using the optimal choices of techniques, we are able to achieve an AUC value of 0.736. Furthermore, these results were only achieved when using the Select RUSBoost technique, demonstrating the importance of feature selection, data sampling, and ensemble learning. Overall, we show that data mining can be a valuable tool for law enforcement and others interested in identifying abnormal psychiatric states from Twitter data.
Randall Wald, Taghi M. Khoshgoftaar, Amri Napolitano, Chris Sumner
ICMLA (2)2
2012 An Empirical Study on the Stability of Feature Selection for Imbalanced Software Engineering Data
abstract
In software quality modeling, software metrics are collected during the software development cycle. However, not all metrics are relevant to the class attribute (software quality). Metric (feature) selection has become the cornerstone of many software quality classification problems. Selecting software metrics that are important for software quality classification is a necessary and critical step before the model training process. Recently, the robustness (e.g., stability) of feature selection techniques has been studied, to examine the sensitivity of these techniques to changes (adding/removing program modules to/from their dataset). This work provides an empirical study regarding the stability of feature selection techniques across six software metrics datasets with varying levels of class balance. In this work eighteen feature selection techniques are evaluated. Moreover, three factors, feature subset size, degree of perturbation, and class balance of datasets, are considered in this study to evaluate stability of feature selection techniques. Experimental results show that these factors affect the stability of feature selection techniques as one might expect. We found that with few exceptions, feature ranking based on highly imbalanced datasets are less stable than based on slightly imbalanced data. Results also show that making smaller changes to the datasets has less impact on the stability of feature ranking techniques. Overall, we conclude that a careful understanding of one's dataset (and certain choices of metric selection technique) can help practitioners build more reliable software quality models.
Huanjing Wang, Taghi M. Khoshgoftaar, Amri Napolitano
ICMLA (1)2
2012 A Comparative Study on the Stability of Software Metric Selection Techniques
abstract
In large software projects, software quality prediction is an important aspect of the development cycle to help focus quality assurance efforts on the modules most likely to contain faults. To perform software quality prediction, various software metrics are collected during the software development cycle, and models are built using these metrics. However, not all features (metrics) make the same contribution to the class attribute (e.g., faulty/not faulty). Thus, selecting a subset of metrics that are relevant to the class attribute is a critical step. As many feature selection algorithms exist, it is important to find ones which will produce consistent results even as the underlying data is changed, this quality of producing consistent results is referred to as "stability." In this paper, we investigate the stability of seven feature selection techniques in the context of software quality classification. We compare four approaches for varying the underlying data to evaluate stability: the traditional approach of generating many sub samples of the original data and comparing the features selected from each, an earlier approach developed by our research group which compares the features selected from sub samples of the data with those selected from the original, and two newly-proposed approaches based on comparing two sub samples which are specifically designed to have same number of instances and a specified level of overlap, with one of these new approaches comparing within each pair while the other compares the generated sub samples with the original dataset. The empirical validation is carried out on sixteen software metrics datasets. Our results show that ReliefF is the most stable feature selection technique. Results also show that the level of overlap, degree of perturbation, and feature subset size do affect the stability of feature selection methods. Finally, we find that all four approaches of evaluating stability produce similar results in terms of which feature selection techniques are best under different circumstances.
Huanjing Wang, Taghi M. Khoshgoftaar, Randall Wald, Amri Napolitano
ICMLA (2)2
2012 Stability of Filter-Based Feature Selection Methods for Imbalanced Software Measurement Data
Kehan Gao, Taghi M. Khoshgoftaar, Amri Napolitano
SEKE2
2012 An Empirical Study of Software Metric Selection Techniques for Defect Prediction
Huanjing Wang, Taghi M. Khoshgoftaar, Randall Wald, Amri Napolitano
SEKE2
2012 Software measurement data reduction using ensemble techniques
Huanjing Wang, Taghi M. Khoshgoftaar, Amri Napolitano
Neurocomputing2
2012 An Empirical Study of Feature Ranking Techniques for Software Quality Prediction
abstract
The primary goal of software quality engineering is to produce a high quality software product through the use of some specific techniques and processes. One strategy is applying data mining techniques to software metric and defect data collected during the software development process to identify potential low-quality program modules. In this paper, we investigate the use of feature selection in the context of software quality estimation (also referred to as software defect prediction), where a classification model is used to predict whether program modules (instances) are fault-prone or not-fault-prone. Seven filter-based feature ranking techniques are examined. Among them, six are commonly used, and the other one, named signal to noise ratio (SNR), is rarely employed. The objective of the paper is to compare these seven techniques for various software data sets and assess their effectiveness for software quality modeling. A case study is performed on 16 software data sets, and classification models are built with five different learners and evaluated with two performance metrics. Our experimental results are summarized based on statistical tests for significance. The main conclusion is that the SNR technique performs as well as the best performer of the six commonly used techniques.
Taghi M. Khoshgoftaar, Kehan Gao, Amri Napolitano
Int. J. Softw. Eng. Knowl. Eng.1
2012 Predicting high-risk program modules by selecting the right software measurements
Kehan Gao, Taghi M. Khoshgoftaar, Naeem Seliya
Softw. Qual. J.2
2011 Stability Analysis of Feature Ranking Techniques on Biological Datasets
abstract
One major problem faced when analyzing DNA microarrays is their high dimensionality (large number of features). Therefore, feature selection is a necessary step when using these datasets. However, the addition or removal of instances can alter the subsets chosen by a feature selection technique. The ideal situation is to choose a feature selection technique that is robust (stable) to changes in the number of instances, with selected features changing little even when instances are added or removed. In this study we test the stability of nineteen feature selection techniques across twenty- six datasets with varying levels of class imbalance. Our results show that the best choice of technique depends on the class balance of the datasets. The top performers are Deviance for balanced datasets, Signal to Noise for slightly unbalanced datasets, and AUC for unbalanced datasets. SVM-RFE was the least stable feature selection technique across the board, while other poor performers include Gain Ratio, Gini Index, Probability Ratio, and Power. We also found that enough changes to the dataset can make any feature selection technique unstable, and that using more features increases the stability of most feature selection techniques. Most intriguing was our finding that the more imbalanced a dataset is, the more stable the feature subsets built for that dataset will be. Overall, we conclude that stability is an important aspect of feature ranking which must be taken into account when planning a feature selection strategy or when adding or removing instances from a dataset.
David J. Dittman, Taghi M. Khoshgoftaar, Randall Wald, Huanjing Wang
BIBM2
2011 Feature Selection on Dynamometer Data for Reliability Analysis
abstract
An ocean turbine extracts the kinetic energy from ocean currents to generate electricity. Vibration signals from the turbine hold a wealth of information regarding its state, and detecting changes in these signals is crucial to the timely detection of faults. Wavelet transforms provide a means of analyzing these complex signals and extracting features which are representative of the signal. Feature selection techniques are needed once these wavelet features are extracted to eliminate redundant or useless features before the data is presented to a machine learning algorithm for pattern recognition and classification. This reduces the quantity of data to be processed and can often even increase the machine learner's ability to detect the current state of the machine. This paper empirically compares eight feature selection algorithms on wavelet transformed vibration data originating from an onshore test platform for an ocean turbine. A case study shows the classification performances of seven machine learners when trained on the datasets with varying numbers of features selected from the original set of all features. Our results highlight that by choosing an appropriate feature selection technique and applying it to selecting just the 3 most important features (3.33% of the original feature set), some classifiers such as the decision tree and random forest can correctly differentiate between faulty and nonfaulty states almost 100% of the time. These results also show the performance differences between different feature selection algorithms and classifier combinations.
Janell Duhaney, Taghi M. Khoshgoftaar, John C. Sloan
ICTAI2
2011 Impact of Data Sampling on Stability of Feature Selection for Software Measurement Data
abstract
Software defect prediction can be considered a binary classification problem. Generally, practitioners utilize historical software data, including metric and fault data collected during the software development process, to build a classification model and then employ this model to predict new program modules as either fault-prone (fp) or not-fault-prone (nfp). Limited project resources can then be allocated according to the prediction results by (for example) assigning more reviews and testing to the modules predicted to be potentially defective. Two challenges often come with the modeling process: (1) high-dimensionality of software measurement data and (2) skewed or imbalanced distributions between the two types of modules (fp and nfp) in those datasets. To overcome these problems, extensive studies have been dedicated towards improving the quality of training data. The commonly used techniques are feature selection and data sampling. Usually, researchers focus on evaluating classification performance after the training data is modified. The present study assesses a feature selection technique from a different perspective. We are more interested in studying the stability of a feature selection method, especially in understanding the impact of data sampling techniques on the stability of feature selection when using the sampled data. Some interesting findings are found based on two case studies performed on datasets from two real-world software projects.
Kehan Gao, Taghi M. Khoshgoftaar, Amri Napolitano
ICTAI2
2011 Feature Selection for Vibration Sensor Data Transformed by a Streaming Wavelet Packet Decomposition
abstract
Vibration signals play a valuable role in the remote monitoring of high-assurance machinery such as ocean turbines. Because they are waveforms, vibration data must be transformed prior to being incorporated into a machine condition monitoring/prognostic health monitoring (MCM/PHM) solution to detect which frequencies of oscillation are most prevalent. One downside of these transformations, especially the streaming version of the wavelet packet decomposition (denoted SWPD), is that they can produce a large number of features, hindering the model building and evaluation process. In this paper we demonstrate how feature selection techniques may be applied to the output of the SWPD transformation, vastly reducing the total number of features used to build models. The resulting data can be used to build more accurate models for use in MCM/PHM while minimizing computation time.
Randall Wald, Taghi M. Khoshgoftaar, John C. Sloan
ICTAI2
2011 Measuring Stability of Threshold-Based Feature Selection Techniques
abstract
Feature selection has been applied in many domains, such as text mining and software engineering. Ideally a feature selection technique should produce consistent outputs regardless of minor variations in the input data. Researchers have recently begun to examine the stability (robustness) of feature selection techniques. The stability of a feature selection method is defined as the degree of agreement between its outputs to randomly-selected subsets of the same input data. This study evaluated the stability of 11 threshold-based feature ranking techniques (rankers) when applied to 16 real-world software measurement datasets of different sizes. Experimental results demonstrate that AUC (Area Under the Receiver Operating Characteristic Curve) and PRC (Area Under the Precision-Recall Curve) performed best among the 11 rankers.
Huanjing Wang, Taghi M. Khoshgoftaar
ICTAI2
2011 Using Classifier-Based Nominal Imputation to Improve Machine Learning
Xiaoyuan Su, Russell Greiner, Taghi M. Khoshgoftaar, Amri Napolitano
PAKDD (1)3
2011 Software Defect Prediction for High-Dimensional and Class-Imbalanced Data
Kehan Gao, Taghi M. Khoshgoftaar
SEKE2
2011 A Comparative Study of Different Strategies for Predicting Software Quality
Taghi M. Khoshgoftaar, Kehan Gao, Amri Napolitano
SEKE1
2011 An Empirical Study of Software Metrics Selection Using Support Vector Machine
Huanjing Wang, Taghi M. Khoshgoftaar, Amri Napolitano
SEKE2
2011 An exploration of learning when data is noisy and imbalanced
abstract
Much of the research literature in data mining and machine learning has focused on developing classification models for various application-specific learning tasks. In contrast, the characteristics of the underlying data, and their impacts on learning, have received much less attention. While it is generally understood that imbalanced, noisy and relatively small datasets make classification tasks more difficult, there has been, to our knowledge, no comprehensive examination of the impacts of these important and commonly-encountered dataset characteristics on the learning process. In this work, we present a comprehensive empirical analysis of learning from imbalanced, limited and noisy data. We present the performance of 11 commonly used learning algorithms and the effects of dataset size, class distribution, noise level and noise distribution on each learner. In this work, for which over one million classification models were built, we identify which learners are most robust to changing each of these experimental factors using two different performance metrics. Our results show that each of these factors plays a critical role in learner performance, with some learners exhibiting much greater stability than others.
Jason Van Hulse, Taghi M. Khoshgoftaar, Amri Napolitano
Intell. Data Anal.2
2011 Metric Selection for Software Defect Prediction
abstract
Real-world software systems are becoming larger, more complex, and much more unpredictable. Software systems face many risks in their life cycles. Software practitioners strive to improve software quality by constructing defect prediction models using metric (feature) selection techniques. Finding faulty components in a software system can lead to a more reliable final system and reduce development and maintenance costs. This paper presents an empirical study of six commonly used filter-based software metric rankers and our proposed ensemble technique using rank ordering of the features (mean or median), applied to three large software projects using five commonly used learners. The classification accuracy was evaluated in terms of the AUC (Area Under the ROC (Receiver Operating Characteristic) Curve) performance metric. Results demonstrate that the ensemble technique performed better overall than any individual ranker and also possessed better robustness. The empirical study also shows that variations among rankers, learners and software projects significantly impacted the classification outcomes, and that the ensemble method can smooth out performance.
Huanjing Wang, Taghi M. Khoshgoftaar, Jason Van Hulse, Kehan Gao
Int. J. Softw. Eng. Knowl. Eng.2
2011 Choosing software metrics for defect prediction: an investigation on feature selection techniques
abstract
Abstract The selection of software metrics for building software quality prediction models is a search‐based software engineering problem. An exhaustive search for such metrics is usually not feasible due to limited project resources, especially if the number of available metrics is large. Defect prediction models are necessary in aiding project managers for better utilizing valuable project resources for software quality improvement. The efficacy and usefulness of a fault‐proneness prediction model is only as good as the quality of the software measurement data. This study focuses on the problem of attribute selection in the context of software quality estimation. A comparative investigation is presented for evaluating our proposed hybrid attribute selection approach, in which feature ranking is first used to reduce the search space, followed by a feature subset selection. A total of seven different feature ranking techniques are evaluated, while four different feature subset selection approaches are considered. The models are trained using five commonly used classification algorithms. The case study is based on software metrics and defect data collected from multiple releases of a large real‐world software system. The results demonstrate that while some feature ranking techniques performed similarly, the automatic hybrid search algorithm performed the best among the feature subset selection methods. Moreover, performances of the defect prediction models either improved or remained unchanged when over 85were eliminated. Copyright © 2011 John Wiley & Sons, Ltd.
Kehan Gao, Taghi M. Khoshgoftaar, Huanjing Wang, Naeem Seliya
Softw. Pract. Exp.2
2011 Comparing Boosting and Bagging Techniques With Noisy and Imbalanced Data
abstract
This paper compares the performance of several boosting and bagging techniques in the context of learning from imbalanced and noisy binary-class data. Noise and class imbalance are two well-established data characteristics encountered in a wide range of data mining and machine learning initiatives. The learning algorithms studied in this paper, which include SMOTEBoost, RUSBoost, Exactly Balanced Bagging, and Roughly Balanced Bagging, combine boosting or bagging with data sampling to make them more effective when data are imbalanced. These techniques are evaluated in a comprehensive suite of experiments, for which nearly four million classification models were trained. All classifiers are assessed using seven different performance metrics, providing a complete perspective on the performance of these techniques, and results are tested for statistical significance via analysis-of-variance modeling. The experiments show that the bagging techniques generally outperform boosting, and hence in noisy data environments, bagging is the preferred method for handling class imbalance.
Taghi M. Khoshgoftaar, Jason Van Hulse, Amri Napolitano
IEEE Trans. Syst. Man Cybern. Part A1
2011 Ontology-Based Business Process Customization for Composite Web Services
abstract
A key goal of the Semantic Web is to shift social interaction patterns from a producer-centric paradigm to a consumer-centric one. Treating customers as the most valuable assets and making the business models work better for them are at the core of building successful consumer-centric business models. It follows that customizing business processes constitutes a major concern in the realm of a knowledge-pull-based human semantic Web. This paper conceptualizes the customization of service-based business processes leveraging the existing knowledge of Web services and business processes. We represent this conceptualization as a new Extensible Markup Language (XML) markup language Web Ontology Language-Business Process Customization (OWL-BPC), based on the de facto semantic markup language for Web-based information [Web Ontology Language (OWL)]. Furthermore, we report a framework, built on OWL-BPC, for customizing service-based business processes, which supports customization detection and enactment. Customization detection is enabled by a business-goal analysis, and customization enactment is enabled via event-condition-action rule inference. Our solution and framework have the following capabilities in dealing with inconsistencies and misalignments in business process interactions: 1) resolve semantic mismatch of process parameters; 2) handle behavioral mismatches which may or may not be compatible; and 3) process misaligned rendezvous requirements. Such capabilities are applicable to business processes with heterogeneous domain ontology. We present an architectural description of the implementation and a walk-through of an example of solving a customization problem as a validation of the proposed approach.
Qianhui Althea Liang, Xindong Wu 0001, E. K. Park, Taghi M. Khoshgoftaar, Chihung Chi
IEEE Trans. Syst. Man Cybern. Part A4
2010 Comparative Analysis of DNA Microarray Data through the Use of Feature Selection Techniques
abstract
One of today's most important scientific research topics is discovering the genetic links between cancers. This paper contains the results of a comparison of three different cancers (breast, colon, and lung) based on the results of feature selection techniques on a data set created from DNA micro array data consisting of samples from all three cancers. The data was run through a set of eighteen feature rankers which ordered the genes by importance with respect to a targeted cancer. This process was repeated three times, each time with a different target cancer. The rankings were then compared, keeping each feature ranker static while varying the cancers being compared. The cancers were evaluated both in pairs and all together, for matching genes. The results of the comparison show a large correlation between the two known hereditary cancers, breast and colon, and little correlation between lung cancer and the other cancers. This is the first study to apply eighteen different feature rankers in a bioinformatics case study, eleven of which were recently proposed and implemented by our research team.
David J. Dittman, Taghi M. Khoshgoftaar, Randall Wald, Jason Van Hulse
ICMLA2
2010 A Novel Noise Filtering Algorithm for Imbalanced Data
abstract
Noise filtering is a commonly-used methodology to improve the performance of learners built using low-quality data. A common type of noise filtering is a data preprocessing technique called classification filtering. In classification filtering, a classifier is built and evaluated on the training dataset (typically using cross-validation) and any misclassified instances are considered noisy. The strategies employed with classification filters are not ideal, particularly when learning from class-imbalanced data. To address this deficiency, we propose an alternative method for classification filtering called the threshold-adjusted classification filter. This methodology is compared with the standard classification filter, and the results clearly demonstrate the efficacy of our technique.
Jason Van Hulse, Taghi M. Khoshgoftaar, Amri Napolitano
ICMLA2
2010 A Comparative Study of Ensemble Feature Selection Techniques for Software Defect Prediction
abstract
Feature selection has become the essential step in many data mining applications. Using a single feature subset selection method may generate local optima. Ensembles of feature selection methods attempt to combine multiple feature selection methods instead of using a single one. We present a comprehensive empirical study examining 17 different ensembles of feature ranking techniques (rankers) including six commonly-used feature ranking techniques, the signal-to-noise filter technique, and 11 threshold-based feature ranking techniques. This study utilized 16 real-world software measurement data sets of different sizes and built 13,600 classification models. Experimental results indicate that ensembles of very few rankers are very effective and even better than ensembles of many or all rankers.
Huanjing Wang, Taghi M. Khoshgoftaar, Amri Napolitano
ICMLA2
2010 Attribute Selection and Imbalanced Data: Problems in Software Defect Prediction
abstract
The data mining and machine learning community is often faced with two key problems: working with imbalanced data and selecting the best features for machine learning. This paper presents a process involving a feature selection technique for selecting the important attributes and a data sampling technique for addressing class imbalance. The application domain of this study is software engineering, more specifically, software quality prediction using classification models. When using feature selection and data sampling together, different scenarios should be considered. The four possible scenarios are: (1) feature selection based on original data, and modeling (defect prediction) based on original data; (2) feature selection based on original data, and modeling based on sampled data; (3) feature selection based on sampled data, and modeling based on original data; and (4) feature selection based on sampled data, and modeling based on sampled data. The research objective is to compare the software defect prediction performances of models based on the four scenarios. The case study consists of nine software measurement data sets obtained from the PROMISE software project repository. Empirical results suggest that feature selection based on sampled data performs significantly better than feature selection based on original data, and that defect prediction models perform similarly regardless of whether the training data was formed using sampled or original data.
Taghi M. Khoshgoftaar, Kehan Gao, Naeem Seliya
ICTAI (1)1
2010 Software Engineering with Computational Intelligence and Machine Learning A Novel Software Metric Selection Technique Using the Area Under ROC Curves
Taghi M. Khoshgoftaar, Kehan Gao
SEKE1
2010 Ensemble Feature Selection Technique for Software Quality Classification
Huanjing Wang, Taghi M. Khoshgoftaar, Kehan Gao
SEKE2
2010 Evolutionary data analysis for the class imbalance problem
abstract
Class imbalance, where the classes in a dataset are not represented equally, is a common occurrence in machine learning. Classification models built with such datasets are often not practical since most machine learning algorithms would tend to perfo
Taghi M. Khoshgoftaar, Naeem Seliya, Dennis J. Drown
Intell. Data Anal.1
2010 An Empirical Evaluation of Repetitive Undersampling Techniques
abstract
Class imbalance is a fundamental problem in data mining and knowledge discovery which is encountered in a wide array of application domains. Random undersampling has been widely used to alleviate the harmful effects of imbalance, however, this technique often leads to a substantial amount of information loss. Repetitive undersampling techniques, which generate an ensemble of models, each trained on a different, undersampled subset of the training data, have been proposed to allieviate this difficulty. This work reviews three repetitive undersampling methods currently used to handle imbalance and presents a detailed and comprehensive empirical study using four different learners, four performance metrics and 15 datasets from various application domains. To our knowledge, this work is the most thorough study of repetitive undersampling techniques.
Jason Van Hulse, Taghi M. Khoshgoftaar, Amri Napolitano
Int. J. Softw. Eng. Knowl. Eng.2
2010 Supervised neural network modeling: an empirical investigation into learning from imbalanced data with labeling errors
abstract
Neural network algorithms such as multilayer perceptrons (MLPs) and radial basis function networks (RBFNets) have been used to construct learners which exhibit strong predictive performance. Two data related issues that can have a detrimental impact on supervised learning initiatives are class imbalance and labeling errors (or class noise). Imbalanced data can make it more difficult for the neural network learning algorithms to distinguish between examples of the various classes, and class noise can lead to the formulation of incorrect hypotheses. Both class imbalance and labeling errors are pervasive problems encountered in a wide variety of application domains. Many studies have been performed to investigate these problems in isolation, but few have focused on their combined effects. This study presents a comprehensive empirical investigation using neural network algorithms to learn from imbalanced data with labeling errors. In particular, the first component of our study investigates the impact of class noise and class imbalance on two common neural network learning algorithms, while the second component considers the ability of data sampling (which is commonly used to address the issue of class imbalance) to improve their performances. Our results, for which over two million models were trained and evaluated, show that conclusions drawn using the more commonly studied C4.5 classifier may not apply when using neural networks.
Taghi M. Khoshgoftaar, Jason Van Hulse, Amri Napolitano
IEEE Trans. Neural Networks1
2010 Evolutionary Optimization of Software Quality Modeling with Multiple Repositories
abstract
A novel search-based approach to software quality modeling with multiple software project repositories is presented. Training a software quality model with only one software measurement and defect data set may not effectively encapsulate quality trends of the development organization. The inclusion of additional software projects during the training process can provide a cross-project perspective on software quality modeling and prediction. The genetic-programming-based approach includes three strategies for modeling with multiple software projects: Baseline Classifier, Validation Classifier, and Validation-and-Voting Classifier. The latter is shown to provide better generalization and more robust software quality models. This is based on a case study of software metrics and defect data from seven real-world systems. A second case study considers 17 different (nonevolutionary) machine learners for modeling with multiple software data sets. Both case studies use a similar majority-voting approach for predicting fault-proneness class of program modules. It is shown that the total cost of misclassification of the search-based software quality models is consistently lower than those of the non-search-based models. This study provides clear guidance to practitioners interested in exploiting their organization's software measurement data repositories for improved software quality modeling.
Yi Liu 0023, Taghi M. Khoshgoftaar, Naeem Seliya
IEEE Trans. Software Eng.2
2010 RUSBoost: A Hybrid Approach to Alleviating Class Imbalance
abstract
Class imbalance is a problem that is common to many application domains. When examples of one class in a training data set vastly outnumber examples of the other class(es), traditional data mining algorithms tend to create suboptimal classification models. Several techniques have been used to alleviate the problem of class imbalance, including data sampling and boosting. In this paper, we present a new hybrid sampling/boosting algorithm, called RUSBoost, for learning from skewed training data. This algorithm provides a simpler and faster alternative to SMOTEBoost, which is another algorithm that combines boosting and data sampling. This paper evaluates the performances of RUSBoost and SMOTEBoost, as well as their individual components (random undersampling, synthetic minority oversampling technique, and AdaBoost). We conduct experiments using 15 data sets from various application domains, four base learners, and four evaluation metrics. RUSBoost and SMOTEBoost both outperform the other procedures, and RUSBoost performs comparably to (and often better than) SMOTEBoost while being a simpler and faster technique. Given these experimental results, we highly recommend RUSBoost as an attractive alternative for improving the classification performance of learners built using imbalanced data.
Chris Seiffert, Taghi M. Khoshgoftaar, Jason Van Hulse, Amri Napolitano
IEEE Trans. Syst. Man Cybern. Part A2
2009 Wrapper-Based Feature Ranking for Software Engineering Metrics
abstract
The application of feature ranking to software engineering datasets is rare at best. In this study, we consider wrapper-based feature ranking where nine performance metrics aided by a particular learner are evaluated. We consider five learners and take two different approaches, each in conjunction with one of two different methodologies: 3-fold Cross-Validation (CV) and 3-fold Cross-Validation Risk Impact (CV-R). The classifiers are Naive Bayes (NB), Multi Layer Perceptron (MLP), k- Nearest Neighbors (kNN), Support Vector Machines (SVM), and Logistic Regression (LR). The performance metrics used as ranking techniques are Overall Accuracy (OA), F-Measure(FM), Geometric Mean (GM), Arithmetic Mean (AM), Area under ROC (AUC), Area under PRC (PRC), Best F-Measure (BFM), Best Geometric Mean (BGM), and Best Arithmetic Mean (BAM). To evaluate the classifier performance after feature selection has been applied, we use AUC as the performance evaluator. This paper represents a preliminary report on our proposed wrapper-based feature ranking approach to software defect prediction problems.
Wilker Altidor, Taghi M. Khoshgoftaar, Amri Napolitano
ICMLA2
2009 Feature Selection with Imbalanced Data for Software Defect Prediction
abstract
In this paper, we study the learning impact of data sampling followed by attribute selection on the classification models built with binary class imbalanced data within the scenario of software quality engineering. We use a wrapper-based attribute ranking technique to select a subset of attributes, and the random undersampling technique (RUS) on the majority class to alleviate the negative effects of imbalanced data on the prediction models. The datasets used in the empirical study were collected from numerous software projects. Five data preprocessing scenarios were explored in these experiments, including: (1) training on the original, unaltered fit dataset, (2) training on a sampled version of the fit dataset, (3) training on an unsampled version of the fit dataset using only the attributes chosen by feature selection based on the unsampled fit dataset, (4) training on an unsampled version of the fit dataset using only the attributes chosen by feature selection based on a sampled version of the fit dataset, and (5) training on a sampled version of the fit dataset using only the attributes chosen by feature selection based on the sampled version of the fit dataset. We compared the performances of the classification models constructed over these five different scenarios. The results demonstrate that the classification models constructed on the sampled fit data with or without feature selection (case 2 and case 5) significantly outperformed the classification models built with the other cases (unsampled fit data). Moreover, the two scenarios using sampled data (case 2 and case 5) showed very similar performances, but the subset of attributes (case 5) is only around 15% or 30% of the complete set of attributes (case 2).
Taghi M. Khoshgoftaar, Kehan Gao
ICMLA1
2009 An Empirical Study on Wrapper-Based Feature Ranking
abstract
Feature selection has become the cornerstone of many classification problems. It has been applied in many domains such as Web mining, text categorization, gene expression microarray analysis, image analysis, and combinatorial chemistry. One type of well-studied feature selection methodology is filtering, which is typically divided into ranking and subset evaluation. This work provides an empirical study regarding one type of feature ranking for which very limited research exists, namely wrapper-based feature ranking. Nine performance metrics are evaluated, and while these metrics are commonly used in data mining to evaluate classifier performance, they are rarely used as feature ranking techniques. Moreover, five different learners, 5-nearest neighbors (5NN), logistic regression (LR), multi layer perceptron (MLP), Naive Bayes (NB), and support vector machines (SVM) in conjunction with two different methodologies, 3-fold cross-validation (CV) and 3-fold cross-validation risk impact (CVR) are used in this study to evaluate feature relevancy and to determine ranking similarities among the different ranking techniques.
Wilker Altidor, Taghi M. Khoshgoftaar, Jason Van Hulse
ICTAI2
2009 Exploring Software Quality Classification with a Wrapper-Based Feature Ranking Technique
abstract
Feature selection is a process of selecting a subset of relevant features for building learning models. It is an important activity for data preprocessing used in software quality modeling and other data mining problems. Feature selection algorithms can be divided into two categories, feature ranking and feature subset selection. Feature ranking orders the features by a criterion and a user selects some of the features that are appropriate for a given scenario. Feature subset selection techniques search the space of possible feature subsets and evaluate the suitability of each. This paper investigates performance metric based feature ranking techniques by using the multilayer perceptron (MLP) learner with nine different performance metrics. The nine performance metrics include overall accuracy (OA), default F-measure (DFM), default geometric mean (DGM), default arithmetic mean (DAM), area under ROC (AUC), area under PRC (PRC), best F-measure (BFM), best geometric mean (BGM) and best arithmetic mean (BAM). The goal of the paper is to study the effect of the different performance metrics on the feature ranking results, which in turn influences the classification performance. We assessed the performance of the classification models constructed on those selected feature subsets through an empirical case study that was carried out on six data sets of real-world software systems. The results demonstrate that AUC, PRC, BFM, BGM and BAM as performance metrics for feature ranking outperformed the other performance metrics, OA, DFM, DGMand DAM, unanimously across all the data sets and therefore are recommended based on this study. In addition, the performances of the classification models were maintained or even improved when over 85 percent of the features were eliminated from the original data sets.
Kehan Gao, Taghi M. Khoshgoftaar, Amri Napolitano
ICTAI2
2009 A Study on the Relationships of Classifier Performance Metrics
abstract
There is no general consensus on which classifier performance metrics are better to use as compared to others. While some studies investigate a handful of such metrics in a comparative fashion, an evaluation of specific relationships among a large set of commonly-used performance metrics is much needed in the data mining and machine learning community. This study provides a unique insight into the underlying relationships among classifier performance metrics. We do so with a large case study involving 35 datasets from various domains and the C4.5 decision tree algorithm. A common property of the 35 datasets is that they suffer from the class imbalance problem. Our approach is based on applying factor analysis to the classifier performance space which is characterized by 22 performance metrics. It is shown that such a large number of performance metrics can be grouped into two-to-four relationship-based groups extracted by factor analysis. This work is a step in the direction of providing the analyst with an improved understanding about the different relationships and groupings among the performance metrics, thus facilitating the selection of performance metrics that capture relatively independent aspects of a classifier's performance.
Naeem Seliya, Taghi M. Khoshgoftaar, Jason Van Hulse
ICTAI2
2009 High-Dimensional Software Engineering Data and Feature Selection
abstract
Software metrics collected during project development play a critical role in software quality assurance. A software practitioner is very keen on learning which software metrics to focus on for software quality prediction. While a concise set of software metrics is often desired, a typical project collects a very large number of metrics. Minimal attention has been devoted to finding the minimum set of software metrics that have the same predictive capability as a larger set of metrics - we strive to answer that question in this paper. We present a comprehensive comparison between seven commonly-used filter-based feature ranking techniques (FRT) and our proposed hybrid feature selection (HFS) technique. Our case study consists of a very high-dimensional (42 software attributes) software measurement data set obtained from a large telecommunications system. The empirical analysis indicates that HFS performs better than FRT; however, the Kolmogorov-Smirnov feature ranking technique demonstrates competitive performance. For the telecommunications system, it is found that only 10% of the software attributes are sufficient for effective software quality prediction.
Huanjing Wang, Taghi M. Khoshgoftaar, Kehan Gao, Naeem Seliya
ICTAI2
2009 A Novel Hybrid Search Algorithm for Feature Selection
Pengpeng Lin, Huanjing Wang, Taghi M. Khoshgoftaar
SEKE3
2009 Value-Based Software Quality Modeling
Naeem Seliya, Taghi M. Khoshgoftaar
SEKE2
2009 An Extendible Translation of BPEL to a Machine-verifiable Model
John C. Sloan, Taghi M. Khoshgoftaar, Augusto Varas
SEKE2
2009 Knowledge discovery from imbalanced and noisy data
Jason Van Hulse, Taghi M. Khoshgoftaar
Data Knowl. Eng.2
2009 Software quality analysis by combining multiple projects and learners
Taghi M. Khoshgoftaar, Pierre Rebours, Naeem Seliya
Softw. Qual. J.1
2009 From Web Service Artifact to a Readable and Verifiable Model
abstract
Models of Web service compositions that are both readable and verifiable will benefit organizations that integrate purportedly reusable Web services. Colored Petri nets (CPNs) are at once verifiable and visually expressive, capable of presenting subtle flaws in service composition. Constructing CPN models from business process execution language (BPEL) artifacts had been a manual process requiring human judgment. Building on results from the workflow community, we automate the mapping of artifacts written in BPEL to models used by CPN Tools - a formal verification environment for development, simulation, and model checking of colored Petri nets. We extend related work that already converts BPEL to Petri nets, to reflect hierarchy and data type (color in CPN terminology), while improving model layout. We present a prototype implementation that mines both a BPEL artifact and the Petri net generated from it by an existing tool. The prototype partitions the Petri net into subnets, lays them out, colors them, and generates their XML file for import into CPN tools. Our results include depictions of subnets produced and initial simulation results for a well-known case study.
John C. Sloan, Taghi M. Khoshgoftaar
IEEE Trans. Serv. Comput.2
2009 Evolutionary Sampling and Software Quality Modeling of High-Assurance Systems
abstract
Software quality modeling for high-assurance systems, such as safety-critical systems, is adversely affected by the skewed distribution of fault-prone program modules. This sparsity of defect occurrence within the software system impedes training and performance of software quality estimation models. Data sampling approaches presented in data mining and machine learning literature can be used to address the imbalance problem. We present a novel genetic algorithm-based data sampling method, named evolutionary sampling, as a solution to improving software quality modeling for high-assurance systems. The proposed solution is compared with multiple existing data sampling techniques, including random undersampling, one-sided selection, Wilson's editing, random oversampling, cluster-based oversampling, synthetic minority oversampling technique (SMOTE), and borderline-SMOTE. This paper involves case studies of two real-world software systems and builds C4.5- and RIPPER-based software quality models both before and after applying a given data sampling technique. It is empirically shown that evolutionary sampling improves performance of software quality models for high-assurance systems and is significantly better than most existing data sampling techniques.
Dennis J. Drown, Taghi M. Khoshgoftaar, Naeem Seliya
IEEE Trans. Syst. Man Cybern. Part A2
2009 Empirical Case Studies in Attribute Noise Detection
abstract
The quality of data is an important issue in any domain-specific data mining and knowledge discovery initiative. The validity of solutions produced by data-driven algorithms can be diminished if the data being analyzed are of low quality. The quality of data is often realized in terms of data noise present in the given dataset and can include noisy attributes or labeling errors. Hence, tools for improving the quality of data are important to the data mining analyst. We present a comprehensive empirical investigation of our new and innovative technique for ranking attributes in a given dataset from most to least noisy. Upon identifying the noisy attributes, specific treatments can be applied depending on how the data are to be used. In a classification setting, for example, if the class label is determined to contain the most noise, processes to cleanse this important attribute may be undertaken. Independent variables or predictors that have a low correlation to the class attribute and appear noisy may be eliminated from the analysis. Several case studies using both real-world and synthetic datasets are presented in this study. The noise detection performance is evaluated by injecting noise into multiple attributes at different noise levels. The empirical results demonstrate conclusively that our technique provides a very accurate and useful ranking of noisy attributes in a given dataset.
Taghi M. Khoshgoftaar, Jason Van Hulse
IEEE Trans. Syst. Man Cybern. Part C1
2009 Improving Software-Quality Predictions With Data Sampling and Boosting
abstract
Software-quality data sets tend to fall victim to theclass-imbalanceproblem that plagues so many other application domains. The majority of faults in a software system, particularly high-assurance systems, usually lie in a very small percentage of the software modules. This imbalance between the number of fault-prone (fp) and non-fp (nfp) modules can have a severely negative impact on a data-mining technique's ability to differentiate between the two. This paper addresses the class-imbalance problem as it pertains to the domain of software-quality prediction. We present a comprehensive empirical study examining two different methodologies, data sampling and boosting, for improving the performance of decision-tree models designed to identify fp software modules. This paper applies five data-sampling techniques and boosting to 15 software-quality data sets of different sizes and levels of imbalance. Nearly 50 000 models were built for the experiments contained in this paper. Our results show that while data-sampling techniques are very effective in improving the performance of such models, boosting almost always outperforms even the best data-sampling techniques. This significant result, which, to our knowledge, has not been previously reported, has important consequences for practitioners developing software-quality classification models.
Chris Seiffert, Taghi M. Khoshgoftaar, Jason Van Hulse
IEEE Trans. Syst. Man Cybern. Part A2
2008 Software quality modeling: The impact of class noise on the random forest classifier
abstract
This study investigates the impact of increasing levels of simulated class noise on software quality classification. Class noise was injected into seven software engineering measurement datasets, and the performance of three learners, random forests, C4.5, and Naive Bayes, was analyzed. The random forest classifier was utilized for this study because of its strong performance relative to well-known and commonly-used classifiers such as C4.5 and Naive Bayes. Further, relatively little prior research in software quality classification has considered the random forest classifier. The experimental factors considered in this study were the level of class noise and the percent of minority instances injected with noise. The empirical results demonstrate that the random forest obtained the best and most consistent classification performance in all experiments.
Andres Folleco, Taghi M. Khoshgoftaar, Jason Van Hulse, Lofton A. Bullard
IEEE Congress on Evolutionary Computation2
2008 Comparison of Four Performance Metrics for Evaluating Sampling Techniques for Low Quality Class-Imbalanced Data
abstract
Erroneous attribute values can significantly impact learning from otherwise valuable data. The learning impact can be exacerbated by the class imbalanced training data. We investigate and compare the overall learning impact of sampling such data by using four distinct performance metrics suitable for models built from binary class imbalanced data. Seven relatively free of noise, class imbalanced software engineering measurement datasets were used. A novel noise injection procedure was applied to these datasets. We injected domain realistic noise into the independent and dependent (class) attributes of randomly selected instances to simulate lower quality measurement data. Seven well known data sampling techniques with the benchmark decision-tree learner C4.5 were used. No other related studies were found that have comprehensively investigated learning by sampling low quality binary class imbalanced data containing both independent and dependent corrupted attributes. Two sampling techniques (random undersampling and Wilson's editing) with better and more robust learning performances were identified. In contrast, all metrics concurred on the identification of the worst performing sampling technique (cluster-based oversampling).
Andres Folleco, Taghi M. Khoshgoftaar, Amri Napolitano
ICMLA2
2008 RUSBoost: Improving classification performance when training data is skewed
abstract
Constructing classification models using skewed training data can be a challenging task. We present RUSBoost, a new algorithm for alleviating the problem of class imbalance. RUSBoost combines data sampling and boosting, providing a simple and efficient method for improving classification performance when training data is imbalanced. In addition to performing favorably when compared to SMOTEBoost (another hybrid sampling/boosting algorithm), RUSBoost is computationally less expensive than SMOTEBoost and results in significantly shorter model training times. This combination of simplicity, speed and performance makes RUSBoost an excellent technique for learning from imbalanced data.
Chris Seiffert, Taghi M. Khoshgoftaar, Jason Van Hulse, Amri Napolitano
ICPR2
2008 VoB predictors: Voting on bagging classifications
abstract
Bagging predictors relies on bootstrap sampling to maintain a set of diverse base classifiers constituting the classifier ensemble, where the diversity among base classifiers is ensured through a random sampling (with replacement) process on the original data. In this paper, we propose a random missing value corruption based bootstrap sampling process, where the objective is to enhance the diversity of the learning sets through random missing value injection, such that base classifiers can form an accurate classifier ensemble. Our VoB (voting on bagging classifications) predictors first generate multiple incomplete datasets from a base complete dataset by randomly injecting missing values with a small missing ratio, then apply a bagging predictor trained on each of the incomplete dataset to give classifications. The final prediction of a class is the result of voting on the classifications. Our empirical results show that VoB predictors significantly improve the classification performance on complete data, and perform better than bagging predictors.
Xiaoyuan Su, Taghi M. Khoshgoftaar, Xingquan Zhu 0001
ICPR2
2008 Resampling or Reweighting: A Comparison of Boosting Implementations
abstract
Boosting has been shown to improve the performance of classifiers in many situations, including when data is imbalanced. There are, however, two possible implementations of boosting, and it is unclear which should be used. Boosting by reweighting is typically used, but can only be applied to base learners which are designed to handle example weights. On the other hand, boosting by resampling can be applied to any base learner. In this work, we empirically evaluate the differences between these two boosting implementations using imbalanced training data. Using 10 boosting algorithms, 4 learners and 15 datasets, we find that boosting by resampling performs as well as, or significantly better than, boosting by reweighting (which is often the default boosting implementation). We therefore conclude that in general, boosting by resampling is preferred over boosting by weighting.
Chris Seiffert, Taghi M. Khoshgoftaar, Jason Van Hulse, Amri Napolitano
ICTAI (1)2
2008 Improving Learner Performance with Data Sampling and Boosting
abstract
Learning from imbalanced datasets is a well known problem in the data mining community. Many techniques have been proposed to alleviate the problems associated with class imbalance, including data sampling and boosting. While data sampling has received the bulk of the attention from the research community, our results show that boosting often results in better classification performance than even the best data sampling techniques. In this work, we compare the performance of data sampling and boosting on ten datasets from various application domains using two commonly used learners. In addition, we propose the use of both data sampling and boosting in an attempt to combine the strengths of these techniques and achieve even better classification performance.
Chris Seiffert, Taghi M. Khoshgoftaar, Jason Van Hulse, Amri Napolitano
ICTAI (1)2
2008 Addressing Class Imbalance in Non-binary Classification Problems
abstract
The problem of class imbalance in machine learning is quite real and cumbersome when it comes to building a useful and practical classification model. We present a unique insight into addressing class imbalance for classification problems that involve three or more categories, i.e. non-binary. This study is different than related works in the literature because most works focus on addressing class imbalance only for binary classification problems, even if it means transforming a non-binary dataset into a binary classification problem. We propose an effective, yet simple approach to alleviating class imbalance issues when the classification problem involves more than two classes. The process, with four different methods, is based on applying random undersampling and random oversampling to different parts of the dataset for achieving better classification performance. The proposed data sampling methods are evaluated in the context of two real-world datasets obtained from the UCI Repository for Machine Learning Databases, and two commonly used classification algorithms: C4.5 and RIPPER. Our results demonstrate that the multi-group classification accuracy increases significantly in most cases after the proposed data sampling methods are applied. The positive outcome of this study motivates us to further our research on class imbalance and non-binary classification problems.
Naeem Seliya, Zhiwei Xu 0001, Taghi M. Khoshgoftaar
ICTAI (1)3
2008 Using Imputation Techniques to Help Learn Accurate Classifiers
abstract
It is difficult to learn good classifiers when training data is missing attribute values. Conventional techniques for dealing with such omissions, such as mean imputation, generally do not significantly improve the performance of the resulting classifier. We proposed imputation-helped classifiers, which use accurate imputation techniques, such as Bayesian multiple imputation (BMI), predictive mean matching (PMM), and Expectation Maximization (EM), as preprocessors for conventional machine learning algorithms. Our empirical results show that EM-helped and BMI-helped classifiers work effectively when the data is "missing completely at random", generally improving predictive performance over most of the original machine learned classifiers we investigated.
Xiaoyuan Su, Taghi M. Khoshgoftaar, Russell Greiner
ICTAI (1)2
2008 Analyzing the Impact of Attribute Noise on Software Quality Classification
Andres Folleco, Taghi M. Khoshgoftaar, Lofton A. Bullard
SEKE2
2008 On the Rarity of Fault-prone Modules in Knowledge-based Software Quality Modeling
Taghi M. Khoshgoftaar, Naeem Seliya, Dennis J. Drown
SEKE1
2008 Toward Model Checking Web Services Over the Web
John C. Sloan, Taghi M. Khoshgoftaar
SEKE2
2008 Imputed Neighborhood Based Collaborative Filtering
abstract
Collaborative filtering (CF) is one of the most effective types of recommender systems. As data sparsity remains a significant challenge for CF, we consider basing predictions on imputed data, and find this often improves performance on very sparse rating data. In this paper, we propose two imputed neighborhood based collaborative filtering (INCF) algorithms: imputed nearest neighborhood CF (INN-CF) and imputed densest neighborhood CF (IDN-CF), each of which first imputes the user rating data using an imputation technique, before using a traditional Pearson correlation-based CF algorithm on the resulting imputed data of the most similar neighbors or the densest neighbors to make CF predictions for a specific user. We compared an extension of Bayesian multiple imputation (eBMI) and the mean imputation (MEI) in these INCF algorithms, with the commonly-used neighborhood based CF, Pearson correlation-based CF, as well as a densest neighborhood based CF. Our empirical results show that IDN-CF using eBMI significantly outperforms its rivals and takes less time to make its best predictions.
Xiaoyuan Su, Taghi M. Khoshgoftaar, Russell Greiner
Web Intelligence2
2008 A comprehensive empirical evaluation of missing value imputation in noisy software measurement data
Jason Van Hulse, Taghi M. Khoshgoftaar
J. Syst. Softw.2
2008 Imputation techniques for multivariate missingness in software measurement data
Taghi M. Khoshgoftaar, Jason Van Hulse
Softw. Qual. J.1
2007 Arbitrarily-Shaped Window Based Stereo Matching using the Go-Light Optimization Algorithm
abstract
In this paper, we present a stereo matching algorithm using arbitrarily-shaped windows and a local optimization method called Go-light. The disparity map comes from a five-pixel arbitrarily-shaped window matching and a regular window based matching. It is then optimized by the Go-light optimization method, in which an outlier disparity value is replaced by the average of its surrounding ones when certain constraints are met. Experiments show that the accuracy of our algorithm is comparable to some of the state-of-the-art stereo correspondence algorithms on the Middlebury stereo data.
Xiaoyuan Su, Taghi M. Khoshgoftaar
ICIP (6)2
2007 Experimental perspectives on learning from imbalanced data
abstract
We present a comprehensive suite of experimentation on the subject of learning from imbalanced data. When classes are imbalanced, many learning algorithms can suffer from the perspective of reduced performance. Can data sampling be used to improve the performance of learners built from imbalanced data? Is the effectiveness of sampling related to the type of learner? Do the results change if the objective is to optimize different performance metrics? We address these and other issues in this work, showing that sampling in many cases will improve classifier performance.
Jason Van Hulse, Taghi M. Khoshgoftaar, Amri Napolitano
ICML2
2007 An application of a rule-based model in software quality classification
abstract
A new rule-based classification model (RBCM) and rulebased model selection technique are presented. The RBCM utilizes rough set theory to significantly reduce the number of attributes, discretation to partition the domain of attribute values, and Boolean predicates to generate the decision rules that comprise the model. When the domain values of an attribute are continuous and relatively large, rough set theory requires that they be discretized. The subsequent discretized domain must have the same characteristics as the original domain values. However, this can lead to a large number of partitions of the attribute's domain space, which in turn leads to large rule sets. These rule sets tend to form models that over-fit. To address this issue, the proposed rule-based model adopts a new model selection strategy that minimizes over-fitting for the RBCM. Empirical validation of the RBCM is accomplished through a case study on a large legacy telecommunications system. The results demonstrate that the proposed RBCM and the model selection strategy are effective in identifying the classification model that minimizes over-fitting and high cost classification errors. Keywords: rule-based classification model, rough set, reducts, discretization, software quality classification
Lofton A. Bullard, Taghi M. Khoshgoftaar, Kehan Gao
ICMLA2
2007 Using evolutionary sampling to mine imbalanced data
abstract
Class imbalance tends to cause inferior performance in data mining learners. Evolutionary sampling is a technique which seeks to counter this problem by using genetic algorithms to evolve a reduced sample of a complete dataset to train a classification model. Evolutionary sampling works to remove noisy and duplicate instances so that the sampled training data will produce a superior classifier. We propose this novel technique as a method to handle severe class imbalance in data mining. This paper presents our research into the the use of evolutionary sampling with C4.5 decision trees and compares the technique's performance with random undersamp ling.
Dennis J. Drown, Taghi M. Khoshgoftaar, Ramaswamy Narayanan
ICMLA2
2007 Learning with limited minority class data
abstract
A practical problem in data mining and machine learning is the limited availability of data. For example, in a binary classification problem it is often the case that examples of one class are abundant, while examples of the other class are in short supply. Examples from one class, typically the positive class, can be limited due to the financial cost or time required to collect these examples. This work presents a comprehensive empirical study of learning when examples from one class are extremely rare, but examples of the other class(es) are plentiful. Specifically, we address the issue of how many examples from the abundant class should be used when training a classifier on data where one class is very rare. Nearly one million classifiers were built and evaluated to generate the results presented in this work. Our results demonstrate that the often used 'even distribution' is not optimal when dealing with such rare events.
Taghi M. Khoshgoftaar, Chris Seiffert, Jason Van Hulse, Amri Napolitano, Andres Folleco
ICMLA1
2007 An Empirical Study of Learning from Imbalanced Data Using Random Forest
abstract
This paper discusses a comprehensive suite of experiments that analyze the performance of the random forest (RF) learner implemented in Weka. RF is a relatively new learner, and to the best of our knowledge, only preliminary experimentation on the construction of random forest classifiers in the context of imbalanced data has been reported in previous work. Therefore, the contribution of this study is to provide an extensive empirical evaluation of RF learners built from imbalanced data. What should be the recommended default number of trees in the ensemble? What should the recommended value be for the number of attributes? How does the RF learner perform on imbalanced data when compared with other commonly-used learners? We address these and other related issues in this work.
Taghi M. Khoshgoftaar, Moiz Golawala, Jason Van Hulse
ICTAI (2)1
2007 Mining Data with Rare Events: A Case Study
abstract
The performance of classification models can be negatively impacted if the data on which they are trained contains very rare events. While recent research has investigated the issue of class imbalance, few if any studies address issues related to the handling of extreme imbalance (rare events), where the minority class can account for as little as 0.1% of the training data. This work investigates the effect of dataset size and class distribution on classification performance when examples from the minority class are rare. In addition, we compare the performance improvement achieved by acquiring additional examples to that of applying data sampling. Our results demonstrate that data sampling is very effective at alleviating the problem of rare events.
Chris Seiffert, Taghi M. Khoshgoftaar, Jason Van Hulse, Amri Napolitano
ICTAI (2)2
2007 An Empirical Study of the Noise Impact on Cost-Sensitive Learning
Xingquan Zhu 0001, Xindong Wu 0001, Taghi M. Khoshgoftaar, Yong Shi 0001
IJCAI3
2007 Learning from Software Quality Data with Class Imbalance and Noise
Andres Folleco, Taghi M. Khoshgoftaar, Jason Van Hulse, Chris Seiffert
SEKE2
2007 Hybrid Collaborative Filtering Algorithms Using a Mixture of Experts
abstract
Collaborative filtering (CF) is one of the most successful approaches for recommendation. In this paper, we propose two hybrid CF algorithms, sequential mixture CF and joint mixture CF, each combining advice from multiple experts for effective recommendation. These proposed hybrid CF models work particularly well in the common situation when data are very sparse. By combining multiple experts to form a mixture CF, our systems are able to cope with sparse data to obtain satisfactory performance. Empirical studies show that our algorithms outperform their peers, such as memory-based, pure model-based, pure content-based CF algorithms, and the content- boosted CF (a representative hybrid CF algorithm), especially when the underlying data are very sparse.
Xiaoyuan Su, Russell Greiner, Taghi M. Khoshgoftaar, Xingquan Zhu 0001
Web Intelligence3
2007 The multiple imputation quantitative noise corrector
Taghi M. Khoshgoftaar, Jason Van Hulse, Chris Seiffert
Intell. Data Anal.1
2007 Improving Software Quality Prediction by Noise Filtering Techniques
Taghi M. Khoshgoftaar, Pierre Rebours
J. Comput. Sci. Technol.1
2007 The pairwise attribute noise detection algorithm
Jason Van Hulse, Taghi M. Khoshgoftaar
Knowl. Inf. Syst.2
2007 Editorial: Special issue on mining low-quality data
Xingquan Zhu 0001, Taghi M. Khoshgoftaar, Ian Davidson, Shichao Zhang 0001
Knowl. Inf. Syst.2
2007 Software quality estimation with limited fault data: a semi-supervised learning perspective
Naeem Seliya, Taghi M. Khoshgoftaar
Softw. Qual. J.2
2007 A Comprehensive Empirical Study of Count Models for Software Fault Prediction
abstract
Count models, such as the Poisson regression model, and the negative binomial regression model, can be used to obtain software fault predictions. With the aid of such predictions, the development team can improve the quality of operational software. The zero-inflated, and hurdle count models may be more appropriate when, for a given software system, the number of modules with faults are very few. Related literature lacks quantitative guidance regarding the application of count models for software quality prediction. This study presents a comprehensive empirical investigation of eight count models in the context of software fault prediction. It includes comparative hypothesis testing, model selection, and performance evaluation for the count models with respect to different criteria.
Kehan Gao, Taghi M. Khoshgoftaar
IEEE Trans. Reliab.2
2007 Count Models for Software Quality Estimation
abstract
Identifying which software modules, during the software development process, are likely to be faulty is an effective technique for improving software quality. Such an approach allows a more focused software quality & reliability enhancement endeavor. The development team may also like to know the number of faults that are likely to exist in a given program module, i.e., a quantitative quality prediction. However, classification techniques such as the logistic regression model (lrm) cannot be used to predict the number of faults. In contrast, count models such as the Poisson regression model (prm), and the zero-inflated Poisson (zip) regression model can be used to obtain both a qualitative classification, and a quantitative prediction for software quality. In the case of the classification models, a classification rule based on our previously developed generalized classification rule is used. In the context of count models, this study is the first to propose a generalized classification rule. Case studies of two industrial software systems are examined, and for each we developed two count models, (prm, and zip), and a classification model (lrm). Evaluating the predictive capabilities of the models, we concluded that the prm, and the zip models have similar classification accuracies as the lrm. The count models are also used to predict the number of faults for the two case studies. The zip model yielded better fault prediction accuracy than the prm. As compared to other quantitative prediction models for software quality, such as multiple linear regression (mlr), the prm, and zip models have a unique property of yielding the probability that a given number of faults will occur in any module
Taghi M. Khoshgoftaar, Kehan Gao
IEEE Trans. Reliab.1
2007 A Multi-Objective Software Quality Classification Model Using Genetic Programming
abstract
A key factor in the success of a software project is achieving the best-possible software reliability within the allotted time & budget. Classification models which provide a risk-based software quality prediction, such as fault-prone & not fault-prone, are effective in providing a focused software quality assurance endeavor. However, their usefulness largely depends on whether all the predicted fault-prone modules can be inspected or improved by the allocated software quality-improvement resources, and on the project-specific costs of misclassifications. Therefore, a practical goal of calibrating classification models is to lower the expected cost of misclassification while providing a cost-effective use of the available software quality-improvement resources. This paper presents a genetic programming-based decision tree model which facilitates a multi-objective optimization in the context of the software quality classification problem. The first objective is to minimize the "Modified Expected Cost of Misclassification", which is our recently proposed goal-oriented measure for selecting & evaluating classification models. The second objective is to optimize the number of predicted fault-prone modules such that it is equal to the number of modules which can be inspected by the allocated resources. Some commonly used classification techniques, such as logistic regression, decision trees, and analogy-based reasoning, are not suited for directly optimizing multi-objective criteria. In contrast, genetic programming is particularly suited for the multi-objective optimization problem. An empirical case study of a real-world industrial software system demonstrates the promising results, and the usefulness of the proposed model
Taghi M. Khoshgoftaar, Yi Liu 0023
IEEE Trans. Reliab.1
2007 Software Quality Analysis of Unlabeled Program Modules With Semisupervised Clustering
abstract
Software quality assurance is a vital component of software project development. A software quality estimation model is trained using software measurement and defect (software quality) data of a previously developed release or similar project. Such an approach assumes that the development organization has experience with systems similar to the current project and that defect data are available for all modules in the training data. In software engineering practice, however, various practical issues limit the availability of defect data for modules in the training data. In addition, the organization may not have experience developing a similar system. In such cases, the task of software quality estimation or labeling modules as fault prone or not fault prone falls on the expert. We propose a semisupervised clustering scheme for software quality analysis of program modules with no defect data or quality-based class labels. It is a constraint-based semisupervised clustering scheme that uses k-means as the underlying clustering algorithm. Software measurement data sets obtained from multiple National Aeronautics and Space Administration software projects are used in our empirical investigation. The proposed technique is shown to aid the expert in making better estimations as compared to predictions made when the expert labels the clusters formed by an unsupervised learning algorithm. In addition, the software quality knowledge learnt during the semisupervised process provided good generalization performance for multiple test data sets. An analysis of program modules that remain unlabeled subsequent to our semisupervised clustering scheme provided useful insight into the characteristics of their software attributes
Naeem Seliya, Taghi M. Khoshgoftaar
IEEE Trans. Syst. Man Cybern. Part A2
2006 A Comparison of Software Fault Imputation Procedures
abstract
This work presents a detailed comparison of three imputation techniques, Bayesian multiple imputation, regression imputation and k nearest neighbor imputation, at various missingness levels. Starting with a complete real-world software measurement dataset called CCCS, missing values were injected into the dependent variable at four levels according to three different missingness mechanisms. The three imputation techniques are evaluated by comparing the imputed and actual values. Our analysis includes a three-way analysis of variance (ANOVA) model, which demonstrates that Bayesian multiple imputation obtains the best performance, followed closely by regression
Jason Van Hulse, Taghi M. Khoshgoftaar, Chris Seiffert
ICMLA2
2006 Assessment of a Multi-Strategy Classifier for an Embedded Software System
abstract
In this paper, a new classification model, RB2CBL, is proposed. Its structure and methodology are described. By cascading a rule-based (RB) model with two case-based learning (CBL) models, RB2CBL possesses the merits of both RB model and CBL model and restrains their drawbacks. In the RB2CBL model, the parameter optimization of the CBL models is essential, and the embedded genetic algorithm optimizer is used. In our case study, a dataset collected from initial releases of two large, Windowscopy-based embedded system applications, which were used primarily for customizing the configuration of wireless telecommunications products, is processed to investigate and evaluate the models. The results show that, by suitably choosing accuracy settings of the RB model, RB2CBL model outperforms the RB model alone without overfitting. In practice, the RB2CBL model effectively reduced the misclassification rates and improved prediction accuracy for the embedded software system
Taghi M. Khoshgoftaar, Kehan Gao
ICTAI1
2006 A Hybrid Approach to Cleansing Software Measurement Data
abstract
Data is extremely important in empirical software engineering. Techniques that provide insight into potential anomalies or inaccuracies in a dataset are becoming an increasingly important way for a data analyst to cope with flawed data. We present a novel hybrid procedure for quantitative outcome correction along with controlled experiments using a real-world software measurement dataset to demonstrate the usefulness of our technique. Instances that are deemed to be noisy relative to the dependent variable, which represents the number of faults recorded in the program module, are cleansed by replacing the original value with a more appropriate alternative value
Taghi M. Khoshgoftaar, Jason Van Hulse, Chris Seiffert
ICTAI1
2006 Collaborative Filtering for Multi-class Data Using Belief Nets Algorithms
abstract
As one of the most successful recommender systems, collaborative filtering (CF) algorithms can deal with high sparsity and high requirement of scalability amongst other challenges. Bayesian belief nets (BNs), one of the most frequently used classifiers, can be used for CF tasks. Previous works of applying BNs to CF tasks were mainly focused on binary-class data, and used simple or basic Bayesian classifiers (Miyahara and Pazzani, 2002; Breese et al., 1998). In this work, we apply advanced BNs models to CF tasks instead of simple ones, and work on real-world multi-class CF data instead of synthetic binary-class data. Empirical results show that with their ability to deal with incomplete data, extended logistic regression on naive Bayes and tree augmented naive Bayes (NB-ELR and TAN-ELR) models (Greiner et al., 2005) consistently perform better than the state-of-the-art Pearson correlation-based CF algorithm. In addition, the ELR-optimized BNs CF models are robust in terms of the ability to make predictions, while the robustness of the Pearson correlation-based CF algorithm degrades as the sparseness of the data increases
Xiaoyuan Su, Taghi M. Khoshgoftaar
ICTAI2
2006 Multiple Imputation of Software Measurement Data: A Case Study
Taghi M. Khoshgoftaar, Jason Van Hulse
SEKE1
2006 Polishing Noise in Continuous Software Measurement Data
Taghi M. Khoshgoftaar, Chris Seiffert, Jason Van Hulse
SEKE1
2006 Class noise detection using frequent itemsets
Jason Van Hulse, Taghi M. Khoshgoftaar
Intell. Data Anal.2
2006 Determining noisy instances relative to attributes of interest
Taghi M. Khoshgoftaar, Jason Van Hulse
Intell. Data Anal.1
2006 Detecting Noisy Instances with the Ensemble Filter: a Study in Software Quality Estimation
abstract
The performance of a classification model is invariably affected by the characteristics of the measurement data it is built upon. If the quality of the data is generally poor, then the classification model will demonstrate poor performance. The detection and removal of noisy instances will improve quality of the data, and consequently, the performance of the classification model. We investigate a noise handling technique that attempts to improve the quality of datasets for classification purposes by eliminating instances that are likely to be noise. Our approach uses twenty five different classification techniques to create an ensemble filter for eliminating likely noise. The basic assumption is that if a given majority of classifiers in the ensemble misclassify an instance, then it is likely to be a noisy instance. Using a relatively large number of base-level classifiers in the ensemble filter facilitates in achieving the desired level of noise removal conservativeness with several possible levels of filtering. It also provides a higher degree of confidence in the noise elimination procedure as the results are less likely to get influenced by (possibly) inappropriate learning bias of a few algorithms with twenty five base-level classifiers than with relatively smaller number of base-level classifiers. Empirical case studies of two high assurance software projects demonstrates the effectiveness of our noise elimination approach by the significant improvement achieved in classification accuracies at various levels of noise filtering.
Taghi M. Khoshgoftaar, Vedang H. Joshi, Naeem Seliya
Int. J. Softw. Eng. Knowl. Eng.1
2006 Resource oriented selection of rule-based classification models: An empirical case study
Taghi M. Khoshgoftaar, Angela Herzberg, Naeem Seliya
Softw. Qual. J.1
2006 An empirical study of predicting software faults with case-based reasoning
Taghi M. Khoshgoftaar, Naeem Seliya, Nandini Sundaresh
Softw. Qual. J.1
2006 Unsupervised multiscale color image segmentation based on MDL principle
abstract
We present an unsupervised multiscale color image segmentation algorithm. The basic idea is to apply mean shift clustering to obtain an over-segmentation and then merge regions at multiple scales to minimize the minimum description length criterion. The performance on the Berkeley segmentation benchmark campares favorably with some existing approaches.
Qiming Luo, Taghi M. Khoshgoftaar
IEEE Trans. Image Process.2
2005 Identifying noise in an attribute of interest
abstract
One of the most significant issues facing the data mining community is that of low-quality data. Real-world datasets are often inundated with various types of data integrity issues, particularly noisy data. In response to the difficulties created by low-quality data, we propose a novel technique to detect noisy instances relative to an attribute of interest (AOI). Any attribute in the dataset can be defined by the user as the attribute of interest. A noise ranking of instances relative to the chosen attribute is output. This approach can be iterated for any number of user-specified attributes of interest. The case study described in this work demonstrates how our technique may be used to detect class noise, which occurs when errors are present in the class or dependent variable. In this scenario the class is declared to be the attribute of interest and an instance noise ranking relative to the class is provided. Our technique is compared to the well-known ensemble and classification filters which have been previously proposed for class noise detection. The results of this study demonstrate the effectiveness of our approach and show that our procedure is a useful tool for improving data quality.
Taghi M. Khoshgoftaar, Jason Van Hulse
ICMLA1
2005 Intrusion detection in wireless networks using clustering techniques with expert analysis
abstract
The increasing reliance upon wireless networks has put tremendous' emphasis on wireless network security. While considerable attention has been given to data mining for intrusion detection in wired networks, limited focus has been devoted to data mining for intrusion detection in wireless networks. This study presents a clustering approach with tracers and expert analysis for intrusion detection in a real-world wireless network. Security vulnerabilities of 802.11 wireless networks are investigated, leading to a summary of network traffic metrics relevant to modeling the security of wireless networks. The proposed approach utilizes a simple distance-based heuristic measure to label clusters as either normal or intrusive. The classification of network traffic instances is further enhanced with the aid of tracers, i.e., a small set of instances with known labels - normal or intrusive. Our study demonstrates the usefulness and promise of the proposed approach, laying the groundwork for a clustering-based framework for intrusion detection in wireless computer networks.
Taghi M. Khoshgoftaar, Shyam Varan Nath, Shi Zhong 0001, Naeem Seliya
ICMLA1
2005 A Clustering Approach to Wireless Network Intrusion Detection
abstract
Intrusion detection in wireless networks has become an indispensable component of any useful wireless network security systems, and has recently gained attention in both research and industry communities due to widespread use of wireless local area networks (WLANs). This paper focuses on detecting intrusions or anomalous behaviors in WLANs with data clustering techniques. We first explore the security vulnerabilities of 802.11 or Wi-Fi networks and summarize the network traffic metrics that are important to model the security of wireless networks. Based on the metrics studied we propose a clustering-based intrusion detection approach and evaluate it on a real-world large wireless network traffic dataset. The evaluation results demonstrate the effectiveness of our proposed intrusion detection approach for wireless networks.
Shi Zhong 0001, Taghi M. Khoshgoftaar, Shyam Varan Nath
ICTAI2
2005 Assessment of a New Three-Group Software Quality Classification Technique: An Empirical Case Study
Taghi M. Khoshgoftaar, Naeem Seliya, Kehan Gao
Empir. Softw. Eng.1
2005 Evaluating indirect and direct classification techniques for network intrusion detection
Taghi M. Khoshgoftaar, Kehan Gao, Nawal H. Ibrahim
Intell. Data Anal.1
2005 Identifying noisy features with the Pairwise Attribute Noise Detection Algorithm
Taghi M. Khoshgoftaar, Jason Van Hulse
Intell. Data Anal.1
2005 Evaluating noise elimination techniques for software quality estimation
Taghi M. Khoshgoftaar, Pierre Rebours
Intell. Data Anal.1
2005 Detecting noisy instances with the rule-based classification model
Taghi M. Khoshgoftaar, Naeem Seliya, Kehan Gao
Intell. Data Anal.1
2005 Enhancing software quality estimation using ensemble-classifier based noise filtering
Taghi M. Khoshgoftaar, Shi Zhong 0001, Vedang H. Joshi
Intell. Data Anal.1
2005 Resource-oriented software quality classification models
Taghi M. Khoshgoftaar, Naeem Seliya, Angela Herzberg
J. Syst. Softw.1
2004 Efficient Image Segmentation by Mean Shift Clustering and MDL-Guided Region Merging
abstract
We present an efficient color and texture segmentation algorithm by combining two statistical techniques: mean shift clustering and minimum description length (MDL) principle. Mean shift clustering is proven in generating robust and accurate segmentation results for color images, but the selection of the two scale parameters remains a challenging problem for images with texture. Optimization based on MDL principle requires little parameter tuning, but the initial input has a strong impact on its efficiency and effectiveness. Our approach is to apply mean shift clustering to generate an initial over-segmentation and then merge regions based on MDL principle. Objects with texture can be extracted with reasonable accuracy by merging regions under the guidance of MDL principle, without the need of convolving the image with a bank of filters. Experimental results on a variety of natural scene images are reported and compared with the JSEG algorithm. It takes about 1 second for our algorithm to process a 320/spl times/240 color image on a conventional PC.
Qiming Luo, Taghi M. Khoshgoftaar
ICTAI2
2004 Semi-Supervised Learning for Software Quality Estimation
abstract
A software quality estimation model is often built using known software metrics and fault data obtained from program modules of previously developed releases or similar projects. Such a supervised learning approach to software quality estimation assumes that fault data is available for all the previously developed modules. Considering the various practical issues in software project development, fault data may not be available for all the software modules in the training data. More specifically, the available labeled training data is such that a supervised learning approach may not yield good software quality prediction. In contrast, a supervised classification scheme aided by unlabeled data, i.e., semisupervised learning, may yield better results. This work investigates semisupervised learning with the expectation maximization (EM) algorithm for the software quality classification problem. Case studies of software measurement data obtained from two NASA software projects, JM1 and KC2, are used in our empirical investigation. A small portion of the JM1 dataset is randomly extracted and used as the labeled data, while the remaining JM1 instances are used as unlabeled data. The performance of the semisupervised classification models built using the EM algorithm is evaluated by using the KC2 project as a test dataset. It is shown that the EM-based semisupervised learning scheme improves the predictive accuracy of the software quality classification models.
Naeem Seliya, Taghi M. Khoshgoftaar, Shi Zhong 0001
ICTAI2
2004 Noise Identification with the k-Means Algorithm
abstract
The presence of noise in a measurement dataset can have a negative effect on the classification model built. More specifically, the noisy instances in the dataset can adversely affect the learnt hypothesis. Removal of noisy instances will improve the learnt hypothesis; thus, improving the classification accuracy of the model. A clustering-based noise detection approach using the k-means algorithm is presented. We present a new metric for measuring the potentiality (noise factor) of an instance being noisy. Based on the computed noise factor values of the instances, the clustering-based algorithm is then used to identify and eliminate p% of the instances in the dataset. These p% of instances are considered the most likely to be noisy among the instances in the dataset - the p% value is varied from 1% to 40%. The noise detection approach is investigated with respect to two case studies of software measurement data obtained from NASA software projects. The two datasets are characterized by the same thirteen software metrics and a class label that classifies the program modules as fault-prone and not fault-prone. It is shown that as more noisy instances are removed, classification accuracy of the C4>5 learner improves. This indicates that the removed instances are most likely noisy instances that attributed to poor classification accuracy.
Taghi M. Khoshgoftaar
ICTAI2
2004 Noise Elimination with Ensemble-Classifier Filtering: A Case-Study in Software Quality Engineerin
Taghi M. Khoshgoftaar, Vedang H. Joshi
SEKE1
2004 Multi-Objective Optimization by CBR GA-Optimizer for Module-Order Modeling
Taghi M. Khoshgoftaar, Yudong Xiao, Kehan Gao
SEKE1
2004 Comparative Assessment of Software Quality Classification Techniques: An Empirical Case Study
Taghi M. Khoshgoftaar, Naeem Seliya
Empir. Softw. Eng.1
2004 Identification of fuzzy models of software cost estimation
Zhiwei Xu 0001, Taghi M. Khoshgoftaar
Fuzzy Sets Syst.2
2004 A multiobjective module-order model for software quality enhancement
abstract
The knowledge, prior to system operations, of which program modules are problematic is valuable to a software quality assurance team, especially when there is a constraint on software quality enhancement resources. A cost-effective approach for allocating such resources is to obtain a prediction in the form of a quality-based ranking of program modules. Subsequently, a module-order model (MOM) is used to gauge the performance of the predicted rankings. From a practical software engineering point of view, multiple software quality objectives may be desired by a MOM for the system under consideration: e.g., the desired rankings may be such that 100% of the faults should be detected if the top 50% of modules with highest number of faults are subjected to quality improvements. Moreover, the management team for the same system may also desire that 80% of the faults should be accounted if the top 20% of the modules are targeted for improvement. Existing work related to MOM(s) use a quantitative prediction model to obtain the predicted rankings of program modules, implying that only the fault prediction error measures such as the average, relative, or mean square errors are minimized. Such an approach does not provide a direct insight into the performance behavior of a MOM. For a given percentage of modules enhanced, the performance of a MOM is gauged by how many faults are accounted for by the predicted ranking as compared with the perfect ranking. We propose an approach for calibrating a multiobjective MOM using genetic programming. Other estimation techniques, e.g., multiple linear regression and neural networks cannot achieve multiobjective optimization for MOM(s). The proposed methodology facilitates the simultaneous optimization of multiple performance objectives for a MOM. Case studies of two industrial software systems are presented, the empirical results of which demonstrate a new promise for goal-oriented software quality modeling.
Taghi M. Khoshgoftaar, Yi Liu 0023, Naeem Seliya
IEEE Trans. Evol. Comput.1
2003 Building Decision Tree Software Quality Classification Models Using Genetic Programming
Yi Liu 0023, Taghi M. Khoshgoftaar
GECCO2
2003 Detecting Outliers Using Rule-Based Modeling for Improving CBR-Based Software Quality Classification Models
Taghi M. Khoshgoftaar, Lofton A. Bullard, Kehan Gao
ICCBR1
2003 Genetic Programming-Based Decision Trees for Software Quality Classification
abstract
The knowledge of the likely problematic areas of a software system is very useful for improving its overall quality. Based on such information, a more focused software testing and inspection plan can be devised. Decision trees are attractive for a software quality classification problem which predicts the quality of program modules in terms of risk-based classes. They provide a comprehensible classification model which can be directly interpreted by observing the tree-structure. A simultaneous optimization of the classification accuracy and the size of the decision tree is a difficult problem, and very few studies have addressed the issue. This paper presents an automated and simplified genetic programming (gp) based decision tree modeling technique for the software quality classification problem. Genetic programming is ideally suited for problems that require optimization of multiple criteria. The proposed technique is based on multi-objective optimization using strongly typed GP. In the context of an industrial high-assurance software system, two fitness functions are used for the optimization problem: one for minimizing the average weighted cost of misclassification, and one for controlling the size of the decision tree. The classification performances of the GP-based decision trees are compared with those based on standard GP, i.e., S-expression tree. It is shown that the GP-based decision tree technique yielded better classification models. As compared to other decision tree-based methods, such as C4.5, GP-based decision trees are more flexible and can allow optimization of performance objectives other than accuracy. Moreover, it provides a practical solution for building models in the presence of conflicting objectives, which is commonly observed in software development practice.
Taghi M. Khoshgoftaar, Yi Liu 0023, Naeem Seliya
ICTAI1
2003 Application of an Attribute Selection Method to CBR-Based Software Quality Classification
abstract
This study investigates the attribute selection problem for reducing the number of software metrics (program attributes) used by a case-based reasoning (CBR) software quality classification model. The metrics are selected using the Kolmogorov-Smirnov (K-S) two sample test. The "modified expected cost of misclassification" measure, recently proposed by our research team, is used as a performance measure to select, evaluate, and compare classification models. The attribute selection procedure presented in this paper can assist a software development organization in determining the software metrics that are better indicators of software quality. By reducing the number of software metrics to be collected during the development process, the metrics data collection task can be simplified. Moreover, reducing the number of metrics would result in reducing the computation time of a CBR model. Using an empirical case study of a real-world software system, it is shown that with a reduced number of metrics the CBR technique is capable of yielding useful software quality classification models. Moreover, their performances were better than or similar to CBR models calibrated without attribute selection.
Taghi M. Khoshgoftaar, Laurent A. Nguyen, Kehan Gao, Jayanth Rajeevalochanam
ICTAI1
2003 Fault Prediction Modeling for Software Quality Estimation: Comparing Commonly Used Techniques
Taghi M. Khoshgoftaar, Naeem Seliya
Empir. Softw. Eng.1
2003 Analogy-Based Practical Classification Rules for Software Quality Estimation
Taghi M. Khoshgoftaar, Naeem Seliya
Empir. Softw. Eng.1
2003 Application of fuzzy expert systems in assessing operational risk of software
Zhiwei Xu 0001, Taghi M. Khoshgoftaar, Edward B. Allen
Inf. Softw. Technol.2
2003 Introduction to the Special Issue on Quality Engineering with Computational Intelligence
Taghi M. Khoshgoftaar
Softw. Qual. J.1
2003 Ordering Fault-Prone Software Modules
Taghi M. Khoshgoftaar, Edward B. Allen
Softw. Qual. J.1
2002 Can neural networks be easily interpreted in software cost estimation?
abstract
Software development effort estimation with the aid of neural networks has generally been viewed with skepticism by a majority of the software cost estimation community. Although, neural networks have shown their strengths in solving complex problems, their shortcoming of being 'black boxes' models has prevented them from being accepted as a common practice for cost estimation. In this paper, we study the interpretation of cost estimation models based on a backpropagation three layer perceptron network. Our proposed idea comprises mainly of the use of a method that maps this neural network to a fuzzy rule based system. Consequently, if the obtained fuzzy rules are easily interpreted, the neural network will also be easy to interpret. Our case study is based on the COCOMO'81 dataset.
Ali Idri, Taghi M. Khoshgoftaar, Alain Abran
FUZZ-IEEE2
2002 Software Quality Classification Modeling Using The SPRINT Decision Tree Algorithm
abstract
Predicting the quality of system modules prior to software testing and operations can benefit the software development team. Such a timely reliability estimation can be used to direct cost-effective quality improvement efforts to the high-risk modules. Tree-based software quality classification models based on software metrics are used to predict whether a software module is fault-prone or not fault-prone. They are white box quality estimation models with good accuracy, and are simple and easy to interpret. This paper presents an in-depth study of calibrating classification trees for software quality estimation using the SPRINT decision tree algorithm. Many classification algorithms have memory limitations including the requirement that data sets be memory resident. SPRINT removes all of these limitations and provides a fast and scalable analysis. It is an extension of a commonly used decision tree algorithm, CART, and provides a unique tree-pruning technique based on the minimum description length (MDL) principle. Combining the MDL pruning technique and the modified classification algorithm, SPRINT yields classification trees with useful prediction accuracy. The case study used comprises of software metrics and fault data collected over four releases from a very large telecommunications system. It is observed that classification trees built by SPRINT are more balanced and demonstrate better stability in comparison to those built by CART.
Taghi M. Khoshgoftaar, Naeem Seliya
ICTAI1
2002 Improving Usefulness of Software Quality Classification Models Based on Boolean Discriminant Functions
abstract
BDF (Boolean discriminant functions) are an attractive technique for software quality estimation. Software quality classification models based on BDF provide stringent rules for classifying not fault-prone modules (nfp), thereby predicting a large number of modules as fp. Such models are practically not useful from software quality assurance and software management points of view. This is because, given the large number of modules predicted as fp, project management will face a difficult task of deploying, cost-effectively, the always-limited reliability improvement resources to all the fp modules. This paper proposes the use of generalized Boolean discriminant functions (GBDF) as a solution for improving the practical and managerial usefulness of classification models based on BDF. In addition, the use of GBDF avoids the need to build complex hybrid classification models in order to improve usefulness of models based on BDF. A case study of a full-scale industrial software system is presented to illustrate the promising results obtained from using the proposed classification technique using GBDF.
Taghi M. Khoshgoftaar
ISSRE1
2002 Uncertain Classification of Fault-Prone Software Modules
Taghi M. Khoshgoftaar, Xiaojing Yuan, Edward B. Allen, Wendell D. Jones, John P. Hudepohl
Empir. Softw. Eng.1
2002 Predicting Fault-Prone Modules in Embedded Systems Using Analogy-Based Classification Models
abstract
Embedded systems have become ubiquitous and essential entities in our ever growing high-tech world. The backbone of today's information-highway infrastructure are embedded systems such as telecommunication systems. They demand high reliability, so as to prevent severe consequences of failures including costly repairs at remote sites. Technology changes mandate that embedded systems evolve, resulting in a demand for techniques for improving reliability of their future system releases. Reliability models based on software metrics can be effective tools for software engineering of embedded systems, because quality improvements are so resource-consuming that it is not feasible to apply them to all modules. Identification of the likely fault-prone modules before system testing, can be effective in reducing the likelihood of faults discovered during operations. A software quality classification model is calibrated using software metrics from a past release, and is then applied to modules currently under development to estimate which modules are likely to be fault-prone. This paper presents and demonstrates an effective case-based reasoning approach for calibrating such classification models. It is attractive for software engineering of embedded systems, because it can be used to develop software reliability models using a faster, cheaper, and easier method. We illustrate our approach with two large-scale case studies obtained from embedded systems. They involve data collected from telecommunication systems including wireless systems. It is indicated that the level of classification accuracy observed in both case studies would be beneficial in achieving high software reliability of subsequent releases of the embedded systems.
Taghi M. Khoshgoftaar, Bojan Cukic, Naeem Seliya
Int. J. Softw. Eng. Knowl. Eng.1
2002 Using regression trees to classify fault-prone software modules
abstract
Software faults are defects in software modules that might cause failures. Software developers tend to focus on faults, because they are closely related to the amount of rework necessary to prevent future operational software failures. The goal of this paper is to predict which modules are fault-prone and to do it early enough in the life cycle to be useful to developers. A regression tree is an algorithm represented by an abstract tree, where the response variable is a real quantity. Software modules are classified as fault-prone or not, by comparing the predicted value to a threshold. A classification rule is proposed that allows one to choose a preferred balance between the two types of misclassification rates. A case study of a very large telecommunications systems considered software modules to be fault-prone, if any faults were discovered by customers. Our research shows that classifying fault-prone modules with regression trees and the using the classification rule in this paper, resulted in predictions with satisfactory accuracy and robustness.
Taghi M. Khoshgoftaar, Edward B. Allen, Jianyu Deng
IEEE Trans. Reliab.1
2001 An Application of Zero-Inflated Poisson Regression for Software Fault Prediction
abstract
Poisson regression model is widely used in software quality modeling. When the response variable of a data set includes a large number of zeros, Poisson regression model will underestimate the probability of zeros. A zero-inflated model changes the mean structure of the pure Poisson model. The predictive quality is therefore improved. In this paper, we examine a full-scale industrial software system and develop two models, Poisson regression and zero-inflated Poisson regression. To our knowledge, this is the first study that introduces the zero-inflated Poisson regression model in software reliability. Comparing the predictive qualities of the two competing models, we conclude that for this system, the zero-inflated Poisson regression model is more appropriate in theory and practice.
Taghi M. Khoshgoftaar, Kehan Gao, Robert M. Szabo
ISSRE1
2001 Software Quality Prediction for High-Assurance Network Telecommunications Systems
abstract
Modern high-assurance network telecommunications systems often must have high software reliability. Software quality models yield timely predictions of quality indicators on a module-by-module basis, enabling one to target enhancement techniques. This paper introduces fuzzy nonlinear regression (FNR) to practitioners in high-assurance systems engineering. FNR integrates fuzzy logic and neural networks techniques to generate good results. We present an innovative method that differs from other FNRs such that the statistics of the dependent variable are used to build the FNR instead of focusing on the cost function. We tested our model in a large network telecommunications system written in a Pascal-like proprietary language. Data splitting was applied in the case study. In order to avoid the unfair partition of the data set, we randomly split the data thirty times and performed the experiment for each iteration. We also conducted our experiment using multiple linear regression (MLR) and we found that the results of FNR were significantly better than those using MLR
Zhiwei Xu 0001, Taghi M. Khoshgoftaar
Comput. J.2
2001 Controlling Overfitting in Classification-Tree Models of Software Quality
Taghi M. Khoshgoftaar, Edward B. Allen
Empir. Softw. Eng.1
2001 Empirical Assessment of a Software Metric: The Information Content of Operators
Taghi M. Khoshgoftaar, Edward B. Allen
Softw. Qual. J.1
2001 Cost-Benefit Analysis of Software Quality Models
Taghi M. Khoshgoftaar, Edward B. Allen, Wendell D. Jones, John P. Hudepohl
Softw. Qual. J.1
2001 Data Mining of Software Development Databases
Taghi M. Khoshgoftaar, Edward B. Allen, Wendell D. Jones, John P. Hudepohl
Softw. Qual. J.1
2000 Modeling software quality: the Software Measurement Analysis and Reliability Toolkit
abstract
The paper presents the Software Measurement Analysis and Reliability Toolkit (SMART) which is a research tool for software quality modeling using case based reasoning (CBR) and other modeling techniques. Modern software systems must have high reliability. Software quality models are tools for guiding reliability enhancement activities to high risk modules for maximum effectiveness and efficiency. A software quality model predicts a quality factor, such as the number of faults in a module, early in the life cycle in time for effective action. Software product and process metrics can be the basis for such fault predictions. Moreover, classification models can identify fault prone modules. CBR is an attractive modeling method based on automated reasoning processes. However, to our knowledge, few CBR systems for software quality modeling have been developed. SMART addresses this area. There are currently three types of models supported by SMART: classification based on CBR, CBR classification extended with cluster analysis, and module-order models, which predict the rank-order of modules according to a quality factor. An empirical case study of a military command, control, and communications applied SMART at the end of coding. The models built by SMART had a level of accuracy that could be very useful to software developers.
Taghi M. Khoshgoftaar, Edward B. Allen, Jason C. Busboom
ICTAI1
2000 Improving Tree-Based Models of Software Quality with Principal Components Analysis
abstract
Software quality classification models can predict which modules are to be considered fault-prone, and which are not, based on software product metrics, process metrics and execution metrics. Such predictions can be used to target improvement efforts to those modules that need them the most. Classification-tree modeling is a robust technique for building such software quality models. However, the model structure may be unstable, and accuracy may suffer when the predictors are highly correlated. This paper presents an empirical case study of four releases of a very large telecommunications system, which shows that the tree-based models can be improved by transforming the predictors with principal components analysis, so that the transformed predictors are not correlated. The case study used the regression-tree algorithm in the S-Plus package and then applied a general decision rule to classify the modules.
Taghi M. Khoshgoftaar, Ruqun Shan, Edward B. Allen
ISSRE1
2000 Modeling Fault-Prone Modules of Subsystems
abstract
Software developers are very interested in targeting software enhancement activities prior to release, so that reworking of faulty modules can be avoided. Credible predictions of which modules are likely to have faults discovered by customers can be the basis for selecting modules for enhancement. Many case studies in the literature build models to predict which modules will be fault-prone without regard to the subsystems defined by the system's functional architecture. Our hypothesis is this: models that are specially built for subsystems will be more accurate than a system-wide model applied to each subsystem's modules. In other words, the subsystem that a module belongs to can be valuable information in software quality modeling. This paper presents an empirical case study which compared software quality models of an entire system to models of a major functional subsystem. The study, modeled a very large telecommunications system with classification trees built by the CART (classification and regression trees) algorithm. For predicting subsystem quality, we found that a model built with training data on the subsystem alone was more accurate than a similar model built with training data on the entire system. We concluded that the characteristics of the subsystem's modules were not similar to those of the system as a whole, and thus, information on subsystems can be valuable.
Taghi M. Khoshgoftaar, Vishal Thaker, Edward B. Allen
ISSRE1
2000 Balancing Misclassification Rates in Classification-Tree Models of Software Quality
Taghi M. Khoshgoftaar, Xiaojing Yuan, Edward B. Allen
Empir. Softw. Eng.1
2000 Case-Based Software Quality Prediction
abstract
Highly reliable software is becoming an essential ingredient in many systems. However, assuring reliability often entails time-consuming costly development processes. One cost-effective strategy is to target reliability-enhancement activities to those modules that are likely to have the most problems. Software quality prediction models can predict the number of faults expected in each module early enough for reliability enhancement to be effective. This paper introduces a case-based reasoning technique for the prediction of software quality factors. Case-based reasoning is a technique that seeks to answer new problems by identifying similar "cases" from the past. A case-based reasoning system can function as a software quality prediction model. To our knowledge, this study is the first to use case-based reasoning systems for predicting quantitative measures of software quality. A case study applied case-based reasoning to software quality modeling of a family of full-scale industrial software systems. The case-based reasoning system's accuracy was much better than a corresponding multiple linear regression model in predicting the number of design faults. When predicting faults in code, its accuracy was significantly better than a corresponding multiple linear regression model for two of three test data sets and statistically equivalent for the third.
K. Ganesan 0001, Taghi M. Khoshgoftaar, Edward B. Allen
Int. J. Softw. Eng. Knowl. Eng.2
2000 A practical classification-rule for software-quality models
abstract
A practical classification rule for a SQ (software quality) model considers the needs of the project to use a model to guide targeting software RE (reliability enhancement) efforts, such as extra reviews early in development. Such a rule is often more useful than alternative rules. This paper discusses several classification rules for SQ models, and recommends a generalized classification rule, where the effectiveness and efficiency of the model for guiding software RE efforts can be explicitly considered. This is the first application of this rule to SQ modeling that we know of. Two case studies illustrate application of the generalized classification rule. A telecommunication-system case-study models membership in the class of fault-prone modules as a function of the number of interfaces to other modules. A military-system case-study models membership in the class of fault-prone modules as a function of a set of process metrics that depict the development history of a module. These case studies are examples where balanced misclassification rates resulted in more useful and practical SQ models than other classification rules.
Taghi M. Khoshgoftaar, Edward B. Allen
IEEE Trans. Reliab.1
2000 Classification-tree models of software-quality over multiple releases
abstract
This paper presents an empirical study that evaluates software-quality models over several releases, to address the question, "How long will a model yield useful predictions?" The classification and regression trees (CART) algorithm is introduced, CART can achieve a preferred balance between the two types of misclassification rates. This is desirable because misclassification of fault-prone modules often has much more severe consequences than misclassification of those that are not fault-prone. The case-study developed 2 classification-tree models based on 4 consecutive releases of a very large legacy telecommunication system. Forty-two software product, process and execution metrics were candidate predictors. Model 1 used measurements of the first release as the training data set; this model had 11 important predictors. Model 2 used measurements of the second release as the training data set; this model had 15 important predictors. Measurements of subsequent releases were evaluation data sets. Analysis of the models' predictors yielded insights into various software development practices. Both models had accuracy that would be useful to developers. One might suppose that software-quality models lose their value very quickly over successive releases due to evolution of the product and the underlying development processes. The authors found the models remained useful over all the releases studied.
Taghi M. Khoshgoftaar, Edward B. Allen, Wendell D. Jones, John P. Hudepohl
IEEE Trans. Reliab.1
1999 Modelling software quality with GP
Matthew P. Evett, Taghi M. Khoshgoftaar, Pei-der Chien, Edward B. Allen
GECCO2
1999 Experience Paper: Preparing Measurements of Legacy Software for Predicting Operational Faults
abstract
Software quality modeling can be used by a software maintenance project to identify a limited set of software modules that probably need improvement. A model's goal is to recommend a set of modules to receive special treatment. The purpose of the paper is to report our experiences modeling software quality with classification trees, including necessary preprocessing of data. We conducted a case study on two releases of a very large legacy telecommunications system. A module was considered fault-prone if any faults were discovered by customers, and not fault-prone otherwise. Software product, process, and execution metrics were the basis for predictors. The TREEDISC algorithm for building classification trees was investigated, because it emphasizes statistical significance. Numeric data, such as software metrics, are not suitable for TREEDISC. Consequently, we transformed measurements into discrete ordinal predictors by grouping. This case study investigated the sensitivity of modeling results to various groupings. We found that robustness, accuracy, and parsimony of the models were influenced by the maximum number of groups. Models based on two sets of candidate predictors had similar sensitivity.
Taghi M. Khoshgoftaar, Edward B. Allen, Xiaojing Yuan, Wendell D. Jones, John P. Hudepohl
ICSM1
1999 Classification tree models of software quality over multiple releases
abstract
Software quality models are tools for focusing software enhancement efforts. Such efforts are essential for mission-critical embedded software, such as telecommunications systems, because customer-discovered faults have very serious consequences and are very expensive to repair. We present an empirical study that evaluated software quality models over several releases to address the question, "How long will a model yield useful predictions?" We also introduce the Classification And Regression Trees (CART) algorithm to software reliability engineering practitioners. We present our method for exploiting CART features to achieve a preferred balance between the two types of misclassification rates. This is desirable because misclassifications of fault-prone modules often have much more severe consequences than misclassifications of those that are not fault-prone. We developed two classification-tree models based on four consecutive releases of a very large legacy telecommunications system. Forty-two software product, process, and execution metrics were candidate predictors. The first software quality model used measurements of the first release as the training data set and measurements of the subsequent three releases as evaluation data sets. The second model used measurements of the second release as the training data set and measurements of the subsequent two releases as evaluation data sets. Both models had accuracy that would be useful to developers.
Taghi M. Khoshgoftaar, Edward B. Allen, Wendell D. Jones, John P. Hudepohl
ISSRE1
1999 A Comparative Study of Ordering and Classification of Fault-Prone Software Modules
Taghi M. Khoshgoftaar, Edward B. Allen
Empir. Softw. Eng.1
1999 Data Mining for Predictors of Software Quality
abstract
"Knowledge discovery in data bases" (KDD) for software engineering is a process for finding useful information in the large volumes of data that are a byproduct of software development, such as data bases for configuration management and for problem reporting. This paper presents guidelines for extracting innovative process metrics from these commonly available data bases. This paper also adapts the Classification And Regression Trees algorithm, CART, to the KDD process for software engineering data. To our knowledge, this algorithm has not been used previously for empirical software quality modeling. In particular, we present an innovative way to control the balance between misclassification rates. A KDD case study of a very large legacy telecommunications software system found that variables derived from source code, configuration management transactions, and problem reporting transactions can be useful predictors of software quality. The KDD process discovered that for this software development environment, out of forty software attributes, only a few of the predictor variables were significant. This resulted in a model that predicts whether modules are likely to have faults discovered by customers. Software developers need such predictions early in development to target software enhancement techniques to the modules that need improvement the most.
Taghi M. Khoshgoftaar, Edward B. Allen, Wendell D. Jones, John P. Hudepohl
Int. J. Softw. Eng. Knowl. Eng.1
1999 Using Classification Trees for Software Quality Models: Lessons Learned
abstract
High software reliability is an important attribute of high-assurance systems. Software quality models yield timely predictions of quality indicators on a module-by-module basis, enabling one to focus on finding faults early in development. This paper introduces the Classification And Regression Trees (CART) a algorithm to practitioners in high-assurance systems engineering. This paper presents practical lessons learned on building classification trees for software quality modeling, including an innovative way to control the balance between misclassification rates. A case study of a very large telecommunications system used CART to build software quality models. The models predicted whether or not modules would have faults discovered by customers, based on various sets of software product and process metrics as independent variables. We found that a model based on two software product metrics had comparable accuracy to a model based on forty product and process metrics.
Taghi M. Khoshgoftaar, Edward B. Allen, Archana Naik, Wendell D. Jones, John P. Hudepohl
Int. J. Softw. Eng. Knowl. Eng.1
1999 Which Software Modules have Faults which will be Discovered by Customers?
abstract
Software is the medium for implementing increasingly sophisticated features during the maintenance phase as successive releases are developed. Software quality models can predict which software modules are likely to have faults that will be discovered by customers. Such models are key components of a system such as Enhanced Measurement for Early Risk Assessment of Latent Defects (EMERALD). It is a sophisticated system of decision support tools used by software designers and managers at Nortel to assess risk and improve software quality and reliability of legacy software systems. This paper reports an approach to software quality modelling that is suitable for industrial systems such as EMERALD. We conducted a case study of a large legacy telecommunications system in the maintenance phase to predict whether each module will be considered fault-prone. The case study is distinctive in the following respects. (1) Fault-prone modules were defined in terms of faults discovered by customers, which represent only a small fraction of the modules in the system. (2) We developed models based on software product and process metrics that can make useful predictions at the end of the coding phase and at the time of release. (3) The modelling approach is suitable for very large systems. We anticipate that refinements of this case study's models will be incorporated into EMERALD. A similar approach could be taken for other systems. Copyright © 1999 John Wiley & Sons, Ltd.
Taghi M. Khoshgoftaar, Edward B. Allen, Wendell D. Jones, John P. Hudepohl
J. Softw. Maintenance Res. Pract.1
1998 Hitting the Moving Target: Trials and Tribulations of Modeling Quality in Evolving Software Systems
Wendell D. Jones, Taghi M. Khoshgoftaar, John C. Munson, T. Troy Pearse, George E. Stark
ICSM2
1998 Can a Software Quality Model Hit a Moving Target?
abstract
This paper examines factors that make accurate quality modeling of an evolving software system challenging. The context of our discussion is development of a sequence of software releases. Our goal is to predict software quality of a release early enough co make significant quality improvements prior to release.
Taghi M. Khoshgoftaar, Edward B. Allen
ICSM1
1998 Predicting the order of fault-prone modules in legacy software
abstract
A goal of software quality modeling is to recommend modules for reliability enhancement early enough to prevent poor quality. Reliability improvement techniques include more rigorous design and code reviews and more extensive testing. This paper introduces the concept of module-order models for guiding software reliability enhancement and provides an empirical case study that shows how such models can be used. A module-order model predicts the rank-order of modules according to a quantitative quality factor. The case study examined a large legacy telecommunications system. We found that the amount of new and changed code due to the development of a release can be a better predictor of code churn due to subsequent bug fixes, compared to software product metrics alone. In such projects, process-related measures derived from configuration management data may be adequate for software quality modeling, without resorting to software product measurement tools and expertise.
Taghi M. Khoshgoftaar, Edward B. Allen
ISSRE1
1998 Classification of Fault-Prone Software Modules: Prior Probabilities, Costs, and Model Evaluation
Taghi M. Khoshgoftaar, Edward B. Allen
Empir. Softw. Eng.1
1997 Evolutionary neural networks: a robust approach to software reliability problems
abstract
In this empirical study, from a large data set of software metrics for program modules, thirty distinct partitions into training and validation sets are automatically generated with approximately equal distributions of fault prone and not fault prone modules. Thirty classification models are built for each of the two approaches considered-discriminant analysis and the evolutionary neural network (ENN) approach-and their performances on corresponding data sets are compared. The lower error proportions for ENNs on fault prone, not fault prone, and overall classification were found to be statistically significant. The robustness of ENNs follows from their superior performance on the range of data configurations used. It is suggested that ENNs can be effective in other software reliability problem domains, where they have been largely ignored.
Robert Hochman, Taghi M. Khoshgoftaar, Edward B. Allen, John P. Hudepohl
ISSRE2
1997 Predicting fault-prone modules with case-based reasoning
abstract
Software quality classification models seek to predict quality factors such as whether a module will be fault prone, or not. Case based reasoning (CBR) is a modeling technique that seeks to answer new questions by identifying similar "cases" from the past. When applied to software reliability, the working hypothesis of our approach is this: a module currently under development is probably fault prone if a module with similar product and process attributes in an earlier release was fault prone. The contribution of the paper is application of case based reasoning to software quality modeling. To the best of our knowledge, this is the first time that case based reasoning has been used to identify fault prone modules. A case study illustrates our approach and provides evidence that case based reasoning can be the basis for useful software quality classification models that are competitive with discriminant models. The case study revisits data from a previously published nonparametric discriminant analysis study. The Type II misclassification rate of the CBR model was substantially better than that of the discriminant model. Although the Type I misclassification rate was slightly greater and the overall misclassification rate was only slightly less, the CBR model was preferred when costs of misclassification were considered.
Taghi M. Khoshgoftaar, K. Ganesan 0001, Edward B. Allen, Fletcher D. Ross, Rama Munikoti, Nishith Goel, Amit Nandi
ISSRE1
1997 An Information Theory-Based Approach to Quantifying the Contribution of a Software Metric
Taghi M. Khoshgoftaar, Edward B. Allen, David L. Lanning
J. Syst. Softw.1
1997 Application of neural networks to software quality modeling of a very large telecommunications system
abstract
Society relies on telecommunications to such an extent that telecommunications software must have high reliability. Enhanced measurement for early risk assessment of latent defects (EMERALD) is a joint project of Nortel and Bell Canada for improving the reliability of telecommunications software products. This paper reports a case study of neural-network modeling techniques developed for the EMERALD system. The resulting neural network is currently in the prototype testing phase at Nortel. Neural-network models can be used to identify fault-prone modules for extra attention early in development, and thus reduce the risk of operational problems with those modules. We modeled a subset of modules representing over seven million lines of code from a very large telecommunications software system. The set consisted of those modules reused with changes from the previous release. The dependent variable was membership in the class of fault-prone modules. The independent variables were principal components of nine measures of software design attributes. We compared the neural-network model with a nonparametric discriminant model and found the neural-network model had better predictive accuracy.
Taghi M. Khoshgoftaar, Edward B. Allen, John P. Hudepohl, Stephen J. Aud
IEEE Trans. Neural Networks1
1996 Detection of Fault-Prone Software Modules During a Spiral Life Cycle
abstract
The article is an experience report on identifying fault prone modules in a subsystem of the Joint Surveillance Target Attack Radar System, JSTARS, a large tactical military system. The project followed the spiral life cycle model. The iterations of the system were developed in FORTRAN about one year apart. We developed a discriminant analysis model using software metrics from one iteration to predict whether or not each module in the next would be considered fault prone. Tactical military software is required to have high reliability. Each software function is often considered mission critical, and the lives of military personnel often depend on mission success. In our project, each iteration of a spiral life cycle development produced a system that was suitable for operational testing. A risk analysis based on operational testing guided development of the next iteration. Identifying fault prone modules early in the development of an iteration can lead to better reliability, The results confirm previously published studies that discriminant analysis can be a useful tool in identification of fault prone software modules. This study used consecutive iterations, first, to build, and then to evaluate the model. This model validation approach is more realistic than earlier studies which split data from one project to simulate two iterations. Model results could be used to identify those modules that would probably benefit from earlier reviews and testing, and thus, reduce the risk of unexpected problems with those modules.
Taghi M. Khoshgoftaar, Edward B. Allen, Robert Halstead 0001, Gary P. Trio
ICSM1
1996 Using the genetic algorithm to build optimal neural networks for fault-prone module detection
abstract
The genetic algorithm is applied to developing optimal or near optimal backpropagation neural networks for fault-prone/not-fault-prone classification of software modules. The algorithm considers each network in a population of neural networks as a potential solution to the optimal classification problem. Variables governing the learning and other parameters and network architecture are represented as substrings (genes) in a machine-level bit string (chromosome). When the population undergoes simulated evolution using genetic operators-selection based on a fitness function, crossover, and mutation-the average performance increases in successive generations. We found that, on the same data, compared with the best manually developed networks, evolved networks produced improved classifications in considerably less time, with no human effort, and with greater confidence in their optimality or near optimality. Strategies for devising a fitness function specific to the problem are explored and discussed.
Robert Hochman, Taghi M. Khoshgoftaar, Edward B. Allen, John P. Hudepohl
ISSRE2
1996 Integrating metrics and models for software risk assessment
abstract
Enhanced Measurement for Early Risk Assessment of Latent Defects (EMERALD) is a decision support system for assessing reliability risk. It is used by software developers and managers to improve telecommunications software service quality as perceived by the customer and the end user. Risk models are based on static characteristics of source code. This paper shows how a system such as EMERALD can enhance software development, testing, and maintenance by integration of: a software quality improvement strategy; measurements and models; and delivery of results to the desktop of developers in a timely manner. This paper also summarizes empirical experiments with EMERALD's models using data from large industrial telecommunications software systems. EMERALD has been applied to a very large system with over 12 million lines of source code within procedures. Experience and lessons learned are also discussed.
John P. Hudepohl, Stephen J. Aud, Taghi M. Khoshgoftaar, Edward B. Allen, Jean Mayrand
ISSRE3
1996 Detection of software modules with high debug code churn in a very large legacy system
abstract
Society has become so dependent on reliable telecommunications, that failures can risk loss of emergency service, business disruptions, or isolation from friends. Consequently, telecommunications software is required to have high reliability. Many previous studies define the classification fault prone in terms of fault counts. This study defines fault prone as exceeding a threshold of debug code churn, defined as the number of lines added or changed due to bug fixes. Previous studies have characterized reuse history with simple categories. This study quantified new functionality with lines of code. The paper analyzes two consecutive releases of a large legacy software system for telecommunications. We applied discriminant analysis to identify fault prone modules based on 16 static software product metrics and the amount of code changed during development. Modules from one release were used as a fit data set and modules from the subsequent release were used as a test data set. In contrast, comparable prior studies of legacy systems split the data to simulate two releases. We validated the model with a realistic simulation of utilization of the fitted model with the test data set. Model results could be used to give extra attention to fault prone modules and thus, reduce the risk of unexpected problems.
Taghi M. Khoshgoftaar, Edward B. Allen, Nishith Goel, Amit Nandi, John McMullan
ISSRE1
1996 The impact of software evolution and reuse on software quality
Taghi M. Khoshgoftaar, Edward B. Allen, Kalai Kalaichelvan, Nishith Goel
Empir. Softw. Eng.1
1996 Analysis and differentiation of software system environments
Taghi M. Khoshgoftaar, David L. Lanning
Softw. Qual. J.1
1996 Using neural networks to predict software faults during testing
abstract
This paper investigates the application of principal components analysis to neural-network modeling. The goal is to predict the number of faults. Ten software product measures were gathered from a large commercial software system. Principal components were then extracted from these measures. We trained two neural networks, one with the observed (raw) data, and one with principal components. We compare the predictive quality of the two competing models using data collected from two similar systems. These systems were developed by the same organization, and used the same development process. For the environment we studied, applying principal-components analysis to the raw data yields a neural-network model whose predictive quality is statistically better than a neural-network model developed using the raw data alone. The improvement in model predictive quality is appreciable from a practitioner's point of view. We concur with published literature regarding the number of hidden layers needed in a neural-network model. A single hidden layer of neurons yielded a network of sufficient generality to be useful when predicting faults. This is important, because networks with more hidden layers take correspondingly more time to train. The application of alternative network architectures and training algorithms in software engineering should continue to be investigated.
Taghi M. Khoshgoftaar, Robert M. Szabo
IEEE Trans. Reliab.1
1995 Multivariate assessment of complex software systems: a comparative study
abstract
Assessment of large complex systems requires robust modeling techniques. Multivariate models can be misleading if the underlying metrics are highly correlated. Munson and Khoshgoflaar propose using principal components analysis to avoid such problems. Even though many have used the technique, the advantages have not previously been empirically demonstrated, especially for large complex systems. Our case study illustrates that principal components analysis can substantially improve the predictive quality of a software quality model. This paper presents a case study of a sample of modules representing about 1.3 million lines of code, taken from a much larger real-time telecommunications system. This study used discriminant analyse's for classification of fault-prone modules, based on measurements of software design attributes and categorical variables indicating new, changed, and reused modules. Quality of fit and predictive quality were evaluated.
Taghi M. Khoshgoftaar, Edward B. Allen
ICECCS1
1995 Detecting program modules with low testability
abstract
We model the relationship between static software product measures and a dynamic quality measure, testability. To our knowledge, this is the first time a dynamic quality measure has been modeled using static software product measures. We first give an overview of testability analysis and discriminant modeling. Using static software product measures collected from a real time avionics software system, we develop two discriminant models and classify the component program modules as having low or high testability. The independent variables are principal components derived from the observed software product measures. One model is used to evaluate the quality of fit and one is used to assess classification performance. We show that for this study, the quality of fit and classification performance of the discriminant modeling methodology are excellent and yield a potentially useful insight into the relationship between static software measures and testability.
Taghi M. Khoshgoftaar, Robert M. Szabo, Jeffrey M. Voas
ICSM1
1995 Detection of fault-prone program modules in a very large telecommunications system
abstract
Telecommunications software is known for its high reliability. Society has become so accustomed to reliable telecommunications, that failures can cause major disruptions. This is an experience report on application of discriminant analysis based on 20 static software product metrics, to identify fault prone modules in a large telecommunications system, so that reliability may be improved. We analyzed a sample of 2000 modules representing about 1.3 million lines of code, drawn from a much larger system. Sample modules were randomly divided into a fit data set and a test data set. We simulated utilization of the fitted model with the test data set. We found that identifying new modules and changed modules mere significant components of the discriminant model, and improved its performance. The results demonstrate that data on module reuse is a valuable input to quality models and that discriminant analysis can be a useful tool in early identification of fault prone software modules in large telecommunications systems. Model results could be used to identify those modules that would probably benefit from extra attention, and thus, reduce the risk of unexpected problems with those modules.
Taghi M. Khoshgoftaar, Edward B. Allen, Kalai Kalaichelvan, Nishith Goel, John P. Hudepohl, Jean Mayrand
ISSRE1
1995 An assessment of software quality in a C++ environment
abstract
In this experience report, we discuss the application of a three group discriminant model to an object oriented software system. The model is used to classify software modules as high, medium, or low risk with respect to the number of faults. To our knowledge, this is among the first empirical validations of a true three group software quality model. Two models are presented. One model incorporates only traditional measures while the other model includes both traditional and object oriented measures. We show that for this system, the addition of the object oriented measures enhances the model by reducing the overall misclassification rate and significantly reducing the misclassification in the medium group.
Robert M. Szabo, Taghi M. Khoshgoftaar
ISSRE2
1995 A neural network approach for early detection of program modules having high risk in the maintenance phase
Taghi M. Khoshgoftaar, David L. Lanning
J. Syst. Softw.1
1995 A Performance Analysis of an Object-Based I/O Architecture in a Video Server Environment
Khoa D. Huynh, Taghi M. Khoshgoftaar
Multim. Syst.2
1995 Performance Analysis of a Peer-to-Peer I/O Architecture in Video Server Environments
Khoa D. Huynh, Taghi M. Khoshgoftaar
Multim. Tools Appl.2
1995 Investigating ARIMA models of software system quality
Taghi M. Khoshgoftaar, Robert M. Szabo
Softw. Qual. J.1
1994 Improving Code Churn Predictions During the System Test and Maintenance Phases
abstract
We show how to improve the prediction of gross change using neural networks. We select a multiple regression quality model from the principal components of software complexity metrics collected from a large commercial software system at the beginning of the testing phase. Our measure of quality is based on gross change, and is collected at the end of the maintenance phase. This quality measure is attractive for study as it is both objective and easily obtained directly from the source code. Then, we train a neural network with the complete set of principal components. Comparisons of the two models, gathered from eight related software systems, shows that the neural network offers much improved predictive quality over the multiple regression model.>
Taghi M. Khoshgoftaar, Robert M. Szabo
ICSM1
1994 Canonical Modeling of Software Complexity and Fault Correction Activity
abstract
The study applies canonical correlation analysis to investigate the relationship between source code complexity and fault correction activity. Product and process measures collected during the development of a commercial real-time product provide the data for this analysis. Sets of variables represent source code complexity and fault correction activity. Significant canonical correlations along two dimensions support the hypothesis that source code complexity exerted a causal influence on fault correction activity during the system test phase of the real-time product. Interpretation of the two significant canonical correlations reveals relationships between the sets of variables that are not immediately apparent from their simple correlations.>
David L. Lanning, Taghi M. Khoshgoftaar
ICSM2
1994 On the impact of software product dissimilarity on software quality models
abstract
The current software market favors software development organizations that apply software quality models. Software engineers fit quality models to data collected from past projects. Predictions from these models provide guidance in setting schedules and allocating resources for new and ongoing development projects. To improve model stability and predictive quality, engineers select models from the orthogonal linear combinations produced using principal components analysis. However, recent research revealed that the principal components underlying source code measures are not necessarily stable across software products. Thus, the principal components underlying the product used to fit a regression model can vary from the principal components underlying the product for which we desire predictions. We investigate the impact of this principal components instability on the predictive quality of regression models. To achieve this, we apply an analytical technique for accessing the aptness of a given model to a particular application.>
Taghi M. Khoshgoftaar, David L. Lanning
ISSRE1
1994 A comparative study of pattern recognition techniques for quality evaluation of telecommunications software
abstract
The extreme risks of software faults in the telecommunications environment justify the costs of data collection and modeling of software quality. Software quality models based on data drawn from past projects can identify key risk or problem areas in current similar development efforts. Once these problem areas are identified, the project management team can take actions to reduce the risks. Studies of several telecommunications systems have found that only 4-6% of the system modules were complex /spl lsqb/LeGall et al. 1990/spl rsqb/. Since complex modules are likely to contain a large proportion of a system's faults, the approach of focusing resources on high-risk modules seems especially relevant to telecommunications software development efforts. A number of researchers have recognized this, and have applied modeling techniques to isolate fault-prone or high-risk program modules. A classification model based upon discriminant analytic techniques has shown promise in performing this task. The authors introduce a neural network classification model for identifying high-risk program modules, and compare the quality of this model with that of a discriminant classification model fitted with the same data. They find that the neural network techniques provide a better management tool in software engineering environments. These techniques are simpler, produce more accurate models, and are easier to use.>
Taghi M. Khoshgoftaar, David L. Lanning, Abhijit S. Pandya
IEEE J. Sel. Areas Commun.1
1994 Alternative approaches for the use of metrics to order programs by complexity
Taghi M. Khoshgoftaar, John C. Munson, David L. Lanning
J. Syst. Softw.1
1994 Performance Analysis of Advanced I/O Architectures for PC-Based Video Servers
Khoa D. Huynh, Taghi M. Khoshgoftaar
Multim. Syst.2
1994 A Performance Analysis of Personal Computers in a Video Conferencing Environment
Khoa D. Huynh, Taghi M. Khoshgoftaar
Multim. Syst.2
1993 A Comparative Study of Predictive Models for Program Changes During System Testing and Maintenance
abstract
By modeling the relationship between software complexity attributes and software quality attributes, software engineers can take actions early in the development cycle to control the cost of the maintenance phase. The effectiveness of these model-based actions depends heavily on the predictive quality of the model. An enhanced modeling methodology that shows significant improvements in the predictive quality of regression models developed to predict software changes during maintenance is applied here. The methodology reduces software complexity data to domain metrics by applying principal components analysis. It then isolates clusters of similar program modules by applying cluster analysis to these derived domain metrics. Finally, the methodology develops individual regression models for each cluster. These within-cluster models have better predictive quality than a general model fitted to all of the observations.>
Taghi M. Khoshgoftaar, John C. Munson, David L. Lanning
ICSM1
1993 A neural network modeling methodology for the detection of high-risk programs
abstract
The profitability of a software development effort is highly dependent on both timely market entry and the reliability of the released product. To get a highly reliable product to the market on schedule, software engineers must allocate resources appropriately across the development effort. Software quality models based upon data drawn from past projects can identify key risk or problem areas in current similar development efforts. Knowing the high-risk modules in a software design is a key to good design and staffing decisions. A number of researchers have recognized this, and have applied modeling technqiues to isolate fault-prone or high-risk program modules early in the development cycle. Discriminant analytic classification models have shown promise in performing this task. We introduce a neural network classification model for identifying high-risk program modules, and we compare the quality of this model with that of a discriminant classification model fitted with the same data. We find that the neural network techniques provide a better management tool in software engineering environments.
Taghi M. Khoshgoftaar, David L. Lanning, Abhijit S. Pandya
ISSRE1
1993 A Performance Analysis of the IBM Subsystem Control Block Architecture in a Video Conferencing Environment
abstract
Article A performance analysis of the IBM subsystem control block architecture in a video conferencing environment Share on Authors: Khoa D. Huynh View Profile , Taghi M. Khoshgoftaar View Profile Authors Info & Claims MULTIMEDIA '93: Proceedings of the first ACM international conference on MultimediaSeptember 1993 Pages 321–330https://doi.org/10.1145/166266.168413Online:01 September 1993Publication History 4citation405DownloadsMetricsTotal Citations4Total Downloads405Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Khoa D. Huynh, Taghi M. Khoshgoftaar
ACM Multimedia2
1993 A high-level performance analysis of the IBM subsystem control block (SCB) architecture
Khoa D. Huynh, Taghi M. Khoshgoftaar, Gerald Marazas
Microprocess. Microprogramming2
1993 Measurement of data structure complexity
John C. Munson, Taghi M. Khoshgoftaar
J. Syst. Softw.2
1992 Software measurement for the space shuttle HAL/S maintenance environment
abstract
The authors discuss the development of a complexity domain model for the HAL/S programming language in which all of the space shuttle software is written. Their research has indicated that some aspects of software complexity in other programming language environments are strongly related to software problems incurred during software maintenance. The goal of the current research effort is to establish the particular relationship between the complexity domain model and measurable aspects of software maintenance. From a maintenance perspective, the specific objective is to identify those software complexity attributes that are most closely related to or highly correlated with measures related to software maintenance. This project is expected to lead to the creation of a large database of space system metrics data, the development of a metrics capture tool, and the identification of leading indicators of software maintenance problems.>
John C. Munson, Taghi M. Khoshgoftaar
ICSM2
1992 A neural network approach for predicting software development faults
abstract
Accurately predicting the number of faults in program modules is a major problem in the quality control of a large scale software system. In this paper, the use of the neural networks as a tool for predicting the number of faults in programs is explored. Software complexity metrics have been shown to be closely related to the distribution of faults in program modules. The objective in the construction of models of software quality is to use measures that may be obtained relatively early in the software development life cycle to provide reasonable initial estimates of quality of an evolving software system. Measures of software quality and software complexity to be used in this modeling process exhibit systematic departures of normality assumptions of regression modeling. This paper introduces a new approach for static reliability modeling and compares its performance in the modeling of software reliability from software complexity in terms of the predictive quality and the quality of fit with more traditional regression modeling techniques. The neural networks did produce models with better quality of fit and predictive quality when applied to one data set obtained from a large commercial system.>
Taghi M. Khoshgoftaar, Abhijit S. Pandya, Hemant B. More
ISSRE1
1992 A workload model for frame-based real-time applications on distributed systems
Khoa D. Huynh, Eduardo B. Fernández, Taghi M. Khoshgoftaar
J. Syst. Softw.3
1992 Predictive Modeling Techniques of Software Quality from Software Measures
abstract
The objective in the construction of models of software quality is to use measures that may be obtained relatively early in the software development life cycle to provide reasonable initial estimates of the quality of an evolving software system. Measures of software quality and software complexity to be used in this modeling process exhibit systematic departures of the normality assumptions of regression modeling. Two new estimation procedures are introduced, and their performances in the modeling of software quality from software complexity in terms of the predictive quality and the quality of fit are compared with those of the more traditional least squares and least absolute value estimation techniques. The two new estimation techniques did produce regression models with better quality of fit and predictive quality when applied to data obtained from two software development projects.>
Taghi M. Khoshgoftaar, John C. Munson, Bibhuti B. Bhattacharya, Gary D. Richardson
IEEE Trans. Software Eng.1
1992 The Detection of Fault-Prone Programs
abstract
The use of the statistical technique of discriminant analysis as a tool for the detection of fault-prone programs is explored. A principal-components procedure was employed to reduce simple multicollinear complexity metrics to uncorrelated measures on orthogonal complexity domains. These uncorrelated measures were then used to classify programs into alternate groups, depending on the metric values of the program. The criterion variable for group determination was a quality measure of faults or changes made to the programs. The discriminant analysis was conducted on two distinct data sets from large commercial systems. The basic discriminant model was constructed from deliberately biased data to magnify differences in metric values between the discriminant groups. The technique was successful in classifying programs with a relatively low error rate. While the use of linear regression models has produced models of limited value, this procedure shows great promise for use in the detection of program modules with potential for faults.>
John C. Munson, Taghi M. Khoshgoftaar
IEEE Trans. Software Eng.2
1991 Software reliability model selection: a cast study
abstract
Predicting the remaining errors in a software system historically has been difficult to do with accuracy. The models used to predict future events have often worked well on one system or collection of data, and not at all well on another. Much of the recent work in the software reliability field has been on model selection and identifying which model would work well with which software system. The Akaike Information Criterion can be used to select the best model from among several models. A case study is given of an application of this technique to an ongoing software project. The Akaike Information Criterion was used to select the best model for a system and then that model was used to predict the number of remaining errors.>
Taghi M. Khoshgoftaar, Timothy G. Woodcock
ISSRE1
1991 The use of software complexity metrics in software reliability modeling
abstract
The central theme of the study is the creation of a suitable complexity measure for use in software reliability models. Factor analytic techniques are employed to reduce the dimensionality of the complexity problem space to produce a set of reduced metrics. The reduced metrics are subsequently combined into a single relative complexity measure. Program complexity varies dynamically as a function of inputs to the system. Hence, the notion of relative complexity is extended to a dynamic or functional complexity metric for use in proposed modifications to existing reliability models.>
John C. Munson, Taghi M. Khoshgoftaar
ISSRE2
1990 The lines of code metric as a predictor of program faults: a critical analysis
abstract
The relationship between measures of software complexity and programming errors is explored. Four distinct regression models were developed for an experimental set of data to create a predictive model from software complexity metrics to program errors. The lines of code metric, traditionally associated with programming errors in predictive models, was found to be less valuable as a criterion measure in these models than measures of software control complexity. A factor analytic technique used to construct a linear compound of lines of code with control metrics was found to yield models of superior predictive quality.>
Taghi M. Khoshgoftaar, John C. Munson
COMPSAC1
1990 Predicting Software Development Errors Using Software Complexity Metrics
abstract
Predictive models that incorporate a functional relationship of program error measures with software complexity metrics and metrics based on factor analysis of empirical data are developed. Specific techniques for assessing regression models are presented for analyzing these models. Within the framework of regression analysis, the authors examine two separate means of exploring the connection between complexity and errors. First, the regression models are formed from the raw complexity metrics. Essentially, these models confirm a known relationship between program lines of code and program errors. The second methodology involves the regression of complexity factor measures and measures of errors. These complexity factors are orthogonal measures of complexity from an underlying complexity domain model. From this more global perspective, it is believed that there is a relationship between program errors and complexity domains of program structure and size (volume). Further, the strength of this relationship suggests that predictive models are indeed possible for the determination of program errors from these orthogonal complexity domains.>
Taghi M. Khoshgoftaar, John C. Munson
IEEE J. Sel. Areas Commun.1
1990 Applications of a relative complexity metric for software project management
John C. Munson, Taghi M. Khoshgoftaar
J. Syst. Softw.2
1989 The Dimensionality of Program Complexity
abstract
Software complexity metrics attempt to define the unique characteristics of computer programs in an analytical way.Many such metrics have been developed to explain various perceived differences among programs.Many studies have been conducted to show the similarity among classes of these metrics.What is lacking in this body of literature is a technique which will aid in the establishment of the true dimensionality of the complexity problem space.The objective of this paper is to examine some recent investigations in the area of software complexity using factor analysis to begin an exploration of the actual dimensionality of the complexity metrics.This technique can expose the relationships of these many metrics, one to another.Some correlation coefficients from recent empirical studies on software metrics were factor analyzed, showing the probable existence of five complexity dimensions within thirty five different complexity measures.
John C. Munson, Taghi M. Khoshgoftaar
ICSE2