EDBT 2026 Demo / reviewers in the wild / expert
Kwabena Ebo Bennin
dblp:166/2603
· DBLP profile ↗
49ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0001-9140-9271ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 42 · 9 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Advancing research software engineering with AI: a research frameworkabstractAbstract The rapid adoption of Artificial Intelligence (AI) and Generative AI (GenAI) tools is transforming the creation, maintenance, and dissemination of research software. Despite their growing prevalence, the implications of these technologies for Research Software Engineering (RSE) practices remain underexplored. This work introduces AI4RSE , an emerging research domain focused on the integration of AI into the development lifecycle of research software. To investigate current trends in AI-augmented RSE, we conducted an empirical study of more than 1,500 open-source research software repositories hosted on Zenodo. Each repository was assessed using a quadrant-based typology defined by two key dimensions: software engineering maturity and the level of AI integration. Our analysis combined static and semantic code inspection, evaluation of alignment with the FAIR Principles for Research Software (FAIR4RS), and heuristic classification of generative AI usage and MLOps adoption. Repositories are categorized into four development modes: Exploratory Coding , Vibe Coding , RSE , and AI4RSE , which reflect different levels of process rigor and AI tool integration. While many projects exhibit informal development patterns, a growing subset demonstrates mature, AI-assisted workflows. This landscape reveals key challenges, such as reproducibility risks and licensing ambiguity, while also highlighting emerging opportunities, including AI-assisted testing and intelligent documentation generation. The findings support a research agenda for AI4RSE, outlining benchmarks, guidelines, and community standards to promote responsible, reproducible, and scalable adoption of AI in scientific software development. Siamak Farshidi, Kwabena Ebo Bennin, Önder Babur, June Sallou, Ayalew Kassahun, Bedir Tekinerdogan |
Autom. Softw. Eng. | 2 |
| 2025 | Energy Efficient Device-to-Device Routing in Three-Tier Symbiotic Radio IoT NetworksabstractThis paper presents a three-tier (3T) symbiotic radio device-to-device (D2D) routing Internet-of-Things network. The primary network (PN) consists of a base station routing data through multiple single antenna decode-and-forward sensor node(s) (DSNs) and to the destination sensor node. The secondary network (SN) consists of single-antenna active sensor tag nodes (ASN) transmitting sensed data to a multi-antenna reader using the backscatter (BC) technique. The tertiary network (TN) consists of single antenna semi-passive energy harvesting sensor tags (SSN) that transmit sensed data using BC to the same multi-antenna reader. The symbiotic nature exists where the ASNs and SSNs utilize the DSNs' radio frequency for data transfer. At the same time, the PN benefits from ASNs and SSNs assistive data transfer in terms of spatial diversity. In this work, we focus on maximizing the EE for the proposed 3T SRN D2D-routing by optimizing the DSN and ASN power allocation and the SSN reflection co-efficient. Based on the optimization problem, a minimum distance-based and a maximum channel gain-based D2D-routing algorithms are proposed. The proposed channel-gain-based 3T SRN D2D-routing algorithm is shown to be superior to the distance-based algorithm. Derek Kwaku Pobi Asiedu, Kwabena Ebo Bennin, Mustapha Benjillali, Samir Saoudi |
WCNC | 2 |
| 2025 | An Energy-Efficient Precision Agriculture Communication Embedded System CoAP DesignabstractPrecision agriculture in rural areas often faces challenges in communication due to the limited availability, robustness and power-constrained devices of existing infrastructure. The Constrained Application Protocol (CoAP), particularly when it is run over the User Datagram Protocol (UDP) transport layer, offers a lightweight and efficient communication solution for resource-constrained environments prevalent in precision agriculture. Therefore, this work focuses on the implementation of CoAP over UDP to bypass modem firmware and rely on the Zephyr operating system, providing a more accessible and scalable communication model for energy-efficient precision agricultural systems. The CoAP over UDP protocol implementation results highlight the practical advantages of CoAP for energy efficiency and reduced latency, particularly in scenarios with limited connectivity or high packet loss conditions. In addition, byte overhead analysis reveals a substantial reduction in overhead with CoAP, especially in small payload scenarios Efraim P. N. Manurung, Mattia Fiumara, Kwabena Ebo Bennin, Derek Kwaku Pobi Asiedu |
IEEE Internet Things J. | 3 |
| 2024 | Agile Requirements Engineering in a Distributed Environment: Experiences from the Software Industry During Unprecedented Global ChallengesabstractUnprecedented global challenges such as the COVID-19 pandemic necessitated a widespread transition to Work-From-Home (WFH) arrangements for project teams, posing significant challenges in conveying requirements within agile Requirements Engineering (RE). While numerous studies have examined the impact of transitioning work routines during the pandemic, limited research exists on the specific challenges of agile RE operating within the WFH context. Given the pervasive shift in the software development ecosystem worldwide, where WFH is projected to persist even in the post-COVID era, it is imperative to ascertain the challenges associated with WFH-based agile RE. During the pandemic, we collaborated with startups to conduct an industry-academia project. By adopting the methodology of action research, this study comprehensively analyzed agile RE practices and reported the key challenges encountered within the WFH context. To mitigate these challenges, several collaborative RE techniques were employed in three intervention cycles. Interviews were conducted to thoroughly analyze the results. This study also provides insights into collaborative RE techniques and valuable lessons learned. Considering the increasing prevalence of WFH as a working mode in the post-pandemic era, this study equips the community with practical strategies to navigate agile RE challenges and better prepare for unprecedented challenges in the future. Yishu Li, Jacky W. Keung, Kwabena Ebo Bennin, Zhen Yang 0022 |
COMPSAC | 3 |
| 2024 | An Empirical Study of the Impact of Test Strategies on Online Optimization for Ensemble-Learning Defect PredictionabstractEnsemble learning methods have been used to enhance the reliability of defect prediction models. However, there is an inconclusive stability of a single method attaining the highest accuracy among various software projects. This work aims to improve the performance of ensemble-learning defect prediction among such projects by helping select the highest accuracy ensemble methods. We employ bandit algorithms (BA), an online optimization method, to select the highest-accuracy ensemble method. Each software module is tested sequentially, and bandit algorithms utilize the test outcomes of the modules to evaluate the performance of the ensemble learning methods. The test strategy followed might impact the testing effort and prediction accuracy when applying online optimization. Hence, we analyzed the test order's influence on BA's performance. In our experiment, we used six popular defect prediction datasets, four ensemble learning methods such as bagging, and three test strategies such as testing positive-prediction modules first (PF). Our results show that when BA is applied with PF, the prediction accuracy improved on average, and the number of found defects increased by 7% on a minimum of five out of six datasets (although with a slight increase in the testing effort by about 4% from ordinal ensemble learning). Hence, BA with PF strategy is the most effective to attain the highest prediction accuracy using ensemble methods on various projects. Kensei Hamamoto, Masateru Tsunoda, Amjed Tahir, Kwabena Ebo Bennin, Akito Monden, Koji Toda, Keitaro Nakasai, Ken-ichi Matsumoto |
ICSME | 4 |
| 2024 | A systematic review on food recommender systemsabstractThe Internet has revolutionized the way information is retrieved, and the increase in the number of users has resulted in a surge in the volume and heterogeneity of available data. Recommender systems have become popular tools to help users retrieve relevant information quickly. Food Recommender Systems (FRS), in particular, have proven useful in overcoming the overload of information present in the food domain. However, the recommendation of food is a complex domain with specific characteristics causing many challenges. Additionally, very few systematic literature reviews have been conducted in the domain on FRS. This paper presents a systematic literature review that summarizes the current state-of-the-art in FRS. Our systematic review examines the different methods and algorithms used for recommendation, the data and how it is processed, and evaluation methods. It also presents the advantages and disadvantages of FRS. To achieve this, a total of 67 high-quality studies were selected from a pool of 2,738 studies using strict quality criteria. The review provides valuable information to the research field, helping researchers in the domain to select a strategy to develop FRS. This review can help improve the efficiency of development, thus closing the gap between the development of FRS and other recommender systems. Jon Nicolas Bondevik, Kwabena Ebo Bennin, Önder Babur, Carsten Ersch |
Expert Syst. Appl. | 2 |
| 2023 | Towards Requirements Engineering Activities for Machine Learning-Enabled FinTech ApplicationsabstractThe complexity required in the software development of machine learning (ML) applications introduces additional challenges to requirement engineering (RE) activities. RE researchers expressed concerns and the need for more discussions on RE for ML, requiring additional real-world case studies to evaluate RE activities for practical ML-enabled applications. This study aims to observe the RE activities for ML-enabled systems in a real-world context, taking action research in the ML-enabled FinTech project where the RE activities are being adjusted by engaging the data scientists to help and clarify ML-related requirements. This paper discussed the difficulties of RE activities from the perspectives of the data scientist and requirement engineer. Considering data and model relevance in developing the ML-enabled FinTech application, a RE framework iteratively made active changes according to the parameters is proposed, which includes the selected ML-related requirement characteristics to pursue and complete RE activities for ML-enabled application development. The feedback from the practitioners indicates that such practices address the difficulties of improving data quality and verifying model requirements in RE activities. The lessons learned by researchers and practitioners are also presented, which provides practical suggestions to the SE and RE communities with similar concerns in the related context. Yishu Li, Jacky W. Keung, Kwabena Ebo Bennin, Yangyang Huang |
APSEC | 3 |
| 2023 | Reference architecture design for computer-based speech therapy systemsabstractWith the current international shortage of speech-language pathologists (SLPs), there is a demand for online tools to support SLPs with their daily tasks. For this purpose, several online speech therapy systems (OSTSs) have been proposed and discussed in the literature. However, developing these OSTSs is not trivial since it involves the consideration of various functional and quality concerns. Hence, for communicating the design decisions and guiding the development and analysis of these systems, a proper architecture design is important. Unfortunately, the architecture design of OSTSs has not been explicitly addressed in the literature. To this end, we present a reference architecture for OSTSs which has been designed following well-established architecture design methods. The reference architecture captures the reusable design elements of OSTSs and can be used to derive various different application architectures. A case study approach is used to illustrate and validate the use of the presented reference architecture. Geertruida Aline Attwell, Kwabena Ebo Bennin, Bedir Tekinerdogan |
Comput. Speech Lang. | 2 |
| 2023 | Developers talking about code qualityabstractAbstract There are many aspects of code quality, some of which are difficult to capture or to measure. Despite the importance of software quality, there is a lack of commonly accepted measures or indicators for code quality that can be linked to quality attributes. We investigate software developers’ perceptions of source code quality and the practices they recommend to achieve these qualities. We analyze data from semi-structured interviews with 34 professional software developers, programming teachers and students from Europe and the U.S. For the interviews, participants were asked to bring code examples to exemplify what they consider good and bad code, respectively. Readability and structure were used most commonly as defining properties for quality code. Together with documentation, they were also suggested as the most common target properties for quality improvement. When discussing actual code, developers focused on structure, comprehensibility and readability as quality properties. When analyzing relationships between properties, the most commonly talked about target property was comprehensibility. Documentation, structure and readability were named most frequently as source properties to achieve good comprehensibility. Some of the most important source code properties contributing to code quality as perceived by developers lack clear definitions and are difficult to capture. More research is therefore necessary to measure the structure, comprehensibility and readability of code in ways that matter for developers and to relate these measures of code structure, comprehensibility and readability to common software quality attributes. Jürgen Börstler, Kwabena Ebo Bennin, Sara Hooshangi, Johan Jeuring, Hieke Keuning, Carsten Kleiner, Bonnie K. MacKellar, Rodrigo Duran 0001, Harald Störrle, Daniel Toll, Jelle van Assema |
Empir. Softw. Eng. | 2 |
| 2023 | Finding the best learning to rank algorithms for effort-aware defect predictionabstractContext: Effort-Aware Defect Prediction (EADP) ranks software modules or changes based on their predicted number of defects (i.e., considering modules or changes as effort) or defect density (i.e., considering LOC as effort) by using learning to rank algorithms . Ranking instability refers to the inconsistent conclusions produced by existing empirical studies of EADP. The major reason is the poor experimental design , such as comparison of few learning to rank algorithms, the use of small number of datasets or datasets without indicating numbers of defects, and evaluation with inappropriate or few metrics. Objective: To find a stable ranking of learning to rank algorithms to investigate the best ones for EADP, Method: We examine the practical effects of 34 algorithms on 49 datasets for EADP. We measure the performance of these algorithms using 7 module-based and 7 LOC-based metrics and run experiments under cross-release and cross-project settings, respectively. Finally, we obtain the ranking of these algorithms by performing the Scott-Knott ESD test. Results: When module is used as effort, random forest regression performs the best under cross-release setting, and linear regression performs the best under cross-project setting among the learning to rank algorithms; (2) when LOC is used as effort, LTR-linear (Learning-to-Rank with the linear model) performs the best under cross-release setting, and Ranking SVM performs the best under cross-project setting. Conclusion: This comprehensive experimental procedure allows us to discover a stable ranking of the studied algorithms to select the best ones according to the requirement of software projects. Xiao Yu 0008, Heng Dai, Li Li 0029, Xiaodong Gu 0002, Jacky W. Keung, Kwabena Ebo Bennin, Jin Liu 0016 |
Inf. Softw. Technol. | 6 |
| 2023 | On the use of deep learning in software defect predictionabstractAutomated software defect prediction (SDP) methods are increasingly applied, often with the use of machine learning (ML) techniques. Yet, the existing ML-based approaches require manually extracted features, which are cumbersome, time consuming and hardly capture the semantic information reported in bug reporting tools. Deep learning (DL) techniques provide practitioners with the opportunities to automatically extract and learn from more complex and high-dimensional data. The purpose of this study is to systematically identify, analyze, summarize, and synthesize the current state of the utilization of DL algorithms for SDP in the literature. We systematically selected a pool of 102 peer-reviewed studies and then conducted a quantitative and qualitative analysis using the data extracted from these studies. Main highlights include: (1) most studies applied supervised DL; (2) two third of the studies used metrics as an input to DL algorithms; (3) Convolutional Neural Network is the most frequently used DL algorithm. Based on our findings, we propose to (1) develop more comprehensive DL approaches that automatically capture the needed features; (2) use diverse software artifacts other than source code; (3) adopt data augmentation techniques to tackle the class imbalance problem; (4) publish replication packages. Görkem Giray, Kwabena Ebo Bennin, Ömer Köksal, Önder Babur, Bedir Tekinerdogan |
J. Syst. Softw. | 2 |
| 2022 | Preliminary Analysis of Review Method Selection Based on Bandit AlgorithmsabstractTo enhance the reliability of software, it is important is to review all software artifacts (e.g., design documents) to remove defects as earlier as possible. There are various review methods available, and project managers face the challenge of choosing a suitable method for their current projects. One of approaches to support the selection of review methods is to evaluate review methods beforehand, to identify the most effective method on average. However, past studies have not evaluated review methods thoroughly as the process can be time-consuming. We propose a bandit-algorithm (BA) based method to evaluate and then dynamically select a suitable review method (from a list of candidates). In our experiments, we assume that the proposed method is applied to design document review on basic design phase. We performed experiments based on a simulation, instead of using an actual dataset. On our simulation, when a review method is selected by our BA method, productivity (i.e., total development time) was improved by about 1.25 times, and it was the second highest among candidates of review methods. Takuto Kudo, Masateru Tsunoda, Amjed Tahir, Kwabena Ebo Bennin, Koji Toda, Keitaro Nakasai, Akito Monden, Ken-ichi Matsumoto |
APSEC | 4 |
| 2022 | Using Bandit Algorithms for Selecting Feature Reduction Techniques in Software Defect PredictionabstractBackground: Selecting a suitable feature reduction technique. when building a defect prediction model, can be challenging. Different techniques can result in the selection of different independent variables which have an impact on the overall performance of the prediction model. To help in the selection, previous studies have assessed the impact of each feature reduction technique using different datasets. However, there are many reduction techniques, and therefore some of the well-known techniques have not been assessed by those studies. Aim: The goal of the study is to select a high-accuracy reduction technique from several candidates without preliminary assessments. Method: We utilized bandit algorithm (BA) to help with the selection of best features reduction technique for a list of candidates. To select the best feature reduction technique, BA evaluates the prediction accuracy of the candidates, comparing testing results of different modules with their prediction results. By substituting the reduction technique for the prediction method, BA can then be used to select the best reduction technique. In the experiment, we evaluated the performance of BA to select suitable reduction technique. We performed cross version defect prediction using 14 datasets. As feature reduction techniques, we used two assessed and two non-assessed techniques. Results: Using BA, the prediction accuracy was higher or equivalent than existing approaches on average, compared with techniques selected based on an assessment. Conclusions: BA can have larger impact on improving prediction models by helping not only on selecting suitable models, but also in selecting suitable feature reduction techniques. Masateru Tsunoda, Akito Monden, Koji Toda, Amjed Tahir, Kwabena Ebo Bennin, Keitaro Nakasai, Masataka Nagura, Ken-ichi Matsumoto |
MSR | 5 |
| 2022 | An empirical study on the effectiveness of data resampling approaches for cross-project software defect predictionabstractAbstract Cross‐project defect prediction (CPDP), where data from different software projects are used to predict defects, has been proposed as a way to provide data for software projects that lack historical data. Evaluations of CPDP models using the Nearest Neighbour (NN) Filter approach have shown promising results in recent studies. A key challenge with defect‐prediction datasets is class imbalance, that is, highly skewed datasets where non‐buggy modules dominate the buggy modules. In the past, data resampling approaches have been applied to within‐projects defect prediction models to help alleviate the negative effects of class imbalance in the datasets. To address the class imbalance issue in CPDP, the authors assess the impact of data resampling approaches on CPDP models after the NN Filter is applied. The impact on prediction performance of five oversampling approaches (MAHAKIL, SMOTE, Borderline‐SMOTE, Random Oversampling and ADASYN) and three undersampling approaches (Random Undersampling, Tomek Links and One‐sided selection) is investigated and results are compared to approaches without data resampling. The authors examined six defect prediction models on 34 datasets extracted from the PROMISE repository. The authors' results show that there is a significant positive effect of data resampling on CPDP performance, suggesting that software quality teams and researchers should consider applying data resampling approaches for improved recall ( pd ) and g ‐measure prediction performance. However, if the goal is to improve precision and reduce false alarm ( pf ) then data resampling approaches should be avoided. Kwabena Ebo Bennin, Amjed Tahir, Stephen G. MacDonell, Jürgen Börstler |
IET Softw. | 1 |
| 2022 | Further investigation of the survivability of code technical debt itemsabstractAbstract Context: Technical debt (TD) discusses the negative impact of sub‐optimal decisions to cope with the need‐for‐speed in software development. Code technical debt items (TDI) are atomic elements of TD that can be observed in code artifacts. Empirical results on open‐source systems demonstrated how code‐smells, which are just one type of TDIs, are introduced and “survive” during release cycles. However, little is known about whether the results on the survivability of code‐smells hold for other types of code TDIs (i.e., bugs and vulnerabilities) and in industrial settings. Goal: Understanding the survivability of code TDIs by conducting an empirical study analyzing two industrial cases and 31 open‐source systems from Apache Foundation. Method: We analyzed 133,670 code TDIs (35,703 from the industrial systems) detected by SonarQube (in 193,196 commits) to assess their survivability using survivability models. Results: In general, code TDIs tend to remain and linger for long periods in open‐source systems, whereas they are removed faster in industrial systems. Code TDIs that survive over a certain threshold tend to remain much longer, which confirms previous results. Our results also suggest that bugs tend to be removed faster, while code smells and vulnerabilities tend to survive longer. Ehsan Zabardast, Kwabena Ebo Bennin, Javier Gonzalez-Huerta |
J. Softw. Evol. Process. | 2 |
| 2021 | Using Bandit Algorithms for Project Selection in Cross-Project Defect PredictionabstractBackground: defect prediction model is built using historical data from previous versions/releases of the same project. However, such historical data may not exist in case of newly developed projects. Alternatively, one can train a model using data obtained from external projects. This approach is known as cross-project defect prediction (CPDP). In CPDP, it is still difficult to utilize external projects' data or decide which particular project to use to train a model. Aim: to address this issue, we apply bandit algorithm (BA) to CPDP in order to select the most suitable training project from a set of projects. Method: BA-based prediction iteratively reselects the project after each module is tested, considering the accuracy of the predictions. As baselines, we used simple CPDP methods such as training a model with randomly selected project. All models were built using logistic regression. Results: We experimented our approach on two datasets (NASA and DAMB, with a total of 12 projects). The BA-based defect prediction models resulted in, on average, a higher accuracy (AUC and F1 score) than the baselines. Conclusion: in this preliminarily study, we demonstrate the feasibility of using BA in the context of CPDP. Our initial assessment shows that the use BA for predicting defects in CPDP is promising and may outperform existing approaches. Takuya Asano, Masateru Tsunoda, Koji Toda, Amjed Tahir, Kwabena Ebo Bennin, Keitaro Nakasai, Akito Monden, Ken-ichi Matsumoto |
ICSME | 5 |
| 2021 | Does class size matter? An in-depth assessment of the effect of class size in software defect prediction
Amjed Tahir, Kwabena Ebo Bennin, Xun Xiao, Stephen G. MacDonell |
Empir. Softw. Eng. | 2 |
| 2021 | COSTE: Complexity-based OverSampling TEchnique to alleviate the class imbalance problem in software defect prediction
Shuo Feng 0003, Jacky W. Keung, Xiao Yu 0008, Yan Xiao 0002, Kwabena Ebo Bennin, Md. Alamgir Kabir, Miao Zhang 0025 |
Inf. Softw. Technol. | 5 |
| 2020 | A Drift Propensity Detection Technique to Improve the Performance for Cross-Version Software Defect PredictionabstractIn cross-version defect prediction (CVDP), historical data is derived from the prior version of the same project to predict defects of the current version. Recent studies in CVDP focus on subset selection to deal with the changes of the data distributions. No prior study has focused on training data arriving in streaming fashion across the versions where the significant differences between versions make the prediction unreliable. We refer to this situation as Drift Propensity (DP). By identifying DP, necessary steps can be taken (e.g., updating or retraining the model) to improve the prediction performance. In this paper, we investigate the chronological defect datasets and identify DP in the datasets. The no-memory data management technique is employed to manage the data distributions and a DP detection technique is proposed. The idea behind the proposed DP detection technique is to monitor the algorithm's error-rate. The DP detector triggers DP, warning, and control flags to take necessary steps. The proposed technique is significantly superior in identifying the distribution differences (p-value <; 0.05). The DP's identified in the data distributions achieve large effect sizes (Hedges' g ≥ 0.80) during the pair-wise comparisons. We observe that if the error-rate exponentially increases, it causes DP, resulting in prediction performance deterioration. We thus recommend researches and practitioners to address DP in the chronological datasets. Due to its potential effects in the datasets, the prediction models could be enhanced to get the best results in CVDP. Md. Alamgir Kabir, Jacky W. Keung, Kwabena Ebo Bennin, Miao Zhang 0025 |
COMPSAC | 3 |
| 2020 | Revisiting the Impact of Concept Drift on Just-in-Time Quality AssuranceabstractThe performance of software defect prediction(SDP) models is known to be dependent on the datasets used for training the models. Evolving data in a dynamic software development environment such as significant refactoring and organizational changes introduces new concept to the prediction model, thus making improved classification performance difficult. In this study, we investigate and assess the existence and impact of concept drift on SDP performances. We empirically asses the prediction performance of five models by conducting cross-version experiments using fifty-five releases of five open-source projects. Prediction performance fluctuated as the training datasets changed over time. Our results indicate that the quality and the reliability of defect prediction models fluctuate over time and that this instability should be considered by software quality teams when using historical datasets. The performance of a static predictor constructed with data from historical versions may degrade over time due to the challenges posed by concept drift. Kwabena Ebo Bennin, Nauman Bin Ali, Jürgen Börstler, Xiao Yu 0008 |
QRS | 1 |
| 2020 | Improving Ranking-Oriented Defect Prediction Using a Cost-Sensitive Ranking SVMabstractContext: Ranking-oriented defect prediction (RODP) ranks software modules to allocate limited testing resources to each module according to the predicted number of defects. Most RODP methods overlook that ranking a module with more defects incorrectly makes it difficult to successfully find all of the defects in the module due to fewer testing resources being allocated to the module, which results in much higher costs than incorrectly ranking the modules with fewer defects, and the numbers of defects in software modules are highly imbalanced in defective software datasets. Cost-sensitive learning is an effective technique in handling the cost issue and data imbalance problem for software defect prediction. However, the effectiveness of cost-sensitive learning has not been investigated in RODP models. Aims: In this article, we propose a cost-sensitive ranking support vector machine (SVM) (CSRankSVM) algorithm to improve the performance of RODP models. Method: CSRankSVM modifies the loss function of the ranking SVM algorithm by adding two penalty parameters to address both the cost issue and the data imbalance problem. Additionally, the loss function of the CSRankSVM is optimized using a genetic algorithm. Results: The experimental results for 11 project datasets with 41 releases show that CSRankSVM achieves 1.12%-15.68% higher average fault percentile average (FPA) values than the five existing RODP methods (i.e., decision tree regression, linear regression, Bayesian ridge regression, ranking SVM, and learning-to-rank (LTR)) and 1.08%-15.74% higher average FPA values than the four data imbalance learning methods (i.e., random undersampling and a synthetic minority oversampling technique; two data resampling methods; RankBoost, an ensemble learning method; IRSVM, a CSRankSVM method for information retrieval). Conclusion: CSRankSVM is capable of handling the cost issue and data imbalance problem in RODP methods and achieves better performance. Therefore, CSRankSVM is recommended as an effective method for RODP. Xiao Yu 0008, Jin Liu 0016, Jacky W. Keung, Qing Li 0001, Kwabena Ebo Bennin, Zhou Xu 0003, Xiaohui Cui |
IEEE Trans. Reliab. | 5 |
| 2019 | Assessing the Significant Impact of Concept Drift in Software Defect PredictionabstractConcept drift is a known phenomenon in software data analytics. It refers to the changes in the data distribution over time. The performance of analytic and prediction models degrades due to the changes in the data over time. To improve prediction performance, most studies propose that the prediction model be updated when concept drift occurs. In this work, we investigate the existence of concept drift and its associated effects on software defect prediction performance. We adopt the strategy of an empirically proven method DDM (Drift Detection Method) and evaluate its statistical significance using the chi-square test with Yates continuity correction. The objective is to empirically determine the concept drift and to calibrate the base model accordingly. The empirical study indicates that the concept drift occurs in software defect datasets, and its existence subsequently degrades the performance of prediction models. Two types of concept drifts (gradual and sudden drifts) were identified using the chi-square test with Yates continuity correction in the software defect datasets studied. We suggest concept drift should be considered by software quality assurance teams when building prediction models. Md. Alamgir Kabir, Jacky W. Keung, Kwabena Ebo Bennin, Miao Zhang 0025 |
COMPSAC (1) | 3 |
| 2019 | An Empirical Study of Learning to Rank Techniques for Effort-Aware Defect PredictionabstractEffort-Aware Defect Prediction (EADP) ranks software modules based on the possibility of these modules being defective, their predicted number of defects, or defect density by using learning to rank algorithms. Prior empirical studies compared a few learning to rank algorithms considering small number of datasets, evaluating with inappropriate or one type of performance measure, and non-robust statistical test techniques. To address these concerns and investigate the impact of learning to rank algorithms on the performance of EADP models, we examine the practical effects of 23 learning to rank algorithms on 41 available defect datasets from the PROMISE repository using a module-based effort-aware performance measure (FPA) and a source lines of code (SLOC) based effort-aware performance measure (Norm(Popt). In addition, we compare the performance of these algorithms when they are trained on a more relevant feature subset selected by the Information Gain feature selection method. In terms of FPA and Norm(Popt), statistically significant differences are observed among these algorithms with BRR (Bayesian Ridge Regression) performing best in terms of FPA, and BRR and LTR (Learning-to-Rank) performing best in terms of Norm (Popt). When these algorithms are trained on a more relevant feature subset selected by Information Gain, LTR and BRR still perform best with significant differences in terms of FPA and Norm(Popt). Therefore, we recommend BRR and LTR for building the EADP model in order to find more defects by inspecting a certain number of modules or lines of codes. Xiao Yu 0008, Kwabena Ebo Bennin, Jin Liu 0016, Jacky W. Keung, Xiaofei Yin, Zhou Xu 0003 |
SANER | 2 |
| 2019 | On the relative value of data resampling approaches for software defect prediction
Kwabena Ebo Bennin, Jacky W. Keung, Akito Monden |
Empir. Softw. Eng. | 1 |
| 2019 | Improving bug localization with word embedding and enhanced convolutional neural networks
Yan Xiao 0002, Jacky W. Keung, Kwabena Ebo Bennin, Qing Mi |
Inf. Softw. Technol. | 3 |
| 2018 | Bug Localization with Semantic and Structural Features using Convolutional Neural Network and Cascade ForestabstractBackground: Correctly localizing buggy files for bug reports together with their semantic and structural information is a crucial task, which would essentially improve the accuracy of bug localization techniques. Aims: To empirically evaluate and demonstrate the effects of both semantic and structural information in bug reports and source files on improving the performance of bug localization, we propose CNN_Forest involving convolutional neural network and ensemble of random forests that have excellent performance in the tasks of semantic parsing and structural information extraction. Method: We first employ convolutional neural network with multiple filters and an ensemble of random forests with multi-grained scanning to extract semantic and structural features from the word vectors derived from bug reports and source files. And a subsequent cascade forest (a cascade of ensembles of random forests) is used to further extract deeper features and observe the correlated relationships between bug reports and source files. CNNLForest is then empirically evaluated over 10,754 bug reports extracted from AspectJ, Eclipse UI, JDT, SWT, and Tomcat projects. Results: The experiments empirically demonstrate the significance of including semantic and structural information in bug localization, and further show that the proposed CNN_Forest achieves higher Mean Average Precision and Mean Reciprocal Rank measures than the best results of the four current state-of-the-art approaches (NPCNN, LR+WE, DNNLOC, and BugLocator). Conclusion: CNNLForest is capable of defining the correlated relationships between bug reports and source files, and we empirically show that semantic and structural information in bug reports and source files are crucial in improving bug localization. Yan Xiao 0002, Jacky W. Keung, Qing Mi, Kwabena Ebo Bennin |
EASE | 4 |
| 2018 | Revisiting the size effect in software fault prediction modelsabstractBACKGROUND: In object oriented (OO) software systems, class size has been acknowledged as having an indirect effect on the relationship between certain artifact characteristics, captured via metrics, and fault-proneness, and therefore it is recommended to control for size when designing fault prediction models. Amjed Tahir, Kwabena Ebo Bennin, Stephen G. MacDonell, Stephen R. Marsland |
ESEM | 2 |
| 2018 | MAHAKIL: diversity based oversampling approach to alleviate the class imbalance issue in software defect predictionabstractThis study presents MAHAKIL, a novel and efficient synthetic over-sampling approach for software defect datasets that is based on the chromosomal theory of inheritance. Exploiting this theory, MAHAKIL interprets two distinct sub-classes as parents and generates a new instance that inherits different traits from each parent and contributes to the diversity within the data distribution. We extensively compare MAHAKIL with five other sampling approaches using 20 releases of defect datasets from the PROMISE repository and five prediction models. Our experiments indicate that MAHAKIL improves the prediction performance for all the models and achieves better and more significant pf values than the other oversampling approaches, based on robust statistical tests. Kwabena Ebo Bennin, Jacky W. Keung, Passakorn Phannachitta, Akito Monden, Solomon Mensah |
ICSE | 1 |
| 2018 | Revisiting the Conclusion Instability Issue in Software Effort Estimation (S)abstractConclusion instability is the absence of observing the same effect under varying experimental conditions.Deep Neural Network (DNN) and ElasticNet software effort estimation (SEE) models were applied to two SEE datasets with the view of resolving the conclusion instability issue and assessing the suitability of ElasticNet as a viable SEE benchmark model.Results were mixed as both model types attain conclusion stability for the Kitchenham dataset whilst conclusion instability existed in the Desharnais dataset.ElasticNet was outperformed by DNN and as such it is not recommended to be used as a SEE benchmark model. Michael Franklin Bosu, Solomon Mensah, Kwabena Ebo Bennin, Diab Abuaiadah |
SEKE | 3 |
| 2018 | Cross project defect prediction using class distribution estimation and oversampling
Nachai Limsettho, Kwabena Ebo Bennin, Jacky W. Keung, Hideaki Hata, Ken-ichi Matsumoto |
Inf. Softw. Technol. | 2 |
| 2018 | Duplex output software effort estimation model with self-guided interpretation
Solomon Mensah, Jacky W. Keung, Michael Franklin Bosu, Kwabena Ebo Bennin |
Inf. Softw. Technol. | 4 |
| 2018 | Machine translation-based bug localization technique for bridging lexical gap
Yan Xiao 0002, Jacky W. Keung, Kwabena Ebo Bennin, Qing Mi |
Inf. Softw. Technol. | 3 |
| 2018 | On the value of a prioritization scheme for resolving Self-admitted technical debt
Solomon Mensah, Jacky W. Keung, Jeffrey Svajlenko, Kwabena Ebo Bennin, Qing Mi |
J. Syst. Softw. | 4 |
| 2018 | Cross-company defect prediction via semi-supervised clustering-based data filtering and MSTrA-based transfer learning
Xiao Yu 0008, Man Wu, Yiheng Jian, Kwabena Ebo Bennin, Mandi Fu, Chuanxiang Ma |
Soft Comput. | 4 |
| 2018 | Investigating the Significance of the Bellwether Effect to Improve Software Effort Prediction: Further Empirical StudyabstractContext: In addressing how best to estimate how much effort is required to develop software, a recent study found that using exemplary and recently completed projects [forming Bellwether moving windows (BMW)] in software effort prediction (SEP) models leads to relatively improved accuracy. More studies need to be conducted to determine whether the BMW yields improved accuracy in general, since different sizing and aging parameters of the BMW are known to affect accuracy. Objective: To investigate the existence of exemplary projects (Bellwethers) with defined window size and age parameters, and whether their use in SEP improves prediction accuracy. Method: We empirically investigate the moving window assumption based on the theory that the prediction outcome of a future event depends on the outcomes of prior events. Sampling of Bellwethers was undertaken using three introduced Bellwether methods (SSPM, SysSam, and RandSam). The ergodic Markov chain was used to determine the stationarity of the Bellwethers. Results: Empirical results show that 1) Bellwethers exist in SEP and 2) the BMW has an approximate size of 50 to 80 exemplary projects that should not be more than 2 years old relative to the new projects to be estimated. Conclusion: The study's results add further weight to the recommended use of Bellwethers for improved prediction accuracy in SEP. Solomon Mensah, Jacky W. Keung, Stephen G. MacDonell, Michael Franklin Bosu, Kwabena Ebo Bennin |
IEEE Trans. Reliab. | 5 |
| 2018 | MAHAKIL: Diversity Based Oversampling Approach to Alleviate the Class Imbalance Issue in Software Defect PredictionabstractHighly imbalanced data typically make accurate predictions difficult. Unfortunately, software defect datasets tend to have fewer defective modules than non-defective modules. Synthetic oversampling approaches address this concern by creating new minority defective modules to balance the class distribution before a model is trained. Notwithstanding the successes achieved by these approaches, they mostly result in over-generalization (high rates of false alarms) and generate near-duplicated data instances (less diverse data). In this study, we introduce MAHAKIL, a novel and efficient synthetic oversampling approach for software defect datasets that is based on the chromosomal theory of inheritance. Exploiting this theory, MAHAKIL interprets two distinct sub-classes as parents and generates a new instance that inherits different traits from each parent and contributes to the diversity within the data distribution. We extensively compare MAHAKIL with SMOTE, Borderline-SMOTE, ADASYN, Random Oversampling and the No sampling approach using 20 releases of defect datasets from the PROMISE repository and five prediction models. Our experiments indicate that MAHAKIL improves the prediction performance for all the models and achieves better and more significant pf values than the other oversampling approaches, based on Brunner's statistical significance test and Cliff's effect sizes. Therefore, MAHAKIL is strongly recommended as an efficient alternative for defect prediction models built on highly imbalanced datasets. Kwabena Ebo Bennin, Jacky W. Keung, Passakorn Phannachitta, Akito Monden, Solomon Mensah |
IEEE Trans. Software Eng. | 1 |
| 2017 | Impact of the Distribution Parameter of Data Sampling Approaches on Software Defect Prediction ModelsabstractSampling methods are known to impact defect prediction performance. These sampling methods have configurable parameters that can significantly affect the prediction performance. It is however, impractical to assess the effect of all the possible different settings in the parameter space for all the several existing sampling methods. A constant and easy to tweak parameter present in all sampling methods is the distribution of the defective and non-defective modules in the dataset known as Pfp (% of fault-prone modules). In this paper, we investigate and assess the performance of defect prediction models where the Pfp parameter of sampling methods are tweaked. An empirical experiment and assessment of seven sampling methods on five prediction models over 20 releases of 10 static metric projects indicate that (1) Area Under the Receiver Operating Characteristics Curve (AUC) performance is not improved after tweaking the Pfp parameter, (2) pf (false alarms) performance degrades as the Pfp is increased. (3) a stable predictor is difficult to achieve across different Pfp rates. Hence, we conclude that the Pfp parameter setting can have a large impact on the performance (except AUC) of defect prediction models. We thus recommend researchers experiment with the Pfp parameter of the sampling method since the distribution of training datasets vary. Kwabena Ebo Bennin, Jacky W. Keung, Akito Monden |
APSEC | 1 |
| 2017 | Correlation between the Frequent Use of Gang-of-Four Design Patterns and Structural ComplexityabstractThe structural complexity of design components (e.g. Classes) is proportional to design quality at the system level and is quantified via the object-oriented metrics. The frequent use of design patterns causes of too much abstraction and can increase the structural complexity of design components. Though, in our previous work, we have empirically investigated the impact of use intensity of design pattern on the system level quality attributes. However, the empirical investigation of the effect of usage of design patterns on the design properties is still required. In this regard, we conduct an empirical study and perform a case study which includes the analysis 1) the existence of a correlation between design pattern usage and design metrics, 2) the confounding effect of system size (number of classes) on the correlation, and 3) how the change in number of employed design pattern instances affects the structural complexity in the subsequent releases of a system. The result of this study suggests that structural complexity associated with aggregation, coupling, functional abstraction design properties has a significant relationship with the employed instances of Template, Adapter-Command, Singleton, and Factory Method design patterns. Shahid Hussain 0001, Jacky W. Keung, Arif Ali Khan, Kwabena Ebo Bennin |
APSEC | 4 |
| 2017 | Improving Bug Localization with an Enhanced Convolutional Neural NetworkabstractBackground: Localizing buggy files automatically speeds up the process of bug fixing so as to improve the efficiency and productivity of software quality teams. There are other useful semantic information available in bug reports and source code, but are mostly underutilized by existing bug localization approaches. Aims: We propose DeepLocator, a novel deep learning based model to improve the performance of bug localization by making full use of semantic information. Method: DeepLocator is composed of an enhanced CNN (Convolutional Neural Network) proposed in this study considering bug-fixing experience, together with a new rTF-IDuF method and pretrained word2vec technique. DeepLocator is then evaluated on over 18,500 bug reports extracted from AspectJ, Eclipse, JDT, SWT and Tomcat projects. Results: The experimental results show that DeepLocator achieves 9.77% to 26.65% higher Fmeasure than the conventional CNN and 3.8% higher MAP than a state-of-the-art method HyLoc using less computation time. Conclusion: DeepLocator is capable of automatically connecting bug reports to the corresponding buggy files and successfully achieves better performance based on a deep understanding of semantics in bug reports and source code. Yan Xiao 0002, Jacky W. Keung, Qing Mi, Kwabena Ebo Bennin |
APSEC | 4 |
| 2017 | The Significant Effects of Data Sampling Approaches on Software Defect Prioritization and ClassificationabstractContext: Recent studies have shown that performance of defect prediction models can be affected when data sampling approaches are applied to imbalanced training data for building defect prediction models. However, the magnitude (degree and power) of the effect of these sampling methods on the classification and prioritization performances of defect prediction models is still unknown. Goal: To investigate the statistical and practical significance of using resampled data for constructing defect prediction models. Method: We examine the practical effects of six data sampling methods on performances of five defect prediction models. The prediction performances of the models trained on default datasets (no sampling method) are compared with that of the models trained on resampled datasets (application of sampling methods). To decide whether the performance changes are significant or not, robust statistical tests are performed and effect sizes computed. Twenty releases of ten open source projects extracted from the PROMISE repository are considered and evaluated using the AUC, pd, pf and G-mean performance measures. Results: There are statistical significant differences and practical effects on the classification performance (pd, pf and G-mean) between models trained on resampled datasets and those trained on the default datasets. However, sampling methods have no statistical and practical effects on defect prioritization performance (AUC) with small or no effect values obtained from the models trained on the resampled datasets. Conclusions: Existing sampling methods can properly set the threshold between buggy and clean samples, while they cannot improve the prediction of defect-proneness itself. Sampling methods are highly recommended for defect classification purposes when all faulty modules are to be considered for testing. Kwabena Ebo Bennin, Jacky W. Keung, Akito Monden, Passakorn Phannachitta, Solomon Mensah |
ESEM | 1 |
| 2017 | Investigating the Significance of Bellwether Effect to Improve Software Effort EstimationabstractBellwether effect refers to the existence of exemplary projects (called the Bellwether) within a historical dataset to be used for improved prediction performance. Recent studies have shown an implicit assumption of using recently completed projects (referred to as moving window) for improved prediction accuracy. In this paper, we investigate the Bellwether effect on software effort estimation accuracy using moving windows. The existence of the Bellwether was empirically proven based on six postulations. We apply statistical stratification and Markov chain methodology to select the Bellwether moving window. The resulting Bellwether moving window is used to predict the software effort of a new project. Empirical results show that Bellwether effect exist in chronological datasets with a set of exemplary and recently completed projects representing the Bellwether moving window. Result from this study has shown that the use of Bellwether moving window with the Gaussian weighting function significantly improve the prediction accuracy. Solomon Mensah, Jacky W. Keung, Stephen G. MacDonell, Michael Franklin Bosu, Kwabena Ebo Bennin |
QRS | 5 |
| 2017 | Cross-Project Defect Prediction Using a Credibility Theory Based Naive Bayes ClassifierabstractSeveral defect prediction models proposed are effective when historical datasets are available. Defect prediction becomes difficult when no historical data exist. Cross-project defect prediction (CPDP), which uses projects from other sources/companies to predict the defects in the target projects proposed in recent studies has shown promising results. However, the performance of most CPDP approaches are still beyond satisfactory mainly due to distribution mismatch between the source and target projects. In this study, a credibility theory based Naïve Bayes (CNB) classifier is proposed to establish a novel reweighting mechanism between the source projects and target projects so that the source data could simultaneously adapt to the target data distribution and retain its own pattern. Our experimental results show that the feasibility of the novel algorithm design and demonstrate the significant improvement in terms of the performance metrics considered achieved by CNB over other CPDP approaches. Wai Nam Poon, Kwabena Ebo Bennin, Jianglin Huang, Passakorn Phannachitta, Jacky W. Keung |
QRS | 2 |
| 2017 | A Stratification and Sampling Model for Bellwether Moving WindowabstractAn effective method for finding the relevant number (window size) and the elapsed time (window age) of recently completed projects has proven elusive in software effort estimation.Although these two parameters significantly affect the prediction accuracy, there is no effective method to stratify and sample chronological projects to improve prediction performance of software effort estimation models.Exemplary projects (Bellwether) representing the training set have been empirically validated to improve the prediction accuracy in the domain of software defect prediction.However, the concept of Bellwether and its effect have not been empirically proven in software effort estimation as a method of selecting exemplary/relevant projects with defined window size and age.In view of this, we introduce a novel method for selecting relevant and recently completed projects referred to as Bellwether moving window for improving the software effort prediction accuracy.We first sort and cluster a pool of N projects and apply statistical stratification based on Markov chain modeling to select the Bellwether moving window.We evaluate the proposed approach using the baseline Automatically Transformed Linear Model on the ISBSG dataset.Results show that (1) Bellwether effect exist in software effort estimation dataset, (2) the Bellwether moving window with a window size of 82 to 84 projects and window age of 1.5 to 2 years resulted in an improved prediction accuracy than the traditional approach. Solomon Mensah, Jacky W. Keung, Michael Franklin Bosu, Kwabena Ebo Bennin, Patrick Kwaku Kudjo |
SEKE | 4 |
| 2016 | Filter-INC: Handling Effort-Inconsistency in Software Effort Estimation DatasetsabstractEffort-inconsistency is a situation where historical software project data used for software effort estimation (SEE) are contaminated by many project cases with similar characteristics but are completed with significantly different amount of effort. Using these data for SEE generally produces inaccurate results; however, an effective technique for its handling is yet made to be available. This study approaches the problem differently from common solutions, where available techniques typically attempt to remove every project case they have detected as outliers. Instead, we hypothesize that data inconsistency is caused by only a few deviant project cases and any attempt to remove those other cases will result in reduced accuracy, largely due to loss of useful information and data diversity. Filter-INC (short for Filtering technique for handling effort-INConsistency in SEE datasets) implements the hypothesis to decide whether a project case being detected by any existing technique should be subject to removal. The evaluation is carried out by comparing the performance of 2 filtering techniques between before and after having Filter-INC applied. The results produced from 8 real-world datasets together with 3 machine-learning models, and evaluated by 4 performance measures show a significant accuracy improvement at the confident interval of 95%. Based on the results, we recommend our proposed hypothesis as an important instrument to design a data preprocessing technique for handling effort-inconsistency in SEE datasets, definitely an important step forward in preprocessing data for a more accurate SEE model. Passakorn Phannachitta, Jacky W. Keung, Kwabena Ebo Bennin, Akito Monden, Ken-ichi Matsumoto |
APSEC | 3 |
| 2016 | Investigating the Effects of Balanced Training and Testing Datasets on Effort-Aware Fault Prediction ModelsabstractTo prioritize software quality assurance efforts, faultprediction models have been proposed to distinguish faulty modules from clean modules. The performances of such models are often biased due to the skewness or class imbalance of the datasets considered. To improve the prediction performance of these models, sampling techniques have been employed to rebalance the distribution of fault-prone and non-fault-prone modules. The effect of these techniques have been evaluated in terms of accuracy/geometric mean/F1-measure in previous studies, however, these measures do not consider the effort needed to fixfaults. To empirically investigate the effect of sampling techniqueson the performance of software fault prediction models in a morerealistic setting, this study employs Norm(Popt), an effort-awaremeasure that considers the testing effort. We performed two setsof experiments aimed at (1) assessing the effects of samplingtechniques on effort-aware models and finding the appropriateclass distribution for training datasets (2) investigating the roleof balanced training and testing datasets on performance ofpredictive models. Of the four sampling techniques applied, the over-sampling techniques outperformed the under-samplingtechniques with Random Over-sampling performing best withrespect to the Norm (Popt) evaluation measure. Also, performanceof all the prediction models improved when sampling techniqueswere applied between the rates of (20-30)% on the trainingdatasets implying that a strictly balanced dataset (50% faultymodules and 50% clean modules) does not result in the bestperformance for effort-aware models. Our results also indicatethat performances of effort-aware models are significantly dependenton the proportions of the two types of the classes in thetesting dataset. Models trained on moderately balanced datasetsare more likely to withstand fluctuations in performance as theclass distribution in the testing data varies. Kwabena Ebo Bennin, Jacky W. Keung, Akito Monden, Yasutaka Kamei, Naoyasu Ubayashi |
COMPSAC | 1 |
| 2016 | A Strategy to Determine When to Stop Using Automatic Bug LocalizationabstractInformation retrieval based automatic bug localization techniques provide developers a ranked list of suspicious buggy source entities to aid locate the ones needed to be modified and to fix the bug. However, it is unavoidable that some buggy entities are ranked low in the result list using these automatic techniques. We assume a bug localization process to address this challenge. Each time a source code entity in the ranked list is examined, the developers will have the option as to whether to continue examining the automatic bug localization result, or simply switch to using a conventional localization approach. We propose a new evaluation metric called ETC (Expected Time Cost) in the localization process, which includes the time cost of using the conventional approach. Under our assumptions, we derived simple criteria to minimize ETC. We compared the time cost of a state-of-art automatic localization method, BugLocator, with and without using our strategy in two projects. The result shows that using our proposed strategy combining both automatic localization technique together with conventional approach performs better than using only either the automatic localization technique or the conventional approach. Zhendong Shi, Jacky W. Keung, Kwabena Ebo Bennin, Nachai Limsettho, Qinbao Song |
COMPSAC | 3 |
| 2016 | Empirical Evaluation of Cross-Release Effort-Aware Defect Prediction ModelsabstractTo prioritize quality assurance efforts, various fault prediction models have been proposed. However, the best performing fault prediction model is unknown due to three major drawbacks: (1) comparison of few fault prediction models considering small number of data sets, (2) use of evaluation measures that ignore testing efforts and (3) use of n-fold cross-validation instead of the more practical cross-release validation. To address these concerns, we conducted cross-release evaluation of 11 fault density prediction models using data sets collected from 2 releases of 25 open source software projects with an effort-aware performance measure known as Norm(Popt). Our result shows that, whilst M5 and K* had the best performances, they were greatly influenced by the percentage of faulty modules present and size of data set. Using Norm(Popt) produced an overall average performance of more than 50% across all the selected models clearly indicating the importance of considering testing efforts in building fault-prone prediction models. Kwabena Ebo Bennin, Koji Toda, Yasutaka Kamei, Jacky W. Keung, Akito Monden, Naoyasu Ubayashi |
QRS | 1 |
| 2016 | Multi-Objective Optimization for Software Testing Effort EstimationabstractSoftware Testing Effort (STE), which contributes about 25-40% of the total development effort, plays a significant role in software development.In addressing the issues faced by companies in finding relevant datasets for STE estimation modeling prior to development, cross-company modeling could be leveraged.The study aims at assessing the effectiveness of cross-company (CC) and within-company (WC) projects in STE estimation.A robust multi-objective Mixed-Integer Linear Programming (MILP) optimization framework for the selection of CC and WC projects was constructed and estimation of STE was done using Deep Neural Networks.Results from our study indicate that the application of the MILP framework yielded similar results for both WC and CC modeling.The modeling framework will serve as a foundation to assist in STE estimation prior to the development of new a software project. Solomon Mensah, Jacky W. Keung, Kwabena Ebo Bennin, Michael Franklin Bosu |
SEKE | 3 |
| 2015 | Effects of Geographical, Socio-cultural and Temporal Distances on Communication in Global Software Development during Requirements Change Management - A Pilot StudyabstractTrend of software development is changing rapidly most of the software development organizations are trying to globalize their activities throughout the world. This trend leads towards a phenomenon called Global Software Development (GSD). The main reason behind the software globalization is its various benefits. Besides these benefits, software organizations are facing various challenges. One of these challenges is communication which is considered a big challenge in GSD and it becomes more complicated during the Requirements Change Management (RCM) process due to three factors, they are Geographical, Socio-cultural and Temporal distances. This paper presents a framework which shows the effect of these factors on communication during RCM process in GSD. Communication is the core function of collaboration which allows information to be exchanged between the team members. A pilot study has been conducted in three GSD organizations. A quantitative research method has been used to collect data. The findings from the survey data show that these three factors have a strong negative impact on communication process in GSD. Arif Ali Khan, Jacky W. Keung, Shahid Hussain 0001, Kwabena Ebo Bennin |
ENASE | 4 |