VLDB 2026 Research / reviewers in the wild / expert
Solomon Mensah
dblp:183/9381
· DBLP profile ↗
20ranked-venue papers
7as first author
3since 2021 · last 2024
0000-0001-5629-2661ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 18 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Object detection in adverse weather condition for autonomous vehicles
Emmanuel Owusu Appiah, Solomon Mensah |
Multim. Tools Appl. | 2 |
| 2023 | An Optimal Spacing Approach for Sampling Small-sized DatasetsabstractContext: There has been a growing research focus in conventional machine learning techniques for software effort estimation (SEE).However, there is a limited number of studies that seek to assess the performance of deep learning approaches in SEE.This is because the sizes of SEE datasets are relatively small.Purpose: This study seeks to define a threshold for small-sized datasets in SEE, and investigates the performance of selected conventional machine learning and deep learning models on small-sized datasets.Method: Plausible SEE datasets with their number of project instances and features are extracted from existing literature and ranked.Eubank's optimal spacing theory is used to discretize the ranking of the project instances into three classes (small, medium and large).Five conventional machine learning models and two deep learning models are trained on each dataset classified as small-sized using the leave-one-out cross-validation.The mean absolute error is used to assess the prediction performance of each model.Result: Findings from the study contradicts existing knowledge by demonstrating that deep learning models provide improved prediction performance as compared to the conventional machine learning models on small-sized datasets.Conclusion: Deep learning can be adopted for SEE with the application of regularisation techniques. Samuel Abedu, Solomon Mensah, Frederick Boafo, Eva Bushel, Elizabeth Akuafum |
SEKE | 2 |
| 2022 | A classification scheme to improve conclusion instability using Bellwether moving windowsabstractAbstract Context The use of a subset of recently completed and exemplary data, namely, Bellwether moving window (BMW) has proven successful to result in improved accuracy in software effort estimation (SEE). These outcomes were achieved based on the theory that estimation outcome of a future event depends on previous events. Thus, the existence of a BMW yield improved prediction accuracy for new project estimation. However, the conclusion instability problem across learners still threatens the reliability of SEE for new projects. Such instability concerns are attributed to the data subset considered for the training and validation needs of learners. Objective To investigate whether the use of BMWs together with an effort classification scheme can minimize the conclusion instability problem across learners. Method We apply a Bellwether method comprising of three operators, namely, SORT+CLUSTER, GENERATE_TPM, and APPLY to sample the BMW from a pool of chronological projects from the Maxwell and International Software Benchmarking Standards Group (ISBSG) datasets. The sampled BMW is benchmarked against the entire collection of preprocessed projects, namely, growing portfolio to evaluate prediction and classification accuracy across a set of learners–ElasticNet regression, deep neural networks, and automatically transformed linear model. Results (1) BMW exists in the studied projects and (2) training the learners with a BMW of average window size 28.5%–75.5% of the growing portfolio (not older than 3 years) relatively minimizes the conclusion instability of prediction results. Conclusion When BMWs are available, we recommend their use for estimating the effort for a new project to minimize the conclusion instability problem. Solomon Mensah, Patrick Kwaku Kudjo |
J. Softw. Evol. Process. | 1 |
| 2020 | An automatic software vulnerability classification framework using term frequency-inverse gravity moment and feature selection
Jinfu Chen 0001, Patrick Kwaku Kudjo, Solomon Mensah, Selasie Brown Aformaley, George Akorfu |
J. Syst. Softw. | 3 |
| 2020 | The effect of Bellwether analysis on software vulnerability severity prediction models
Patrick Kwaku Kudjo, Jinfu Chen 0001, Solomon Mensah, Richard Amankwah, Christopher Kudjo |
Softw. Qual. J. | 3 |
| 2019 | Improving the Accuracy of Vulnerability Report Classification Using Term Frequency-Inverse Gravity MomentabstractSoftware vulnerability analysis is one of the critical issues in the software industry, and vulnerability classification plays a major role in this analysis. A typical vulnerability classification model usually involves a stage of term selection, in which the relevant terms are identified via feature selection. It also involves a stage of term weighting, in which document weights for the selected terms are computed, and a stage for classifier learning. Generally, the term frequency-inverse document frequency (TF-IDF) is the most widely used term-weighting method. However, empirical evidence shows that the TF-IDF is plagued with issues pertaining to its effectiveness. This paper introduces a new approach for vulnerability classification, which is based on term frequency and inverse gravity moment (TF-IGM). The proposed method is validated by empirical experiments using three machine learning algorithms on ten publicly available vulnerability datasets. The result shows that TF-IGM outperforms the benchmark method across the applications studied. Patrick Kwaku Kudjo, Jinfu Chen 0001, Minmin Zhou, Solomon Mensah, Rubing Huang |
QRS | 4 |
| 2018 | An Inception Architecture-Based Model for Improving Code Readability ClassificationabstractThe process of classifying a piece of source code into a Readable or Unreadable class is referred to as Code Readability Classification. To build accurate classification models, existing studies focus on handcrafting features from different aspects that intuitively seem to correlate with code readability, and then exploring various machine learning algorithms based on the newly proposed features. On the contrary, our work opens up a new way to tackle the problem by using the technique of deep learning. Specifically, we propose IncepCRM, a novel model based on the Inception architecture that can learn multi-scale features automatically from source code with little manual intervention. We apply the information of human annotators as the auxiliary input for training IncepCRM and empirically verify the performance of IncepCRM on three publicly available datasets. The results show that: 1) Annotator information is beneficial for model performance as confirmed by robust statistical tests (i.e., the Brunner-Munzel test and Cliff's delta); 2) IncepCRM can achieve an improved accuracy against previously reported models across all datasets. The findings of our study confirm the feasibility and effectiveness of deep learning for code readability classification. Qing Mi, Jacky W. Keung, Yan Xiao 0002, Solomon Mensah, Xiupei Mei |
EASE | 4 |
| 2018 | MAHAKIL: diversity based oversampling approach to alleviate the class imbalance issue in software defect predictionabstractThis study presents MAHAKIL, a novel and efficient synthetic over-sampling approach for software defect datasets that is based on the chromosomal theory of inheritance. Exploiting this theory, MAHAKIL interprets two distinct sub-classes as parents and generates a new instance that inherits different traits from each parent and contributes to the diversity within the data distribution. We extensively compare MAHAKIL with five other sampling approaches using 20 releases of defect datasets from the PROMISE repository and five prediction models. Our experiments indicate that MAHAKIL improves the prediction performance for all the models and achieves better and more significant pf values than the other oversampling approaches, based on robust statistical tests. Kwabena Ebo Bennin, Jacky W. Keung, Passakorn Phannachitta, Akito Monden, Solomon Mensah |
ICSE | 5 |
| 2018 | Revisiting the Conclusion Instability Issue in Software Effort Estimation (S)abstractConclusion instability is the absence of observing the same effect under varying experimental conditions.Deep Neural Network (DNN) and ElasticNet software effort estimation (SEE) models were applied to two SEE datasets with the view of resolving the conclusion instability issue and assessing the suitability of ElasticNet as a viable SEE benchmark model.Results were mixed as both model types attain conclusion stability for the Kitchenham dataset whilst conclusion instability existed in the Desharnais dataset.ElasticNet was outperformed by DNN and as such it is not recommended to be used as a SEE benchmark model. Michael Franklin Bosu, Solomon Mensah, Kwabena Ebo Bennin, Diab Abuaiadah |
SEKE | 2 |
| 2018 | Duplex output software effort estimation model with self-guided interpretation
Solomon Mensah, Jacky W. Keung, Michael Franklin Bosu, Kwabena Ebo Bennin |
Inf. Softw. Technol. | 1 |
| 2018 | Not all bug reopens are negative: A case study on eclipse bug reports
Qing Mi, Jacky W. Keung, Yuqi Huo, Solomon Mensah |
Inf. Softw. Technol. | 4 |
| 2018 | Improving code readability classification using convolutional neural networks
Qing Mi, Jacky W. Keung, Yan Xiao 0002, Solomon Mensah, Yujin Gao |
Inf. Softw. Technol. | 4 |
| 2018 | On the value of a prioritization scheme for resolving Self-admitted technical debt
Solomon Mensah, Jacky W. Keung, Jeffrey Svajlenko, Kwabena Ebo Bennin, Qing Mi |
J. Syst. Softw. | 1 |
| 2018 | Investigating the Significance of the Bellwether Effect to Improve Software Effort Prediction: Further Empirical StudyabstractContext: In addressing how best to estimate how much effort is required to develop software, a recent study found that using exemplary and recently completed projects [forming Bellwether moving windows (BMW)] in software effort prediction (SEP) models leads to relatively improved accuracy. More studies need to be conducted to determine whether the BMW yields improved accuracy in general, since different sizing and aging parameters of the BMW are known to affect accuracy. Objective: To investigate the existence of exemplary projects (Bellwethers) with defined window size and age parameters, and whether their use in SEP improves prediction accuracy. Method: We empirically investigate the moving window assumption based on the theory that the prediction outcome of a future event depends on the outcomes of prior events. Sampling of Bellwethers was undertaken using three introduced Bellwether methods (SSPM, SysSam, and RandSam). The ergodic Markov chain was used to determine the stationarity of the Bellwethers. Results: Empirical results show that 1) Bellwethers exist in SEP and 2) the BMW has an approximate size of 50 to 80 exemplary projects that should not be more than 2 years old relative to the new projects to be estimated. Conclusion: The study's results add further weight to the recommended use of Bellwethers for improved prediction accuracy in SEP. Solomon Mensah, Jacky W. Keung, Stephen G. MacDonell, Michael Franklin Bosu, Kwabena Ebo Bennin |
IEEE Trans. Reliab. | 1 |
| 2018 | MAHAKIL: Diversity Based Oversampling Approach to Alleviate the Class Imbalance Issue in Software Defect PredictionabstractHighly imbalanced data typically make accurate predictions difficult. Unfortunately, software defect datasets tend to have fewer defective modules than non-defective modules. Synthetic oversampling approaches address this concern by creating new minority defective modules to balance the class distribution before a model is trained. Notwithstanding the successes achieved by these approaches, they mostly result in over-generalization (high rates of false alarms) and generate near-duplicated data instances (less diverse data). In this study, we introduce MAHAKIL, a novel and efficient synthetic oversampling approach for software defect datasets that is based on the chromosomal theory of inheritance. Exploiting this theory, MAHAKIL interprets two distinct sub-classes as parents and generates a new instance that inherits different traits from each parent and contributes to the diversity within the data distribution. We extensively compare MAHAKIL with SMOTE, Borderline-SMOTE, ADASYN, Random Oversampling and the No sampling approach using 20 releases of defect datasets from the PROMISE repository and five prediction models. Our experiments indicate that MAHAKIL improves the prediction performance for all the models and achieves better and more significant pf values than the other oversampling approaches, based on Brunner's statistical significance test and Cliff's effect sizes. Therefore, MAHAKIL is strongly recommended as an efficient alternative for defect prediction models built on highly imbalanced datasets. Kwabena Ebo Bennin, Jacky W. Keung, Passakorn Phannachitta, Akito Monden, Solomon Mensah |
IEEE Trans. Software Eng. | 5 |
| 2017 | Identifying Textual Features of High-Quality Questions: An Empirical Study on Stack OverflowabstractBackground: Stack Overflow (SO) is a programming-specific Q&A website that serves as a valuable repository of software engineering knowledge. For SO members, formulating a good question is the first step towards eliciting satisfactory responses. Aims: To guide SO members on how to make a good question, we conduct an empirical study using the publicly available Stack Overflow Data Dump for the period of 2008-2016. Method: We first choose 25 features along 5 dimensions to represent the textual characteristics that we are interested in. Making use of the Boruta algorithm, we then capture all features that are either strongly or weakly relevant to the question quality. Results: The results show that the number of tags and code snippets are the most discriminative features, whereas there is only a weak correlation between the question quality and the sentiment-related factors. Based on the empirical evidence, we provide useful and usable suggestions to SO members on how to optimize their questions. Conclusions: We consider that our findings will provide SO members with a better understanding of the patterns behind high-quality questions, this is to support effective and efficient utilization of Q&A websites as the ultimate goal. Qing Mi, Yujin Gao, Jacky W. Keung, Yan Xiao 0002, Solomon Mensah |
APSEC | 5 |
| 2017 | The Significant Effects of Data Sampling Approaches on Software Defect Prioritization and ClassificationabstractContext: Recent studies have shown that performance of defect prediction models can be affected when data sampling approaches are applied to imbalanced training data for building defect prediction models. However, the magnitude (degree and power) of the effect of these sampling methods on the classification and prioritization performances of defect prediction models is still unknown. Goal: To investigate the statistical and practical significance of using resampled data for constructing defect prediction models. Method: We examine the practical effects of six data sampling methods on performances of five defect prediction models. The prediction performances of the models trained on default datasets (no sampling method) are compared with that of the models trained on resampled datasets (application of sampling methods). To decide whether the performance changes are significant or not, robust statistical tests are performed and effect sizes computed. Twenty releases of ten open source projects extracted from the PROMISE repository are considered and evaluated using the AUC, pd, pf and G-mean performance measures. Results: There are statistical significant differences and practical effects on the classification performance (pd, pf and G-mean) between models trained on resampled datasets and those trained on the default datasets. However, sampling methods have no statistical and practical effects on defect prioritization performance (AUC) with small or no effect values obtained from the models trained on the resampled datasets. Conclusions: Existing sampling methods can properly set the threshold between buggy and clean samples, while they cannot improve the prediction of defect-proneness itself. Sampling methods are highly recommended for defect classification purposes when all faulty modules are to be considered for testing. Kwabena Ebo Bennin, Jacky W. Keung, Akito Monden, Passakorn Phannachitta, Solomon Mensah |
ESEM | 5 |
| 2017 | Investigating the Significance of Bellwether Effect to Improve Software Effort EstimationabstractBellwether effect refers to the existence of exemplary projects (called the Bellwether) within a historical dataset to be used for improved prediction performance. Recent studies have shown an implicit assumption of using recently completed projects (referred to as moving window) for improved prediction accuracy. In this paper, we investigate the Bellwether effect on software effort estimation accuracy using moving windows. The existence of the Bellwether was empirically proven based on six postulations. We apply statistical stratification and Markov chain methodology to select the Bellwether moving window. The resulting Bellwether moving window is used to predict the software effort of a new project. Empirical results show that Bellwether effect exist in chronological datasets with a set of exemplary and recently completed projects representing the Bellwether moving window. Result from this study has shown that the use of Bellwether moving window with the Gaussian weighting function significantly improve the prediction accuracy. Solomon Mensah, Jacky W. Keung, Stephen G. MacDonell, Michael Franklin Bosu, Kwabena Ebo Bennin |
QRS | 1 |
| 2017 | A Stratification and Sampling Model for Bellwether Moving WindowabstractAn effective method for finding the relevant number (window size) and the elapsed time (window age) of recently completed projects has proven elusive in software effort estimation.Although these two parameters significantly affect the prediction accuracy, there is no effective method to stratify and sample chronological projects to improve prediction performance of software effort estimation models.Exemplary projects (Bellwether) representing the training set have been empirically validated to improve the prediction accuracy in the domain of software defect prediction.However, the concept of Bellwether and its effect have not been empirically proven in software effort estimation as a method of selecting exemplary/relevant projects with defined window size and age.In view of this, we introduce a novel method for selecting relevant and recently completed projects referred to as Bellwether moving window for improving the software effort prediction accuracy.We first sort and cluster a pool of N projects and apply statistical stratification based on Markov chain modeling to select the Bellwether moving window.We evaluate the proposed approach using the baseline Automatically Transformed Linear Model on the ISBSG dataset.Results show that (1) Bellwether effect exist in software effort estimation dataset, (2) the Bellwether moving window with a window size of 82 to 84 projects and window age of 1.5 to 2 years resulted in an improved prediction accuracy than the traditional approach. Solomon Mensah, Jacky W. Keung, Michael Franklin Bosu, Kwabena Ebo Bennin, Patrick Kwaku Kudjo |
SEKE | 1 |
| 2016 | Multi-Objective Optimization for Software Testing Effort EstimationabstractSoftware Testing Effort (STE), which contributes about 25-40% of the total development effort, plays a significant role in software development.In addressing the issues faced by companies in finding relevant datasets for STE estimation modeling prior to development, cross-company modeling could be leveraged.The study aims at assessing the effectiveness of cross-company (CC) and within-company (WC) projects in STE estimation.A robust multi-objective Mixed-Integer Linear Programming (MILP) optimization framework for the selection of CC and WC projects was constructed and estimation of STE was done using Deep Neural Networks.Results from our study indicate that the application of the MILP framework yielded similar results for both WC and CC modeling.The modeling framework will serve as a foundation to assist in STE estimation prior to the development of new a software project. Solomon Mensah, Jacky W. Keung, Kwabena Ebo Bennin, Michael Franklin Bosu |
SEKE | 1 |