Akito Monden

dblp:41/2487 · DBLP profile ↗
← Back
81ranked-venue papers
4as first author
18since 2021 · last 2025
0000-0003-4295-207XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 62 · 4 first-author · 16 since 2021Artificial intelligence and machine learning · 12 · 4 since 2021Human-computer interaction and ubiquitous computing · 12 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021
YearPublicationVenuePosition
2025 Evaluating ChatGPT's Ability to Detect Naming Bugs in Java Methods
Kinari Nishiura, Atsuya Matsutomo, Akito Monden
ENASE3
2025 Plaintext in the Wild: Investigating Secure Connection Label Accuracy for Android Apps
abstract
Smartphones have become deeply integrated into daily life, prompting widespread concern over how mobile apps handle user data. The Google Play Store requires Android developers to disclose whether their apps encrypt user data during transmission via a "secure connection" label in the Data Safety section. However, these labels are self-declared, and their consistency with actual app behavior remains unclear. In this study, we empirically evaluate the consistency of secure connection labels with real-world data transmission practices by dynamically analyzing network traffic from over 12,000 top-ranked Android apps. We identify 65 apps transmitting sensitive data without encryption, and they collectively account for over 5.842 billion installs, indicating that a substantial number of users may be affected. A majority of them (i.e., 46 apps) falsely claim to encrypt data while transmitting sensitive information in plaintext, and the others correctly disclose the lack of encryption and transmit sensitive data in plaintext. We contacted developers of the inconsistent apps, resulting in several label updates. Our findings reveal a disconnect between declared and actual security practices and offer concrete recommendations to improve the integrity of privacy disclosures on app marketplaces.
Yusei Sakuraba, Hiroki Inayoshi, Shoichi Saito, Akito Monden
SCAM4
2024 An Empirical Study of the Impact of Test Strategies on Online Optimization for Ensemble-Learning Defect Prediction
abstract
Ensemble learning methods have been used to enhance the reliability of defect prediction models. However, there is an inconclusive stability of a single method attaining the highest accuracy among various software projects. This work aims to improve the performance of ensemble-learning defect prediction among such projects by helping select the highest accuracy ensemble methods. We employ bandit algorithms (BA), an online optimization method, to select the highest-accuracy ensemble method. Each software module is tested sequentially, and bandit algorithms utilize the test outcomes of the modules to evaluate the performance of the ensemble learning methods. The test strategy followed might impact the testing effort and prediction accuracy when applying online optimization. Hence, we analyzed the test order's influence on BA's performance. In our experiment, we used six popular defect prediction datasets, four ensemble learning methods such as bagging, and three test strategies such as testing positive-prediction modules first (PF). Our results show that when BA is applied with PF, the prediction accuracy improved on average, and the number of found defects increased by 7% on a minimum of five out of six datasets (although with a slight increase in the testing effort by about 4% from ordinal ensemble learning). Hence, BA with PF strategy is the most effective to attain the highest prediction accuracy using ensemble methods on various projects.
Kensei Hamamoto, Masateru Tsunoda, Amjed Tahir, Kwabena Ebo Bennin, Akito Monden, Koji Toda, Keitaro Nakasai, Ken-ichi Matsumoto
ICSME5
2024 An Empirical Study on Ambiguous Words in Software Requirements Specifications of Local Government and Library Systems
abstract
Ambiguous words in software requirements specifications (SRSs) can cause serious misunderstanding, confusion and/or troubles between software purchasers and developers due to differences in interpretation of the words. This paper quantitatively analyzes existing SRSs to clarify (1) how many ambiguous words are actually included in SRSs and (2) how many these words require correction. This paper targets the Request For Proposals (RFPs), which describes initial requirements, of ten local government systems and ten library systems in Japan. As candidates of ambigous words, we analyzed ten Japanese words: (1) Ippanteki (common, commonly used, widely used), (2) Juudai (serious, seriously, critical, severe), (3) Juubun (sufficient, sufficiently), (4) Tekido (adequate, adequately, appropriate, reasonable), (5) Kouritsuteki (efficient, efficiently), (6) Juuyou (important), (7) Shunji (immediate, immediately), (8) Sugu (im-mediate, immediately), (9) Sokuji (immediate, immediately) and (10) Tadachi (immediate, immediately). As a result of the analysis, we found that among ten words, juubunnna (sufficient) was most frequently appeared in SRSs, and 43% of cases required correction when this word appeared. In addition, we also found that the number of ambiguous words varied greatly among the SRSs, and that larger SRSs did not necessarily contain more ambiguous words.
Toru Nakamichi, Kinari Nishiura, Mariko Sasakura, Akito Monden
SERA4
2024 Porting a Python Application to the Web Using Django: A Case Study of an Archaeological Image Processing System
abstract
In recent years, desktop applications are often ported to the Web. This is because Web applications running in a cloud environment have many advantages, for example, they can be used by a wide variety of clients over the Internet and can dynamically allocate computing resources according to demand. Such porting is also very important in terms of effectively utilizing existing software assets in a modern environment. However, porting to the Web involves numerous considerations that are not easy for those without the knowledge and skills to perform. In this paper we describe in detail our experience of porting a desktop Python application, which uses the image processing library OpenCV and the GUI library PySimpleGUI, into a web application, which uses the web framework Django, the CSS framework Bootstrap, the database management system MySQL and phpMyAdmin. We also employed Docker and Docker Compose for flexible development and deployment. Through our experience, we have identified six key aspects to consider when porting a desktop application to the web. This paper will elaborate on these aspects and how to deal with them. This paper also reports which parts of source code could be reused, which parts had to be newly developed, and the size and time required to conduct reuse, modification, and additional development.
Hikaru Tomita, Mariko Sasakura, Kinari Nishiura, Hiroki Inayoshi, Akito Monden
SERA5
2024 Identifying Security Bugs in Issue Reports: Comparison of BERT, N-gram IDF and ChatGPT
abstract
In recent software development, which has become increasingly large and complex, a huge number of issues including bugs, improvements, new feature requests are reported on a daily basis, and there is a risk of missing urgent bugs. Security bugs are particularly urgent because they can cause serious problems such as mal ware infections, and must be resolved quickly. Therefore, it is important to develop a technique to automatically identify security bugs in a large number of issue reports. The goal of this study is to empirically evaluate recent machine learning methods to identify security bugs using issue report text written in natural language as input. Specifically, this paper focuses on the two-class classification model using BERT, a language model based on the Transformer architecture. The model is constructed by fine-tuning a pre-trained model of BERT with the text of issue reports. In our experiment, we performed classification of issue reports obtained from four open source software projects. As a comparison method, we employ a classification model using features obtained by N -gram IDF, which is an extension of the conventional Bag-of- Words approach. We also employ ChatGPT, which is a general-purpose chatbot that utilizes a large-scale language model (LLM). As a result of our experiment, the BERT-based model showed the best classification performance in terms of F1 score. ChatGPT was better than the N-gram IDF based model, but far behind the BERT.
Daiki Yokoyama, Kinari Nishiura, Akito Monden
SERA3
2024 Extended Association Rule Mining and Its Application to Software Engineering Data Sets
abstract
Association rule mining is a highly effective approach to data analysis for datasets of varying sizes, accommodating diverse feature values. Nevertheless, deriving practical rules from datasets with numerical variables presents a challenge, as these variables must be discretized beforehand. Quantitative association rule mining addresses this issue, allowing the extraction of valuable rules. This paper introduces an extension to quantitative association rules, incorporating a two-variable function in their consequent part. The use of correlation functions, statistical test functions, and error functions is also introduced. We illustrate the utility of this extension through three case studies employing software engineering datasets. In case study 1, we successfully pinpointed the conditions that result in either a high or low correlation between effort and software size, offering valuable insights for software project managers. In case study 2, we effectively identified the conditions that lead to a high or low correlation between the number of bugs and source lines of code, aiding in the formulation of software test planning strategies. In case study 3, we applied our approach to the two-step software effort estimation process, uncovering the conditions most likely to yield low effort estimation errors.
Hidekazu Saito, Kinari Nishiura, Akito Monden, Shuji Morisaki
Int. J. Softw. Eng. Knowl. Eng.3
2023 Subject Experiments with a Learning Support System for Grover's Algorithm
abstract
An appropriate quantum algorithm is needed for each problem to achieve the full performance of a quantum computer. It is necessary to understand the principles of quantum computation to implement quantum algorithms. Quantum computation simulators can assist in understanding the principles and behavior of quantum computation. Combining this with explanations of quantum algorithms using visualization techniques and interactions can further aid learning. In this study, to support the understanding of quantum computation, a difficult concept for beginners, we develop an interactive system in which diagrams and graphs are added to a quantum computation simulator to support the explanation and understanding of qubits and quantum algorithms. To verify the effectiveness of the interaction in supporting learning, a subject experiment is conducted to compare the amount of knowledge understood by a group of novice quantum computer students who learned using the developed system with a group who learned using paper-based materials.
Hayato Yasunaga, Mariko Sasakura, Akito Monden
IV3
2023 Analysis of Programming Performance Based on 2-grams of Keystrokes and Mouse Operations
Kazuki Matsumoto, Kinari Nishiura, Mariko Sasakura, Akito Monden
SERA4
2022 Preliminary Analysis of Review Method Selection Based on Bandit Algorithms
abstract
To enhance the reliability of software, it is important is to review all software artifacts (e.g., design documents) to remove defects as earlier as possible. There are various review methods available, and project managers face the challenge of choosing a suitable method for their current projects. One of approaches to support the selection of review methods is to evaluate review methods beforehand, to identify the most effective method on average. However, past studies have not evaluated review methods thoroughly as the process can be time-consuming. We propose a bandit-algorithm (BA) based method to evaluate and then dynamically select a suitable review method (from a list of candidates). In our experiments, we assume that the proposed method is applied to design document review on basic design phase. We performed experiments based on a simulation, instead of using an actual dataset. On our simulation, when a review method is selected by our BA method, productivity (i.e., total development time) was improved by about 1.25 times, and it was the second highest among candidates of review methods.
Takuto Kudo, Masateru Tsunoda, Amjed Tahir, Kwabena Ebo Bennin, Koji Toda, Keitaro Nakasai, Akito Monden, Ken-ichi Matsumoto
APSEC7
2022 Gaze Analysis in Spot the Difference
abstract
The ability to see and find things is very important in our daily lives. For example, when looking for mistakes in debugging a program, or when looking for misspellings in documents, etc., the visual sense is mainly used. The search may or may not be successful. Is there any difference in the way of searching when the search is successful or unsuccessful? The aim of this study is to analyse the gaze while searching and to clarify the differences between successful and unsuccessful searches, using ‘spot the difference’ as a subject. We have developed an experimental application to measure people's gaze while they are looking ‘spot the difference'. In the experiment conducted in this study, 29 subjects have asked to perform ‘spot the difference’ of multiple problems and their gaze have been measured. Analysis of the data obtained from this experiment shows that in many cases, subjects who could not find a difference were not looking at the location of the difference. On the other hand, the existence of ‘cases of looking but not finding’, in which the difference is not detected even though the difference is fully looked at, is also identified. In the present experiment, ‘looking but not finding’ cases account for 15% of all non-correct responses in all questions.
Mariko Sasakura, Syouta Toda, Akito Monden
IV3
2022 Using Bandit Algorithms for Selecting Feature Reduction Techniques in Software Defect Prediction
abstract
Background: Selecting a suitable feature reduction technique. when building a defect prediction model, can be challenging. Different techniques can result in the selection of different independent variables which have an impact on the overall performance of the prediction model. To help in the selection, previous studies have assessed the impact of each feature reduction technique using different datasets. However, there are many reduction techniques, and therefore some of the well-known techniques have not been assessed by those studies. Aim: The goal of the study is to select a high-accuracy reduction technique from several candidates without preliminary assessments. Method: We utilized bandit algorithm (BA) to help with the selection of best features reduction technique for a list of candidates. To select the best feature reduction technique, BA evaluates the prediction accuracy of the candidates, comparing testing results of different modules with their prediction results. By substituting the reduction technique for the prediction method, BA can then be used to select the best reduction technique. In the experiment, we evaluated the performance of BA to select suitable reduction technique. We performed cross version defect prediction using 14 datasets. As feature reduction techniques, we used two assessed and two non-assessed techniques. Results: Using BA, the prediction accuracy was higher or equivalent than existing approaches on average, compared with techniques selected based on an assessment. Conclusions: BA can have larger impact on improving prediction models by helping not only on selecting suitable models, but also in selecting suitable feature reduction techniques.
Masateru Tsunoda, Akito Monden, Koji Toda, Amjed Tahir, Kwabena Ebo Bennin, Keitaro Nakasai, Masataka Nagura, Ken-ichi Matsumoto
MSR2
2021 Using Bandit Algorithms for Project Selection in Cross-Project Defect Prediction
abstract
Background: defect prediction model is built using historical data from previous versions/releases of the same project. However, such historical data may not exist in case of newly developed projects. Alternatively, one can train a model using data obtained from external projects. This approach is known as cross-project defect prediction (CPDP). In CPDP, it is still difficult to utilize external projects' data or decide which particular project to use to train a model. Aim: to address this issue, we apply bandit algorithm (BA) to CPDP in order to select the most suitable training project from a set of projects. Method: BA-based prediction iteratively reselects the project after each module is tested, considering the accuracy of the predictions. As baselines, we used simple CPDP methods such as training a model with randomly selected project. All models were built using logistic regression. Results: We experimented our approach on two datasets (NASA and DAMB, with a total of 12 projects). The BA-based defect prediction models resulted in, on average, a higher accuracy (AUC and F1 score) than the baselines. Conclusion: in this preliminarily study, we demonstrate the feasibility of using BA in the context of CPDP. Our initial assessment shows that the use BA for predicting defects in CPDP is promising and may outperform existing approaches.
Takuya Asano, Masateru Tsunoda, Koji Toda, Amjed Tahir, Kwabena Ebo Bennin, Keitaro Nakasai, Akito Monden, Ken-ichi Matsumoto
ICSME7
2021 Human Resource Analysis Based on Used Libraries in Eclipse Projects on GitHub
abstract
As current software development increasingly relies on libraries and frameworks, the knowledge and experience to use libraries is considered an important skill in software development. This paper attempts to analyze and identify the types of skills of contributors in five Eclipse projects, focusing on the libraries used by the contributors. We found that standard util libraries and I/O libraries are used in all projects, while specific libraries such as maven libraries and servlet libraries are used in one of the projects. Also, in library category analysis, we found that two projects require the security specialist that can use security libraries. Finally, in library provider analysis, we found that Oracle and JUnit libraries are used in all projects, which indicates that developers are recommended to learn these libraries.
Wilson Chukwu Emmanuel, Akito Monden
SNPD2
2021 Association Metrics Between Two Continuous Variables for Software Project Data
abstract
The correlation coefficient is commonly used in analyses of software project data sets for the purpose of quantifying the relationship between two variables. However, while there are various types of relationships between two variables, the correlation coefficient cannot distinguish between these types. This study proposes new metrics between two continuous variables that have the potential to characterize the relationship types.
Takumi Kanehira, Akito Monden, Zeynep Yücel
SNPD2
2021 A Simulation Model of Software Quality Assurance in the Software Lifecycle
abstract
Software quality assurance (SQA) is a series of activities within the software development lifecycle that repetitively verify or test the software deliverables to ensure their quality. In this paper, we propose a simulation model of SQA to quantitatively demonstrate the positive effect of adding quality assurance (QA) effort especially in early phases of software development. The proposed model can represent the relationship among the number of bugs in each phase, the amount of QA effort, the expected number of detectable bugs and the amount of bug fixing effort. The model can simulate the different QA strategies in a given software development context; thus, it is useful to identify the best or better strategies to improve software quality with smaller QA and bug fixing effort.
Hiroto Nakahara, Akito Monden, Zeynep Yücel
SNPD2
2021 Effectiveness of Explaining a Program to Others in Finding Its Bugs
abstract
Explaining a program to others helps get others to find bugs and for the explainer him/herself to find bugs. However, to the best of our knowledge, there is no quantitative evidence that explaining a program to others helps the explainer find bugs. This study aims to show quantitatively, using an experimental evaluation, that the explainer himself can find new bugs by explaining the program to others. In the experiment, subjects first review a program that contains many bugs and try to find as many bugs as possible. Next, they are required to explain the program aloud to others. We see if they notice any new bugs themselves during the explanation. As a result of the experiment, five out of the six subjects could find new bugs when explaining the program to others. According to the questionnaire to the subjects, the subjects who find many bugs feel that they can understand the program better by explaining it to others.
Toshihiro Nakamura, Akito Monden, Mariko Sasakura, Hidetake Uwano
SNPD2
2021 Task estimation for software company employees based on computer interaction logs
Florian Pellegrin, Zeynep Yücel, Akito Monden, Pattara Leelaprute
Empir. Softw. Eng.3
2020 Estimating Level of Engagement from Ocular Landmarks
abstract
E-learning offers many advantages like being economical, flexible and customizable, but also has challenging aspects such as lack of – social-interaction, which results in contemplation and sense of remoteness. To overcome these and sustain learners’ motivation, various stimuli can be incorporated. Nevertheless, such adjustments initially require an assessment of engagement level. In this respect, we propose estimating engagement level from facial landmarks exploiting the facts that (i) perceptual decoupling is promoted by blinking during mentally demanding tasks; (ii) eye strain increases blinking rate, which also scales with task disengagement; (iii) eye aspect ratio is in close connection with attentional state and (iv) users’ head position is correlated with their level of involvement. Building empirical models of these actions, we devise a probabilistic estimation framework. Our results indicate that high and low levels of engagement are identified with considerable accuracy, whereas medium levels are inherently more challenging, which is also confirmed by inter-rater agreement of expert coders.
Zeynep Yücel, Serina Koyama, Akito Monden, Mariko Sasakura
Int. J. Hum. Comput. Interact.3
2019 Algorithmic Expressions for Assessing Algorithmic Thinking Ability of Elementary School Children
abstract
This Research to Practice Full Paper presents the development of the algorithmic expressions for the assessment tools for assessing algorithmic thinking ability of elementary school children. In Japan, elementary school children will be required to learn computer programming as an interdisciplinary element appearing throughout the curriculum in 2020. The purpose of this programming education is to nurture Computational Thinking (CT) for elementary school children in Japan. However, almost no discussion has been conducted in Japan on how to measure the level of CT an elementary school child has acquired. Since the definition of CT is not very firm, it is not easy to measure the levels of CT. Therefore, several organizations have issued operational definitions of CT. Among the concepts of CT in those operational definitions, Algorithmic Thinking was chosen as a representative of CT, and the assessment tools for evaluating Algorithmic Thinking ability have been developed in this research. The assessment tool was conducted in the experimental Computer Science Unplugged classes and in the control classes in two elementary schools in Japan. There were in total 152 children in the classes, and all of them were 5th grade children. By answering the questions in the assessment tool, each child got a score between 0 and 15. The scores were statistically analyzed.
Yasumasa Oomori, Hidekuni Tsukamoto, Hideo Nagumo, Yasuhiro Takemura, Kouki Iida, Akito Monden, Ken-ichi Matsumoto
FIE6
2019 Effect of Grasping Uniformity on Estimation of Grasping Region from Gaze Data
abstract
This study explores estimation of grasping region of objects from gaze data. Our study distinguishes from previous works by accounting for "grasping uniformity" of the objects. In particular, we consider three types of graspable objects: (i) with a well-defined graspable part (e.g. handle), (ii) without a grip but with an intuitive grasping region, (iii) without any grip or intuitive grasping region. We assume that these types define how "uniform" grasping region is across different graspers. In experiments, we use "Learning to grasp" data set and apply the method of [Pramot et al. 2018] for estimating grasping region from gaze data. We compute similarity of estimations and ground truth annotations for the three types of objects regarding subjects (a) who perform free viewing and (b) who view the images with the intention of grasping. In line with many previous studies, similarity is found to be higher for non-graspers. An interesting finding is that the difference in similarity (between free viewing and motivated to grasp) is higher for type-iii objects; and comparable for type-i and ii objects. Based on this, we believe that estimation of grasping region from gaze data offers a larger potential to "learn" particularly grasping of type-iii objects.
Pimwalun Witchawanitchanun, Zeynep Yücel, Akito Monden, Pattara Leelaprute
HAI3
2019 Data Smoothing for Software Effort Estimation
abstract
The goal of this paper is to improve the estimation performance of software development effort by mitigating the problem caused by outliers in a historical software project data set, which is used to construct an effort estimation model. To date, outlier removal methods have been proposed to solve this problem; however, they are not always effective because removing outliers reduces the number of data points (= software projects in our case) in a data set, and a model built from a small data set often suffers from lack of generality. In such a case, estimation performance can become even worse. In this paper we propose a method called data smoothing to mitigate the problem of outliers without reducing the number of data points. We consider that data points are outliers if they do not meet the assumption of Analogy-Based Estimation (ABE) such that “projects with similar features require similar development efforts.” The proposed method changes the effort values (person-months or person-hours) in a data set so as to satisfy this assumption; and by this way, all outliers become non-outliers without decreasing the data points. As a result of experimental evaluation using 8 software development data sets, we found that the proposed data smoothing showed the same or higher effort estimation accuracy than the non-smoothing case, while conventional outlier removal method showed worse accuracy in some data set.
Kento Korenaga, Akito Monden, Zeynep Yücel
SNPD2
2019 On Preventing Symbolic Execution Attacks by Low Cost Obfuscation
abstract
While various software obfuscation techniques have been proposed to protect software, new types of threats keep emerging such as the symbolic execution attacks. Such attacks automatically analyze programs and are not accounted for by many of the existing obfuscation methods. Nevertheless, several methods against symbolic execution attacks exist such as linear obfuscation methods relying on Collatz conjuncture or obfuscation methods based on one-way hash functions. However, these methods bear several issues. Namely, linear obfuscation is weak against manual analysis due to its deterministic output. On the other hand, SHA-1 requires significant computational cost; and thus, it can be applied to only a limited number of targets. Therefore, in this research, we propose to employ a combination of several computationally cheap (arithmetic) obfuscating operations for preventing symbolic execution attacks. Through an experiment using angr and KLEE as symbolic execution tools, we demonstrate that obfuscation operation using array reference, bit rotation and XOR effectively prevents symbolic execution attacks at a low computational cost.
Toshiki Seto, Akito Monden, Zeynep Yücel, Yuichiro Kanzaki
SNPD2
2019 Prediction of Software Defects Using Automated Machine Learning
abstract
The effectiveness of defect prediction depends on modeling techniques as well as their parameter optimization, data preprocessing and ensemble development. This paper focuses on auto-sklearn, which is a recently-developed software library for automated machine learning, that can automatically select appropriate prediction models, hyperparameters and data preprocessing techniques for a given data set and develop their ensemble with optimized weights. In this paper we empirically evaluate the effectiveness of auto-sklearn in predicting the number of defects in software modules. In the experiment, we used software metrics of 20 OSS projects for cross-release defect prediction and compared auto-sklearn with random forest, decision tree and linear discriminant analysis by using Norm(Popt) as a performance measure. As a result, auto-sklearn showed similar prediction performance as random forest, which is one of the best prediction models for defect prediction in past studies. This indicates that auto-sklearn can obtain good prediction performance for defect prediction without any knowledge of machine learning techniques and models.
Kazuya Tanaka, Akito Monden, Zeynep Yücel
SNPD2
2019 On the relative value of data resampling approaches for software defect prediction
Kwabena Ebo Bennin, Jacky W. Keung, Akito Monden
Empir. Softw. Eng.3
2018 Kurtosis and Skewness Adjustment for Software Effort Estimation
abstract
To avoid software development project failure, accurate estimation of software development effort is necessary at the beginning of a software project. This paper proposes to adjust the kurtosis and the skewness of project feature variables for better fitting of software estimation models. The proposed method conducts logarithmic transformation of variables, then conducts the kurtosis and skewness transformation to make the variable distribution closer to the normal distribution. To empirically evaluate the effectiveness of the proposed method, we employed three industry data sets and linear regression models with three-fold cross validation. The result of the evaluation showed that the models with the proposed method were better in both the goodness of fit and the estimation accuracy in terms of MMRE compared to log-log regression.
Seiji Fukui, Akito Monden, Zeynep Yücel
APSEC2
2018 MAHAKIL: diversity based oversampling approach to alleviate the class imbalance issue in software defect prediction
abstract
This study presents MAHAKIL, a novel and efficient synthetic over-sampling approach for software defect datasets that is based on the chromosomal theory of inheritance. Exploiting this theory, MAHAKIL interprets two distinct sub-classes as parents and generates a new instance that inherits different traits from each parent and contributes to the diversity within the data distribution. We extensively compare MAHAKIL with five other sampling approaches using 20 releases of defect datasets from the PROMISE repository and five prediction models. Our experiments indicate that MAHAKIL improves the prediction performance for all the models and achieves better and more significant pf values than the other oversampling approaches, based on robust statistical tests.
Kwabena Ebo Bennin, Jacky W. Keung, Passakorn Phannachitta, Akito Monden, Solomon Mensah
ICSE4
2018 MAHAKIL: Diversity Based Oversampling Approach to Alleviate the Class Imbalance Issue in Software Defect Prediction
abstract
Highly imbalanced data typically make accurate predictions difficult. Unfortunately, software defect datasets tend to have fewer defective modules than non-defective modules. Synthetic oversampling approaches address this concern by creating new minority defective modules to balance the class distribution before a model is trained. Notwithstanding the successes achieved by these approaches, they mostly result in over-generalization (high rates of false alarms) and generate near-duplicated data instances (less diverse data). In this study, we introduce MAHAKIL, a novel and efficient synthetic oversampling approach for software defect datasets that is based on the chromosomal theory of inheritance. Exploiting this theory, MAHAKIL interprets two distinct sub-classes as parents and generates a new instance that inherits different traits from each parent and contributes to the diversity within the data distribution. We extensively compare MAHAKIL with SMOTE, Borderline-SMOTE, ADASYN, Random Oversampling and the No sampling approach using 20 releases of defect datasets from the PROMISE repository and five prediction models. Our experiments indicate that MAHAKIL improves the prediction performance for all the models and achieves better and more significant pf values than the other oversampling approaches, based on Brunner's statistical significance test and Cliff's effect sizes. Therefore, MAHAKIL is strongly recommended as an efficient alternative for defect prediction models built on highly imbalanced datasets.
Kwabena Ebo Bennin, Jacky W. Keung, Passakorn Phannachitta, Akito Monden, Solomon Mensah
IEEE Trans. Software Eng.4
2017 Impact of the Distribution Parameter of Data Sampling Approaches on Software Defect Prediction Models
abstract
Sampling methods are known to impact defect prediction performance. These sampling methods have configurable parameters that can significantly affect the prediction performance. It is however, impractical to assess the effect of all the possible different settings in the parameter space for all the several existing sampling methods. A constant and easy to tweak parameter present in all sampling methods is the distribution of the defective and non-defective modules in the dataset known as Pfp (% of fault-prone modules). In this paper, we investigate and assess the performance of defect prediction models where the Pfp parameter of sampling methods are tweaked. An empirical experiment and assessment of seven sampling methods on five prediction models over 20 releases of 10 static metric projects indicate that (1) Area Under the Receiver Operating Characteristics Curve (AUC) performance is not improved after tweaking the Pfp parameter, (2) pf (false alarms) performance degrades as the Pfp is increased. (3) a stable predictor is difficult to achieve across different Pfp rates. Hence, we conclude that the Pfp parameter setting can have a large impact on the performance (except AUC) of defect prediction models. We thus recommend researchers experiment with the Pfp parameter of the sampling method since the distribution of training datasets vary.
Kwabena Ebo Bennin, Jacky W. Keung, Akito Monden
APSEC3
2017 The Significant Effects of Data Sampling Approaches on Software Defect Prioritization and Classification
abstract
Context: Recent studies have shown that performance of defect prediction models can be affected when data sampling approaches are applied to imbalanced training data for building defect prediction models. However, the magnitude (degree and power) of the effect of these sampling methods on the classification and prioritization performances of defect prediction models is still unknown. Goal: To investigate the statistical and practical significance of using resampled data for constructing defect prediction models. Method: We examine the practical effects of six data sampling methods on performances of five defect prediction models. The prediction performances of the models trained on default datasets (no sampling method) are compared with that of the models trained on resampled datasets (application of sampling methods). To decide whether the performance changes are significant or not, robust statistical tests are performed and effect sizes computed. Twenty releases of ten open source projects extracted from the PROMISE repository are considered and evaluated using the AUC, pd, pf and G-mean performance measures. Results: There are statistical significant differences and practical effects on the classification performance (pd, pf and G-mean) between models trained on resampled datasets and those trained on the default datasets. However, sampling methods have no statistical and practical effects on defect prioritization performance (AUC) with small or no effect values obtained from the models trained on the resampled datasets. Conclusions: Existing sampling methods can properly set the threshold between buggy and clean samples, while they cannot improve the prediction of defect-proneness itself. Sampling methods are highly recommended for defect classification purposes when all faulty modules are to be considered for testing.
Kwabena Ebo Bennin, Jacky W. Keung, Akito Monden, Passakorn Phannachitta, Solomon Mensah
ESEM3
2017 Evaluating algorithmic thinking ability of primary schoolchildren who learn computer programming
abstract
In this research, a tool for evaluating algorithmic thinking ability of the primary schoolchildren was developed. This tool is based on the three categories of operations used to construct algorithms, namely, sequential operations, conditional branching operations, and iterative operations. Each question in the tool checks to see if the examinee understands the concept of one of the three categories. The tool was developed to evaluate the educational effect of programming education for middle to upper grade (third to sixth grade) primary schoolchildren. Since both Visual Programming Language (VPL) and Textual Programming Language (TPL) could be used, it was required that the tool could be used by both the group of children who use VPLs and the group of children who use TPLs. To make it possible, no programming language appeared in the questions in the tool. The teaching materials for the programming education were also developed in such a way that the three basic concepts of algorithm, namely, sequential processing, conditional branching, and repetitive processing, were clearly taught. The target VPL in this research was Scratch. The evaluation tool was conducted in a weekend class of programming education for primary schoolchildren, and the algorithmic thinking ability of the schoolchildren was analyzed.
Hidekuni Tsukamoto, Yasumasa Oomori, Hideo Nagumo, Yasuhiro Takemura, Akito Monden, Ken-ichi Matsumoto
FIE5
2017 A stability assessment of solution adaptation techniques for analogy-based software effort estimation
Passakorn Phannachitta, Jacky W. Keung, Akito Monden, Ken-ichi Matsumoto
Empir. Softw. Eng.3
2017 Benchmarking IT operations cost based on working time and unit cost
Masateru Tsunoda, Akito Monden, Ken-ichi Matsumoto, Sawako Ohiwa, Tomoki Oshino
Sci. Comput. Program.2
2016 Influence of outliers on analogy based software development effort estimation
abstract
In a software development project, project management is indispensable, and effort estimation is one of the important factors on the management. To improve estimation accuracy, outliers are often removed from dataset used for estimation. However, the influence of the outliers to the estimation accuracy is not clear. In this study, we added outliers to dataset experimentally, to analyze the influence. In the analysis, we changed the percentage of outliers, the extent of outliers, variable including outliers, and location of outliers on the dataset. After that, effort was estimated using the dataset. In the experiment, the influence of outliers was not very large, when they were included in the software size metric, the percentage of outliers was 10%, and the extent of outliers was 100%.
Kenichi Ono, Masateru Tsunoda, Akito Monden, Ken-ichi Matsumoto
ICIS3
2016 Analysis of information system operation cost based on working time and unit cost
abstract
Recently, information system operation becomes more important because of increasing size of information system and outsourcing the system operation. However, it is not easy for customers to judge the validity of the system operation cost. To provide information which helps the judgment, we analyzed factors which affect system operation cost. Working time of system operation service provider has the strong relationship to the cost. So, if customers know the working time, they estimate the cost properly. However, it is difficult for customers to know the working time generally. So, we assumed that customers estimate unit cost and working time, to speculate total operation cost roughly. To help the estimation, we analyzed factors affected working time and unit cost. The analysis results show that working time is settled based on the software size and the number of users, and the unit cost of the engineers increases when network range of the system is wide.
Masateru Tsunoda, Akito Monden, Ken-ichi Matsumoto, Sawako Ohiwa, Tomoki Oshino
ICIS2
2016 A fuzzy hashing technique for large scale software birthmarks
abstract
Software birthmarks have been proposed as a method for enabling the detection of programs that may have been stolen by measuring the similarity between the two programs. A birthmark is created from each program by extracting its native characteristics. The birthmarks of the programs can then be compared. However, because the extracted birthmarks contain a large amount of information, a large amount of time is needed when using them to compare large programs. This paper describes our work to reduce this comparison time. Achieving faster comparisons will enable the evaluation of large programs and simplify the use of birthmarks. Specifically, our method creates hashes from conventional birthmark information using fuzzy hashing, and then measures the similarity of the programs using the obtained hash values. Using the proposed method, we achieved a major speed increase over the conventional birthmark method with distinction rates of over 90%. On the other hand, because preservation performance decreased substantially, the similarity threshold value needed to be lowered when using the proposed method.
Takehiro Tsuzaki, Teruaki Yamamoto, Haruaki Tamada, Akito Monden
ICIS4
2016 Identifying recurring association rules in software defect prediction
abstract
Association rule mining discovers patterns of co-occurrences of attributes as association rules in a data set. The derived association rules are expected to be recurrent, that is, the patterns recur in future in other data sets. This paper defines the recurrence of a rule, and aims to find a criteria to distinguish between high recurrent rules and low recurrent ones using a data set for software defect prediction. An experiment with the Eclipse Mylyn defect data set showed that rules of lower than 30 transactions showed low recurrence. We also found that the lower bound of transactions to select high recurrence rules is dependent on the required precision of defect prediction.
Akito Monden, Yasutaka Kamei, Shuji Morisaki
ICIS2
2016 Filter-INC: Handling Effort-Inconsistency in Software Effort Estimation Datasets
abstract
Effort-inconsistency is a situation where historical software project data used for software effort estimation (SEE) are contaminated by many project cases with similar characteristics but are completed with significantly different amount of effort. Using these data for SEE generally produces inaccurate results; however, an effective technique for its handling is yet made to be available. This study approaches the problem differently from common solutions, where available techniques typically attempt to remove every project case they have detected as outliers. Instead, we hypothesize that data inconsistency is caused by only a few deviant project cases and any attempt to remove those other cases will result in reduced accuracy, largely due to loss of useful information and data diversity. Filter-INC (short for Filtering technique for handling effort-INConsistency in SEE datasets) implements the hypothesis to decide whether a project case being detected by any existing technique should be subject to removal. The evaluation is carried out by comparing the performance of 2 filtering techniques between before and after having Filter-INC applied. The results produced from 8 real-world datasets together with 3 machine-learning models, and evaluated by 4 performance measures show a significant accuracy improvement at the confident interval of 95%. Based on the results, we recommend our proposed hypothesis as an important instrument to design a data preprocessing technique for handling effort-inconsistency in SEE datasets, definitely an important step forward in preprocessing data for a more accurate SEE model.
Passakorn Phannachitta, Jacky W. Keung, Kwabena Ebo Bennin, Akito Monden, Ken-ichi Matsumoto
APSEC4
2016 Investigating the Effects of Balanced Training and Testing Datasets on Effort-Aware Fault Prediction Models
abstract
To prioritize software quality assurance efforts, faultprediction models have been proposed to distinguish faulty modules from clean modules. The performances of such models are often biased due to the skewness or class imbalance of the datasets considered. To improve the prediction performance of these models, sampling techniques have been employed to rebalance the distribution of fault-prone and non-fault-prone modules. The effect of these techniques have been evaluated in terms of accuracy/geometric mean/F1-measure in previous studies, however, these measures do not consider the effort needed to fixfaults. To empirically investigate the effect of sampling techniqueson the performance of software fault prediction models in a morerealistic setting, this study employs Norm(Popt), an effort-awaremeasure that considers the testing effort. We performed two setsof experiments aimed at (1) assessing the effects of samplingtechniques on effort-aware models and finding the appropriateclass distribution for training datasets (2) investigating the roleof balanced training and testing datasets on performance ofpredictive models. Of the four sampling techniques applied, the over-sampling techniques outperformed the under-samplingtechniques with Random Over-sampling performing best withrespect to the Norm (Popt) evaluation measure. Also, performanceof all the prediction models improved when sampling techniqueswere applied between the rates of (20-30)% on the trainingdatasets implying that a strictly balanced dataset (50% faultymodules and 50% clean modules) does not result in the bestperformance for effort-aware models. Our results also indicatethat performances of effort-aware models are significantly dependenton the proportions of the two types of the classes in thetesting dataset. Models trained on moderately balanced datasetsare more likely to withstand fluctuations in performance as theclass distribution in the testing data varies.
Kwabena Ebo Bennin, Jacky W. Keung, Akito Monden, Yasutaka Kamei, Naoyasu Ubayashi
COMPSAC3
2016 Textual vs. visual programming languages in programming education for primary schoolchildren
abstract
The purpose of this research is to compare textual programming languages and visual programming languages from the aspect of motivation. As a textual programming language, Processing programming language was used, and as visual programming languages, Scratch, a derivation of Scratch, Teaching materials offered by code.org, and LEGO Mindstorms EV3 were used. Teaching materials using the textual programming language, and those using the visual programming languages were developed separately. A trial experiment of programming education with the textual programming language was conducted to a cohort of seven primary schoolchildren. Trial experiments with the visual programming languages were conducted twice. In each of them, a cohort of eight primary schoolchildren participated. The motivation of the children was assessed using the questionnaire based on the ARCS (Attention, Relevance, Confidence, and Satisfaction) motivation model. The results with the visual programming languages suggested that the motivation scores of the children increased as the class progressed when visual programming languages were used. On the other hand, the results with Processing suggested that the variance of Satisfaction factor increased as the class progressed when textual programming languages were used, which further suggested that the Satisfaction scores of the children spread as the class progressed when textual programming languages were used.
Hidekuni Tsukamoto, Yasuhiro Takemura, Yasumasa Oomori, Isamu Ikeda, Hideo Nagumo, Akito Monden, Ken-ichi Matsumoto
FIE6
2016 Empirical Evaluation of Cross-Release Effort-Aware Defect Prediction Models
abstract
To prioritize quality assurance efforts, various fault prediction models have been proposed. However, the best performing fault prediction model is unknown due to three major drawbacks: (1) comparison of few fault prediction models considering small number of data sets, (2) use of evaluation measures that ignore testing efforts and (3) use of n-fold cross-validation instead of the more practical cross-release validation. To address these concerns, we conducted cross-release evaluation of 11 fault density prediction models using data sets collected from 2 releases of 25 open source software projects with an effort-aware performance measure known as Norm(Popt). Our result shows that, whilst M5 and K* had the best performances, they were greatly influenced by the percentage of faulty modules present and size of data set. Using Norm(Popt) produced an overall average performance of more than 50% across all the selected models clearly indicating the importance of considering testing efforts in building fault-prone prediction models.
Kwabena Ebo Bennin, Koji Toda, Yasutaka Kamei, Jacky W. Keung, Akito Monden, Naoyasu Ubayashi
QRS5
2016 Unsupervised Bug Report Categorization Using Clustering and Labeling Algorithm
abstract
Bug reports are one of the most crucial information sources for software engineering offering answers to many questions. Yet, getting these answers is not always easy; the information in bug reports is often implicit and some processes are required to extract the meaning of these reports. Most research in this area employ a supervised learning approach to classify bug reports so that required types of reports could be identified. However, this approach often requires an immense amount of time and effort, the resources that already too scarce in many projects. We aim to develop an automated framework that can categorize bug reports, according to their grammatical structure without the need for labeled data. Our framework categorizes bug reports according to their text similarity using topic modeling and a clustering algorithm. Each group of bug reports are labeled with our new clustering labeling algorithm specifically made for clusters in the topic space. Our framework is highly customizable with a modular approach and options to incorporate available background knowledge to improve its performance, while our cluster labeling approach make use of natural language process (NLP) chunking to create the representative labels. Our experiment results demonstrate that the performance of our unsupervised framework is comparable to a supervised learning one. We also show that our labeling process is capable of labeling each cluster with phrases that are representative for that cluster's characteristics. Our framework can be used to automatically categorize the incoming bug reports without any prior knowledge, as an automated labeling suggestion system or as a tool for obtaining knowledge about the structure of the bug report repository.
Nachai Limsettho, Hideaki Hata, Akito Monden, Ken-ichi Matsumoto
Int. J. Softw. Eng. Knowl. Eng.3
2015 Case consistency: a necessary data quality property for software engineering data sets
abstract
Data quality is an essential aspect in any empirical study, because the validity of models and/or analysis results derived from an empirical data is inherently influenced by its quality. In this empirical study, we focus on data consistency as a critical factor influencing the accuracy of prediction models in software engineering. We propose a software metric called Cases Inconsistency Level (CIL) for analyzing conflicts within software engineering data sets by leveraging probability statistics on project cases and counting the number of conflicting pairs. The result demonstrated that CIL is able to be used as a metric to identify either consistent data sets or inconsistent data sets, which are valuable for building robust prediction models. In addition to measuring the level of consistency, CIL is proved to be applicable to predict whether or not an effort model built from data set can achieve higher accuracy, an important indicator for empirical experiments in software engineering.
Passakorn Phannachitta, Akito Monden, Jacky W. Keung, Ken-ichi Matsumoto
EASE2
2015 Programming education for primary school children using a textual programming language
abstract
In this research, a Textual Programming Language (TPL) is used in programming education for primary schoolchildren because of the following reasons: (1) it is more practical to use the programming languages similar to the ones used for developing real applications, (2) typing statements could be easier for primary schoolchildren than generally thought, (3) there exist programming environments such as Processing that are easy to use and produce very attractive graphical outcomes. Teaching material for programming education with Processing was developed. In this teaching material, cartoons were used to explain difficult concepts. The learners who use this teaching material were supposed to draw some computational figures with chosen colors. Trial experiments of programming education using this teaching material was conducted to a cohort of seven primary schoolchildren (six 4th grade and one 5th grade children) in two consecutive weekend classes (one hour each). Since the authors' aim of this programming education was to create a sense of fun and excitement in the children and inculcate a desire to engage with computing, the motivation of the children was assessed using the questionnaire based on the ARCS (Attention, Relevance, Confidence, and Satisfaction) motivation model. The results were encouraging and suggested that TPLs could be used in programming education for primary schoolchildren.
Hidekuni Tsukamoto, Yasuhiro Takemura, Hideo Nagumo, Isamu Ikeda, Akito Monden, Ken-ichi Matsumoto
FIE5
2014 Prediction of the change of learners' motivation in programming education for non-computing majors
abstract
In the past, the authors had been analyzing motivation of the learners in programming education using the ARCS assessment metric. This metric had been used in the application experiment in 13 programming courses, and about 1,700 sets of data was collected. From these data, the learners' model, characteristics of the change of motivation, and ways of improving teaching materials had been clarified. However, these study results were obtained after the terms, when the programming courses were over, and thus did not contribute much to the ongoing programming education. For this reason, in this research, the methods for predicting the change of learners' motivation were studied so that the learners who may need support could be identified. The idea came from the experiment the authors conducted, in which the motivation of learners was analyzed by plotting the motivation scores of each factor in the ARCS model as a 3D graph. As a result, a decreasing tendency of motivation was observed when the distribution of the plot widened. After studying the tendency in detail, it was thought to be due to the influence of the variance of sub-level category scores. In the proposed method, the motivation of each learner is assessed in each lesson using the ARCS assessment metric. If variance of the motivation scores of a learner in a lesson is above a certain threshold value AND if mean of the scores has not decreased from the previous lesson, then the learner is identified as a candidate of learner who needs support at that lesson. In the application experiment, a programming course with 9 lessons was offered and 9 learners attended all the 9 lessons. In the experiment, 7 cases had been identified as the candidates of learners who need support, and out of those 7 cases, a decrease of motivation to less than average was observed in 5 cases.
Hidekuni Tsukamoto, Yasuhiro Takemura, Hideo Nagumo, Akito Monden, Ken-ichi Matsumoto
FIE4
2013 Patch Reviewer Recommendation in OSS Projects
abstract
In an Open Source Software (OSS) project, many developers contribute by submitting source code patches. To maintain the quality of the code, certain experienced developers review each patch before it can be applied or committed. Ideally, within a short amount of time after its submission, a patch is assigned to a reviewer and reviewed. In the real world, however, many large and active OSS projects evolve at a rapid pace and the core developers can get swamped with a large number of patches to review. Furthermore, since these core members may not always be available or may choose to leave the project, it can be challenging, at times, to find a good reviewer for a patch. In this paper, we propose a graph-based method to automatically recommend the most suitable reviewers for a patch. To evaluate our method, we conducted experiments to predict the developers who will apply new changes to the source code in the Eclipse project. Our method achieved an average recall of 0.84 for top-5 predictions and a recall of 0.94 for top-10 predictions.
John Boaz Lee, Akinori Ihara, Akito Monden, Ken-ichi Matsumoto
APSEC (2)3
2013 Improving Analogy-Based Software Cost Estimation through Probabilistic-Based Similarity Measures
abstract
The performance of software cost estimation based on analogy reasoning depends upon the measures that specifying the similarity between software projects. This paper empirically investigates the use of probabilistic-based distance functions to improve the similarity measurement. The probabilistic-based distance functions are considerably more robust, because they collect the implicit correlation between the occurrences of project feature attributes. This information gain enables the constructed estimation model to be more concise and comprehensible. The study compares 6 probabilistic-based distance functions against the commonly-used Euclidian distance. We empirically evaluate the implemented cost estimation model using 5 real-world datasets collected from the PROMISE repository. The result shows a significant improvement in terms of error reduction, that implies an estimation based on probabilistic-based distance functions achieve higher accuracy on average, and the peak performance significantly outperforms the Euclidian distance based on Wilcox on signed-rank test.
Passakorn Phannachitta, Jacky W. Keung, Akito Monden, Ken-ichi Matsumoto
APSEC (1)3
2013 The effects of teaching material remediation with ARCS-strategies for programming education
abstract
In this paper, a method for improving the teaching materials of programming education is introduced, and the evaluation of the effects of using the strategy is presented. By using this method, the teachers of programming education will be able to assess and improve their teaching materials irrespective of their knowledge and experience of their teaching materials already used. In this method, the teaching materials were improved based on the statistical analysis of the motivation of students. Specifically, the motivation of students was measured for each lower category of ARCS motivation model with the authors' original questionnaire. The lower category in a particular lesson that showed a statistically significant decrease from the previous lesson was identified, and the improvement strategies for the lower category were selected from the list of motivation strategies in the ARCS model. The teaching materials of programming education were then improved based on the strategy. In this research, five lower categories of particular lessons in a programming course were identified, and the teaching materials were improved. The improved teaching materials were used in the following programming course, and the effects of the improvements were seen in three lower categories out of the identified five lower categories.
Hidekuni Tsukamoto, Yasuhiro Takemura, Hideo Nagumo, Akito Monden, Ken-ichi Matsumoto
FIE4
2013 An Instruction Folding Method to Prevent Reverse Engineering in Java Platform
abstract
To improve tamper resistance of programs against illegal modification, this paper proposes instruction folding applicable to Java platform. In the proposed method, at first, similar methods are selected in a Java program. Next, these methods are merged into one method and diffs among these methods are stored in the program. Then, at runtime, when one of the merged methods is executed, diffs are restored by self-modification, which is realized by the Java instrumentation mechanism. The proposed method is resilient against tampering of folded method. Even if an adversary modifies the folded method, the program goes crash because the method is repeatedly modified at runtime.
Tetsuya Ohdo, Haruaki Tamada, Yuichiro Kanzaki, Akito Monden
SNPD4
2013 Assessing the Cost Effectiveness of Fault Prediction in Acceptance Testing
abstract
Until now, various techniques for predicting fault-prone modules have been proposed and evaluated in terms of their prediction performance; however, their actual contribution to business objectives such as quality improvement and cost reduction has rarely been assessed. This paper proposes using a simulation model of software testing to assess the cost effectiveness of test effort allocation strategies based on fault prediction results. The simulation model estimates the number of discoverable faults with respect to the given test resources, the resource allocation strategy, a set of modules to be tested, and the fault prediction results. In a case study applying fault prediction of a small system to acceptance testing in the telecommunication industry, results from our simulation model showed that the best strategy was to let the test effort be proportional to "the number of expected faults in a module × log(module size)." By using this strategy with our best fault prediction model, the test effort could be reduced by 25 percent while still detecting as many faults as were normally discovered in testing, although the company required about 6 percent of the test effort for metrics collection, data cleansing, and modeling. The simulation results also indicate that the lower bound of acceptable prediction accuracy is around 0.78 in terms of an effort-aware measure, Norm(Popt). The results indicate that reduction of the test effort can be achieved by fault prediction only if the appropriate test strategy is employed with high enough fault prediction accuracy. Based on these preliminary results, we expect further research to assess their general validity with larger systems.
Akito Monden, Takuma Hayashi, Shoji Shinoda, Kumiko Shirai, Junichi Yoshida, Mike Barker, Ken-ichi Matsumoto
IEEE Trans. Software Eng.1
2012 A Heuristic Rule Reduction Approach to Software Fault-proneness Prediction
abstract
Background: Association rules are more comprehensive and understandable than fault-prone module predictors (such as logistic regression model, random forest and support vector machine). One of the challenges is that there are usually too many similar rules to be extracted by the rule mining. Aim: This paper proposes a rule reduction technique that can eliminate complex (long) and/or similar rules without sacrificing the prediction performance as much as possible. Method: The notion of the method is to removing long and similar rules unless their confidence level as a heuristic is high enough than shorter rules. For example, it starts with selecting rules with shortest length (length=1), and then it continues through the 2nd shortest rules selection (length=2) based on the current confidence level, this process is repeated on the selection for longer rules until no rules are worth included. Result: An empirical experiment has been conducted with the Mylyn and Eclipse PDE datasets. The result of the Mylyn dataset showed the proposed method was able to reduce the number of rules from 1347 down to 13, while the delta of the prediction performance was only. 015 (from. 757 down to. 742) in terms of the F1 prediction criteria. In the experiment with Eclipsed PDE dataset, the proposed method reduced the number of rules from 398 to 12, while the prediction performance even improved (from. 426 to. 441.) Conclusion: The novel technique introduced resolves the rule explosion problem in association rule mining for software proneness prediction, which is significant and provides better understanding of the causes of faulty modules.
Akito Monden, Jacky W. Keung, Shuji Morisaki, Yasutaka Kamei, Ken-ichi Matsumoto
APSEC1
2012 Incorporating Expert Judgment into Regression Models of Software Effort Estimation
abstract
One of the common problems in building an effort estimation model is that not all the effort factors are suitable as predictor variables. As a supplement of missing information in estimation models, this paper explores the project manager's knowledge about the target project. We assume that the experts can judge the target project's productivity level based on his/her own expert knowledge about the project. We also assume that this judgment can be further improved, because using the expert's judgment solely could incur subjective perception. This paper proposes a regression model building/selection method to address this challenge. In the proposed method, a fit dataset for model building is divided into two or three subsets by project productivity, and an estimation model is built on each data subset. The expert judges the productivity level of the target project and selects one of the models to be used. In the experiment, we used three datasets to evaluate the produced effort estimation models. In the experiment, we adjusted the error rate of the judgment and analyzed the relationship between the error rate and the estimation accuracy. As a result, the judgment-incorporating models produced significantly higher estimation accuracy than the conventional linear regression model, where the expert's error rate is less than 37%.
Masateru Tsunoda, Akito Monden, Jacky W. Keung, Ken-ichi Matsumoto
APSEC2
2012 Handling categorical variables in effort estimation
abstract
Background: Accurate effort estimation is the basis of the software development project management. The linear regression model is one of the widely-used methods for the purpose. A dataset used to build a model often includes categorical variables denoting such as programming languages. Categorical variables are usually handled with two methods: the stratification and dummy variables. Those methods have a positive effect on accuracy but have shortcomings. The other handing method, the interaction and the hierarchical linear model (HLM), might be able to compensate for them. However, the two methods have not been examined in the research area. Aim: giving useful suggestions for handling categorical variables with the stratification, transforming dummy variables, the interaction, or HLM, when building an estimation model. Method: We built estimation models with the four handling methods on ISBSG, NASA, and Desharnais datasets, and compared accuracy of the methods with each other. Results: The most effective method was different for datasets, and the difference was statistically significant on both mean balanced relative error (MBRE) and mean magnitude of relative error (MMRE). The interaction and HLM were effective in a certain case. Conclusions: The stratification and transforming dummy variables should be tried at least, for obtaining an accurate model. In addition, we suggest that the application of the interaction and HLM should be considered when building the estimation model.
Masateru Tsunoda, Sousuke Amasaki, Akito Monden
ESEM3
2012 Evaluation of Non Functional Requirements in a Request for Proposal (RFP)
abstract
In the beginning of a contracted based software development project, the RFP is provided by a software user company and used as an initial system requirements specification to ask software developer companies to propose their technical plans to fulfill the requirements. In this process, it is very important to evaluate the quality of the RFP to make sure that basic user requirements are written enough. Especially, non-functional requirements (NFRs) are important since the system architecture greatly depends on the NFRs such as response time and security issues. This paper proposes a simple evaluation model of NFRs included in the RFP, mainly focusing on the user maintenance and operation issues. This model consists of NFR categories, NFR metrics, description level grading and weight to each NFR. As a case study, RFPs of 29 projects were evaluated by the proposed model. As a result, we confirmed that the model could identify poorly-written NFR aspects in the RFP, which need refinement before asking the developer company for a proposal.
Yasuhiro Saito, Akito Monden, Ken-ichi Matsumoto
IWSM/Mensura2
2012 An Ensemble Approach of Simple Regression Models to Cross-Project Fault Prediction
abstract
In software development, prediction of fault-prone modules is an important challenge for effective software testing. However, high prediction accuracy may not be achieved in cross-project prediction, since there is a large difference in distribution of predictor variables between the base project and the target project.@In this paper we propose an prediction technique called gan ensemble of simple regression modelsh to improve the prediction accuracy of cross-project prediction. The proposed method uses weighted sum of outputs of simple logistic regression models to improve the generalization ability of logistic models. To evaluate the performance of the proposed method, we conducted cross-project prediction using datasets of projects from NASA IV&V Facility Metrics Data Program. As a result, the proposed method outperformed conventional logistic regression models in terms of AUC of the Alberg diagram.
Satoshi Uchigaki, Shinji Uchida, Koji Toda, Akito Monden
SNPD4
2011 An Empirical Study of Fault Prediction with Code Clone Metrics
abstract
In this paper, we present a replicated study to predict fault-prone modules with code clone metrics to follow Baba's experiment. We empirically evaluated the performance of fault prediction models with clone metrics using 3 datasets from the Eclipse project and compared it to fault prediction without clone metrics. Contrary to the original Baba's experiment, we could not significantly support the effect of clone metrics, i.e., the result showed that F1-measure of fault prediction was not improved by adding clone metrics to the prediction model. To explain this result, this paper analyzed the relationship between clone metrics and fault density. The result suggested that clone metrics were effective in fault prediction for large modules but not for small modules.
Yasutaka Kamei, Akito Monden, Shinji Kawaguchi, Hidetake Uwano, Masataka Nagura, Ken-ichi Matsumoto, Naoyasu Ubayashi
IWSM/Mensura3
2011 A Model of Project Supervision for Process Correction and Improvement
abstract
Recently, software functional size becomes larger, and consequently, not only a software developer but also a software purchaser suffers considerable losses by software project failure. So avoiding project failure is also important for purchasers. Project supervision (monitoring and control) is expected for the purchaser to suppress risk of project failure. It is performed by sharing software metrics during the project for the purchaser to grasp the status of the project, and corrective actions are done based on analysis results of the metrics. Although there are some software measurement models, the models are not enough to describe how to confirm effects of project supervision. To acquire the effects certainly, the purchaser and the developer should quantitatively confirm whether the effects are acquired or not by project supervision. In addition, the models cannot represent corrective actions when symptoms of project failure are found. We propose the model for project supervision. The model explains planning, collecting data, transforming data, analyzing data, reaction toward found issues, and confirming effect of project supervision. With our model, project supervision can be described more rigorously.
Masateru Tsunoda, Akito Monden, Tomoko Matsumura, Ken-ichi Matsumoto
IWSM/Mensura2
2011 An Analysis of Cost-Overrun Projects Using Financial Data and Software Metrics
abstract
To clarify the characteristics of cost-overrun software projects, this paper focuses on the cost to sales ratio of software development, computed from financial information of a midsize software company in the embedded systems domain, and analyzes the correlation with outsourcing ratio as well as code reuse ratio and relative effort ratio per development phase. As a result, we found that the lower cost to sales ratio projects had the higher relative effort ratio in external design phase, which indicates that spending less effort in external design can cause decrease of profit. We also found that high outsourcing ratio projects had higher cost to sales ratio, and that projects having moderate code reuse ratio had lower and disperse cost to sales ratio, which suggests troubles in code reuse can damage the profit of a project.
Hidetake Uwano, Yasutaka Kamei, Akito Monden, Ken-ichi Matsumoto
IWSM/Mensura3
2011 An Empirical Study of Development Visualization for Procurement by in-Process Measurement during Integration and Testing
abstract
This study describes a new method of development visualization along with empirical evidence of its usefulness. Typically, development activities such as program design, programming, and unit testing are not disclosed to the procurement organization (project owner). However, during integration and testing, various issues require collaboration between the procurement organization and developers. When this occurs, it is important to make the development process visible. Recent reports indicate the usefulness for project management of various in-process project measurements which allow visualization of the formerly invisible software project progress [1–6]. Based on this background, the authors investigated a case study where in-process measurement during the integration and test phase helped to make development issues visible. In this study, data obtained from the integration and testing phase were compared to a development process model. This model was based on the author's experience, and provided a vivid picture of the development activity. By applying in-process measurements in collaboration during the integration test phase, the development activity was clearly visualized, and the procurement organization understood problems.
Yoshiki Mitani, Hiroyuki Yoshikawa, Seishiro Tsuruho, Akito Monden, Mike Barker, Ken-ichi Matsumoto
Int. J. Softw. Eng. Knowl. Eng.4
2010 Revisiting common bug prediction findings using effort-aware models
abstract
Bug prediction models are often used to help allocate software quality assurance efforts (e.g. testing and code reviews). Mende and Koschke have recently proposed bug prediction models that are effort-aware. These models factor in the effort needed to review or test code when evaluating the effectiveness of prediction models, leading to more realistic performance evaluations. In this paper, we revisit two common findings in the bug prediction literature: 1) Process metrics (e.g., change history) outperform product metrics (e.g., LOC), 2) Package-level predictions outperform file-level predictions. Through a case study on three projects from the Eclipse Foundation, we find that the first finding holds when effort is considered, while the second finding does not hold. These findings validate the practical significance of prior findings in the bug prediction literature and encourage their adoption in practice.
Yasutaka Kamei, Shinsuke Matsumoto, Akito Monden, Ken-ichi Matsumoto, Bram Adams, Ahmed E. Hassan
ICSM3
2009 An Empirical Study of the Feedback of the In-process Measurement in a Japanese Consortium-type Software Project
Yoshiki Mitani, Tomoko Matsumura, Katsuro Inoue, Mike Barker, Akito Monden, Ken-ichi Matsumoto
SEKE5
2008 An over-sampling method for analogy-based software effort estimation
abstract
This paper proposes a novel method to generate synthetic projectcases and add them to a fit dataset for the purpose of improving the performance of analogy-based software effort estimation. The proposed method extends conventional over-sampling method, which is a preprocessing procedure for n-group classification problems, which makes it suitable for any imbalanced dataset to be used in analogy-based system. We experimentally evaluated the effect of the over-sampling method to improve the performance of the analogy-based software effort estimation by using the Desharnais dataset. Results show significant improvement to the estimation accuracy by using our approach.
Yasutaka Kamei, Jacky W. Keung, Akito Monden, Ken-ichi Matsumoto
ESEM3
2008 A hybrid faulty module prediction using association rule mining and logistic regression analysis
abstract
This paper proposes a fault-prone module prediction method that combines association rule mining with logistic regression analysis. In the proposed method, we focus on three key measures of interestingness of an association rule (support, confidence and lift) to select useful rules for the prediction. If a module satisfies the premise (i.e. the condition in the antecedent part) of one of the selected rules, the module is classified by the rule as either fault-prone or not. Otherwise, the module is classified by the logistic model. We experimentally evaluated the prediction performance of the proposed method with different thresholds of each rule interestingness measure (support, confidence and lift) using a module set in the Eclipse project, and compared it with three well-known fault-proneness models (logistic regression model, linear discriminant model and classification tree). The result showed that the improvement of the F1-value of the proposed method was 0.163 at maximum compared to conventional models.
Yasutaka Kamei, Akito Monden, Shuji Morisaki, Ken-ichi Matsumoto
ESEM2
2008 Fit data selection for software effort estimation models
abstract
To construct a better multivariate regression model for software effort estimation, this paper proposes a method to select projects as a fit data from a given project data set based on estimation target's features. While regression models were often constructed from all available project data, this paper showed the necessity of fit data selection, and showed that the proposed method is one of the effective and systematic means to do the selection.
Koji Toda, Akito Monden, Ken-ichi Matsumoto
ESEM2
2008 Are good code reviewers also good at design review?
abstract
ESEM '08 : the Second ACM-IEEE international symposium on Empirical software engineering and measurement, October 09-10, 2008, Kaiserslautern, Germany
Hidetake Uwano, Akito Monden, Ken-ichi Matsumoto
ESEM2
2008 DRESREM 2: An Analysis System for Multi-document Software Review Using Reviewers' Eye Movements
abstract
To build high-reliability software in software development, software review is essential. Typically, software review requires documents from multiple phases such as requirements specification, design document and source code to reveal the inconsistencies among them and to ensure the traceability of deliverables. However, most previous studies on software review (reading) techniques focus on finding defects in a single document in their experiments. In this paper, we propose a multi-document review evaluation system, DRESREM2. This system records reviewers' eye movements and mouse/keyboard operations for analysis. We conducted eye gaze analysis of reviewers in design document review with multiple documents (including requirements specification, design document, etc.) to confirm the usefulness of the system. For the performance analysis, we recorded defect detection ratio, detection time per defect, and fixation ratio of eye movements on each document. As a result, reviewers who concentrated their eye movements on requirements specification found more defects in the design document. We believe this result is good evidence to encourage developers to read high-level documents when reviewing lowlevel documents.
Hidetake Uwano, Akito Monden, Ken-ichi Matsumoto
ICSEA2
2008 Analyzing Factors of Defect Correction Effort in a Multi-Vendor Information System Development
abstract
This paper describes an empirical study to reveal factors influencing defect correction effort in software development. In the study we collected various attributes (metrics) of defects found in a typical medium-scale, multi-vendor information system development project in Japan over a six-month period. We then statistically analyzed the relationship between the defects' attributes and the correction effort. The analysis confirmed the well-known principle “defects are the more expensive the later they are detected” by revealing that defects detected in the “system test” were 4.88 times more expensive than those detected in the “coding/unit test”. Another principle “defects are more expensive the longer they survive in software” was also confirmed by revealing that defects, which survived two or more development phases, were 4.44 times more expensive than those detected immediately. We also identified other factors, such as defect reproducibility, severity, and the cause of detection delay, that had a significant influence on the correction effort.
Tomoko Matsumura, Shuji Morisaki, Akito Monden, Ken-ichi Matsumoto
J. Comput. Inf. Syst.3
2007 The Effects of Over and Under Sampling on Fault-prone Module Detection
abstract
The goal of this paper is to improve the prediction performance of fault-prone module prediction models (fault-proneness models) by employing over/under sampling methods, which are preprocessing procedures for a fit dataset. The sampling methods are expected to improve prediction performance when the fit dataset is unbalanced, i.e. there exists a large difference between the number of fault-prone modules and not-fault-prone modules. So far, there has been no research reporting the effects of applying sampling methods to fault-proneness models. In this paper, we experimentally evaluated the effects of four sampling methods (random over sampling, synthetic minority over sampling, random under sampling and one-sided selection) applied to four fault-proneness models (linear discriminant analysis, logistic regression analysis, neural network and classification tree) by using two module sets of industry legacy software. All four sampling methods improved the prediction performance of the linear and logistic models, while neural network and classification tree models did not benefit from the sampling methods. The improvements of Fl-values in linear and logistic models were 0.078 at minimum, 0.224 at maximum and 0.121 at the mean.
Yasutaka Kamei, Akito Monden, Shinsuke Matsumoto, Takeshi Kakimoto, Ken-ichi Matsumoto
ESEM2
2007 Comparison of Outlier Detection Methods in Fault-proneness Models
abstract
In this paper, we experimentally evaluated the effect of outlier detection methods to improve the prediction performance of fault-proneness models. Detected outliers were removed from a fit dataset before building a model. In the experiment, we compared three outlier detection methods (Mahalanobis outlier analysis (MOA), local outlier factor method (LOFM) and rule based modeling (RBM)) each applied to three well-known fault-proneness models (linear discriminant analysis (LDA), logistic regression analysis (LRA) and classification tree (CT)). As a result, MOA and RBM improved F1-values of all models (0.04 at minimum, 0.17 at maximum and 0.10 at mean) while improvements by LOFM were relatively small (-0.01 at minimum, 0.04 at maximum and 0.01 at mean).
Shinsuke Matsumoto, Yasutaka Kamei, Akito Monden, Ken-ichi Matsumoto
ESEM3
2007 Is This Cost Estimate Reliable? - The Relationship between Homogeneity of Analogues and Estimation Reliability
abstract
Analogy-based cost estimation provides a useful and intuitive means to support decision making in software project management. It derives a cost estimate required for completing a project from information about similar past projects, namely the analogues. While on average this method provides a relatively accurate cost estimate there remains a possibility of large estimation errors. In this paper, we empirically tested the hypothesis that "using more homogeneous analogues produces a more reliable cost estimate" using a software engineering data repository established by the software engineering center (SEC), Information-technology Promotion Agency, Japan. This testing showed that low and high homogeneity projects had a large variation in estimation reliability. For instance, the difference was 22.9% (p = 0.021) in terms of percentage to get accurate estimates (better than Median of Magnitude of Relative Error).
Naoki Ohsugi, Akito Monden, Nahomi Kikuchi, Michael D. Barker, Masateru Tsunoda, Takeshi Kakimoto, Ken-ichi Matsumoto
ESEM2
2006 Analyzing individual performance of source code review using reviewers' eye movement
abstract
This paper proposes to use eye movements to characterize the performance of individuals in reviewing source code of computer programs. We first present an integrated environment to measure and record the eye movements of the code reviewers. Based on the fixation data, the environment computes the line number of the source code that the reviewer is currently looking at. The environment can also record and play back how the eyes moved during the review process. We conducted an experiment to analyze 30 review processes (6 programs, 5 subjects) using the environment. As a result, we have identified a particular pattern, called scan, in the subjects' eye movements. Quantitative analysis showed that reviewers who did not spend enough time for the scan tend to take more time for finding defects.
Hidetake Uwano, Masahide Nakamura, Akito Monden, Ken-ichi Matsumoto
ETRA3
2005 Recommendation of Software Technologies Based on Collaborative Filtering
abstract
Software engineers have to select some appropriate development technologies to use in the work; however, engineers sometimes cannot find the appropriate technologies because there are vast amount of options today. To solve this problem, we propose a software technology recommendation method based on collaborative filtering (CF). In the proposed method, at first, questionnaires are collected from concerned engineers about their technical interest. Next, similarities between an active engineer who gets recommendation and the other engineers are calculated according to the technical interests. Then, some similar engineers are selected for the active engineer. At last, some technologies are recommended which attract the similar engineers. An experimental evaluation showed that the proposed method can make accurate recommendations than that of a naive (non-CF) method.
Tomohiro Akinaga, Naoki Ohsugi, Masateru Tsunoda, Takeshi Kakimoto, Akito Monden, Ken-ichi Matsumoto
APSEC5
2005 Javawock: A Java Class Recommender System Based on Collaborative Filtering
Masateru Tsunoda, Takeshi Kakimoto, Naoki Ohsugi, Akito Monden, Ken-ichi Matsumoto
SEKE4
2005 Software Analysis by Code Clones in Open Source Software
Shinji Uchida, Akito Monden, Naoki Ohsugi, Toshihiro Kamiya, Ken-ichi Matsumoto, Hideo Kudo
J. Comput. Inf. Syst.2
2004 Effort Estimation Based on Collaborative Filtering
Naoki Ohsugi, Masateru Tsunoda, Akito Monden, Ken-ichi Matsumoto
PROFES3
2003 Exploiting Self-Modification Mechanism for Program Protection
abstract
In this paper, we present a new method to protect software against illegal acts of hacking. The key idea is to add a mechanism of self-modifying codes to the original program, so that the original program becomes hard to be analyzed. In the binary program obtained by the proposed method, the original code fragments we want to protect are camouflaged by dummy instructions. Then, the binary program autonomously restores the original code fragments within a certain period of execution, by replacing the dummy instructions with the original ones. Since the dummy instructions are completely different from the original ones, code hacking fails if the dummy instructions are read as they are. Moreover, the dummy instructions are scattered over the program, therefore, they are hard to be identified. As a result, the proposed method helps to construct highly invulnerable software without special hardware.
Yuichiro Kanzaki, Akito Monden, Masahide Nakamura, Ken-ichi Matsumoto
COMPSAC2
2002 A Recommendation System for Software Function Discovery
abstract
Since some application software provides users with too many functions, it is often difficult to find those that are useful. This paper proposes a recommendation system based on a collaborative filtering approach to let users discover useful functions at low cost for the purpose of improving productivity when using application software. The proposed system automatically collects histories of software function execution (usage histories) from many users through the Internet. Based on the collaborative filtering approach, collected histories are used for recommending a set of candidate functions that may be useful to the individual user. This paper illustrates conventional filtering algorithms and proposes a new algorithm suitable for recommendation of software functions. The result of an experiment with a prototype recommendation system showed that the average ndpm of our algorithm was smaller than that of conventional algorithms, and it also showed that the standard deviation of ndpm of our algorithm was smaller than that of conventional algorithms. Furthermore, while every conventional algorithm had a case whose recommendation was worse than the random algorithm, our algorithm did not.
Naoki Ohsugi, Akito Monden, Ken-ichi Matsumoto
APSEC2
2000 Button Selection for General GUIs Using Eye and Hand Together
abstract
This paper proposes an efficient technique for eye gaze interface suitable for the general GUI environments such as Microsoft Windows. Our technique uses an eye and a hand together: the eye for moving cursors onto the GUI button (move operation), and the hand for pushing the GUI button (push operation). We also propose the following two techniques to assist the move operation: (1) Automatic adjustment and (2) Manual adjustment. In the automatic adjustment, the cursor automatically moves to the closest GUI button when we push a mouse button. In the manual adjustment, we can move the cursor roughly by an eye, then move it a little more by the mouse onto the GUI button. In the experiment to evaluate our method, GUI button selection by manual adjustment showed better performance than the selection by a mouse even in the situation that has many small GUI buttons placed very closely each other on the GUI.
Masatake Yamamoto, Akito Monden, Ken-ichi Matsumoto, Katsuro Inoue, Koji Torii
Advanced Visual Interfaces2
2000 A Practical Method for Watermarking Java Programs
abstract
Java programs distributed through the Internet are now suffering from program theft. This is because Java programs can be easily decomposed into reusable class files and even decompiled into source code by program users. We propose a practical method that discourages program theft by embedding Java programs with a digital watermark. Embedding a program developer's copyright notation as a watermark in Java class files will ensure the legal ownership of class files. Our embedding method is indiscernible by program users, yet enables us to identify an illegal program that contains stolen class files. The result of the experiment to evaluate our method showed most of the watermarks (20 out of 23) embedded in class files survived two kinds of attacks that attempt to erase watermarks: an obfuscactor attack, and a decompile-recompile attack.
Akito Monden, Hajimu Iida, Ken-ichi Matsumoto, Koji Torii, Katsuro Inoue
COMPSAC1
2000 Modeling and Analysis of Software Aging Process
Akito Monden, Shin-ichi Sato, Ken-ichi Matsumoto, Katsuro Inoue
PROFES1
1995 Speaker adaptation fitting training data size and contents
Masahiro Tonomura, Tetsuo Kosaka, Shoichi Matsunaga, Akito Monden
EUROSPEECH4