Sousuke Amasaki

dblp:87/4954 · DBLP profile ↗
← Back
58ranked-venue papers
31as first author
17since 2021 · last 2025
0000-0001-8763-3457ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 57 · 30 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1
YearPublicationVenuePosition
2025 An Application of Program Mutations for Generating Negative Test Scripts Mimicking Human Errors on Web Applications
Tomoya Yamashita, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara
PROFES3
2024 A Quantitative Investigation of Trends in Confusing Variable Pairs Through Commits: Do Confusing Variable Pairs Survive?
abstract
Programmers can make the variables easy to understand by choosing meaningful names. However, even though individual names are meaningful, a collection of them might adversely affect the code readability when their names are highly similar, such as “bottomRight” vs. “bottomHeight,” and they may cause mixing up or misreading of variables. Such a pair is referred to as a “confusing variable pair.” This paper conducts a large-scale investigation to examine the change trends of confusing variable pairs over commits, and it reports the following findings. (1) The average number of confusing variable pairs appearing in a source file is 1.4 in Java and 1.3 in Python. (2) About 67–75% of confusing variable pairs survive through commits. (3) Confusing variable pairs tend to appear in a subset of source files within a project (the median rates in Java and Python projects: 26% and 13%) and disappear from fewer files (the median rates in Java and Python projects: 6% and 2%). (4) Although the change trends do not vary among projects, some outlier projects have significantly more confusing variable pairs.
Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara
EASE2
2024 Exploring Benefits of Bellwether Projects in Cross-Project IR-based Fault Localization
abstract
CONTEXT: Information retrieval-based bug localization (IRBL) is a promising approach for efficiently identifying buggy software modules in response to user bug reports. Supervised learning techniques were adopted to improve the bug localization performance, but they brought the cold-start problem due to insufficient training data. Recent studies focused on transfer learning techniques to utilize cross-project data. These techniques improved performance but left a question regarding better cross-project data selection. OBJECTIVE: To evaluate the effectiveness of bellwether projects, which are exemplary cross-project data better than the others for training a cross-project bug localization model. METHOD: With a bug localization method, the performance of cross-project bug localization was observed to find bellwether projects and to evaluate its effective usage for bug localization. RESULTS: One cross-project was dominantly better than the others. Also, it was often helpful to mix a bellwether project with the small available within-project data to improve the localization performance. CONCLUSION: A practical implication is to select cross-project data supported by cross-project bug localization on other projects. Mixing it with target project data is often beneficial at an early phase.
Sousuke Amasaki, Pattara Leelaprute, Hirohisa Aman, Tomoyuki Yokogawa
SEAA1
2024 Fault-Proneness of Python Programs Tested By Smelled Test Code
abstract
Software testing is one of the most crucial quality assurance activities, and test results are of great concern to software developers. However, the quality assurance of the test code (test case) itself also becomes critical because a poor-quality test case may fail to detect latent faults and give developers false comfort regarding the test result. A code smell threatening test code quality has been studied as “test smell.” This paper conducts an investigation of test smells in 775 Python open-source programs and reports the results of a quantitative analysis regarding whether test smells impact the fault-proneness of the product code under test. The analysis results show the following two findings. (1) When a test code has one of the reported ten kinds of test smells, the production code under test is more fault- prone than the others. (2) The fault-proneness of a production code tends to get higher when the corresponding test code has two or more different kinds of test smells-over 75% of test smell combinations showed such a trend of increasing the risk of being faulty production code.
Yuki Fushihara, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara
SEAA3
2024 Toward Individual Fairness Testing with Data Validity
abstract
Individual fairness testing (Ift) is a framework to find discriminatory instances within a given classifier. In this paper, we show our idea of a Ift framework, that integrates the notion of data validity, termed "Individual Fairness Testing with Data Validity (Ift-v)". We develop a solid foundation of Ift-v and demonstrate the feasibility of Ift-v. Our preliminary evaluation with Ift-v reveals the possibility that many of discriminatory instances detected by state-of-the-art Ift algorithms are considered invalid. These findings prompt a re-think of the current Ift framework, suggesting a transition from solely focusing on the discovery of discriminatory instances to the consideration of valid ones.
Takashi Kitamura 0001, Sousuke Amasaki, Jun Inoue 0001, Yoshinao Isobe, Takahisa Toda
ASE2
2024 A Multi - Aspect Evaluation of DL-based SQLi Attack Detection Models
abstract
CONTEXT: Web applications are exposed to malicious accesses through the Internet. SQL injection (SQLi) attacks are still a typical threat to web application providers. Although recent studies proposed deep learning-based SQLi attack detection models with high performance, those studies were not evaluated under the same conditions. Crucial aspects other than the predictive performance were also overlooked. OBJECTIVE: To evaluate SQLi attack detection models from multi-aspects related to its operation. METHOD: Three aspects, namely, predictive performance, detection speed, and operation costs, were applied to deep learning-based SQLi attack detection models. RESULTS: No DL-based model beaten a conventional machine learning-based model, Random Forests. CONCLUSION: Researchers must evaluate DL-based SQLi attack detection models with conventional ones with hyper-parameter tuning.
Pattara Leelaprute, Yuki Kase, Sousuke Amasaki, Hirohisa Aman, Tomoyuki Yokogawa
SERA3
2024 An Application of Program Slicing and CodeBERT to Distill Variables With Inappropriate Names
abstract
Variables are essential for handling objects and data in a program, and their names can provide helpful clues to understanding the program. Well-chosen names enhance code readability. On the other hand, ill-chosen names hinder the comprehension of the program or cause misunderstanding. Al-though a variable's name is worthy of attention, it is challenging to judge whether it is appropriate or not automatically. This paper proposes a method for checking variable names using the program slicing technique and CodeBERT to automate the assessment of variable names. Given a variable in a program, the proposed method extracts the program slice regarding the variable and masks the variable's name. Then, that method tries to predict the masked part using CodeBERT and assesses the original name's adequacy by comparing the predicted names with the original name. A case study shows that the proposed method may detect ill-chosen names with high accuracy (higher than 0.9).
Yahiro Mori, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara
SERA3
2023 A Trend Analysis of Test Smells in Python Test Code Over Commit History
abstract
Software testing is an essential activity for developing and maintaining high-quality software. Unit testing with test code (test cases) is a fundamental testing activity, and developers can test their production code whenever they create or modify the code. However, such quality assurance relies on the correctness of the test code. If a test code had a flaw, it would mislead the developers about the hidden faults and prevent early detection of the faults. This paper focuses on "test smells," which may cause test code flaws in Python programs, and analyzes their changing trends over commit history (code changes) toward better Python test code management. Through an empirical data analysis of 100 open-source projects, the paper reports the following findings: (1) a few kinds of test smells constitute the majority of smells detected in the studied projects, and (2) most kinds of smells tend to increase over commits, i.e., many test smells are likely to have remained in test code as technical debt.
Yuki Fushihara, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara
SEAA3
2023 Large-Scale Evaluation of Method-Level Bug Localization with FinerBench4BL
abstract
Bug localization is an important aspect of software maintenance because it can locate modules that need to be changed to fix a specific bug. Although method-level bug localization is helpful for developers, there are only a few tools and techniques for this task; moreover, there is no large-scale framework for their evaluation. In this paper, we present FinerBench4BL, an evaluation framework for method-level information retrieval-based bug localization techniques, and a comparative study using this framework. This framework was semi-automatically constructed from Bench4BL, a file-level bug localization evaluation framework, using a repository transformation approach. We converted the original file-level version repositories provided by Bench4BL into method-level repositories by repository transformation. Method-level data components such as oracle methods can also be automatically derived by applying the oracle generation approach via bug-commit linking in Bench4BL to the generated method repositories. Furthermore, we tailored existing file-level bug localization technique implementations at the method level. We created a framework for method-level evaluation by merging the generated dataset and implementations. The comparison results show that the method-level techniques decreased accuracy whereas improved debugging efficiency compared to file-level techniques.
Shizuka Tsumita, Shinpei Hayashi, Sousuke Amasaki
SANER3
2023 An automated detection of confusing variable pairs with highly similar compound names in Java and Python programs
abstract
Abstract Variable names represent a significant source of information regarding the source code, and a successful naming of variables is key to producing readable code. Programmers often use a compound variable name by concatenating two or more words to make it more informative and enhance the code readability. While each compound variable name is descriptive, a collection of them sometimes produces “confusing” variable pairs if their names are highly similar, e.g., “shippingHeight,” vs. “shippingWeight.” A confusing variable pair would adversely affect the code readability because it can cause a misreading or mix-up of variables during the programming or code review activities. Toward automated support for enhancing code readability, this paper conducts a large-scale investigation of compound variable names in Java and Python programs. The investigation collects 116,921,127 pairs of compound-named variables from 1,876 open-source Java projects and 106,943,523 pairs of such variables from 2,427 open-source Python projects. Then, this study analyzes those variable pairs from two perspectives of name similarity: string similarity and semantic similarity. Through an evaluation study with 30 human participants, the data analyses show that both string and semantic similarity can help detect confusing variable pairs in Java and Python programs. In order to distill confusing variable pairs automatically, support tools for detecting confusing variable pairs are also developed in this study.
Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara
Empir. Softw. Eng.2
2022 An Evaluation of Effort-Aware Fine-Grained Just-in-Time Defect Prediction Methods
abstract
CONTEXT: Software defect prediction (SDP) is an active research topic to support software quality assurance (SQA) activities. It was observed that unsupervised prediction models were often competitive with supervised ones at release-level and change-level defect prediction. Fine-grained just-in-time defect prediction focuses on defective files in a change, rather than the whole change. A recent study showed that the fine-grained just-in-time defect prediction was cost-effective in terms of effort-aware performance measures. Those studies did not explore the effectiveness of supervised and unsupervised models at that finer level in terms of effort-aware performance measures. OBJECTIVE: To examine the performance of supervised and unsupervised prediction models in the context of fine-grained defect prediction in terms of effort-aware performance measures. METHOD: Experiments with a time-sensitive approach were conducted to evaluate the predictive performance of supervised and unsupervised methods proposed in past studies. Datasets from OSS projects with manually validated defect links were employed from a past study. RESULTS: The use of manually validated links led to low-performance results. No clear difference among supervised and unsupervised methods was found while CBS+, a supervised method, was the best method in terms of F-measure. Even CBS+ did not achieve reasonable performance. A non-linear learning algorithm did not help the performance improvement. CONCLUSION: No clear preference among unsupervised and supervised methods. CBS+ was the best method on average. The predictive performance was still a challenge.
Sousuke Amasaki, Hirohisa Aman, Tomoyuki Yokogawa
SEAA1
2022 Have Java Production Methods Co-Evolved With Test Methods Properly?: A Fine-Grained Repository-Based Co-Evolution Analysis
abstract
Any source code of a software product (production code) is expected to be tested to ensure its correct behavior. Whenever a developer updates production code, the developer should also update or create the corresponding test code to check if the updated parts still work correctly. Such a desirable co-evolution relationship between production and test code forms a logical coupling. Although the logical coupling is detectable through an association analysis on the code repository such as Git, the detection granularity is coarse because the conventional repository is at the file level. For observing those logical couplings as precisely as possible, this paper utilizes the finer-grained, Java method-level repository (FinerGit). Then the paper proposes a metric measuring the extent to which a production method has co-evolved with test methods and conducts a case study using ten open-source projects. The results show that most Java methods (98% on average) have co-evolved with test methods, but some have not; The proposed metric helps detect those methods having the potential risk that the developers might not test adequately.
Tenma Kitai, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara
SEAA3
2022 Verifying Game Logic in Unreal Engine 5 Blueprint Visual Scripting System Using Model Checking
abstract
This paper examines modeling methods for applying model checking to game programs created with Unreal Engine 5 Blueprint scripting system (hereinafter UE5 Blueprint). UE5 Blueprint can visually describe game logic by combining various processing nodes, but as the size of the game grows, it becomes more difficult to find and fix bugs that prevent the game from progressing. In this paper, a formal verification technique, model checking, is used to automatically detect game logic bugs. We convert a game program created in UE5 Blueprint into an input model for the model-checker NuSMV to achieve verification by NuSMV. The proposed framework enables the automatic generation of models by formally defining the semantics of nodes. We also propose methods for data flow optimization and abstraction of variable domain for the purpose of reducing the number of states in the model. We applied the proposed method to a blueprint containing a typical flag management bug and confirmed that the bug was correctly detected by NuSMV. Furthermore, we show that the number of states can be significantly reduced by the optimization and abstraction.
Kazuki Wayama, Tomoyuki Yokogawa, Sousuke Amasaki, Hirohisa Aman, Kazutami Arimoto
ASE3
2022 An Evaluation of Cross-Project Defect Prediction Approaches on Cross-Personalized Defect Prediction
Sousuke Amasaki, Hirohisa Aman, Tomoyuki Yokogawa
PROFES1
2022 A comparative study on vectorization methods for non-functional requirements classification
Pattara Leelaprute, Sousuke Amasaki
Inf. Softw. Technol.2
2021 A Preliminary Evaluation of CPDP Approaches on Just-in-Time Software Defect Prediction
abstract
CONTEXT: Just-in-Time defect prediction is to specify the suspicious code commits that might make a product cause defects. Building JIT defect prediction models require a commit history and their fixed defect records. The shortage of commits of new projects motivated research of JIT cross-project defect prediction (CPDP). CPDP approaches proposed for component-level defect prediction were barely evaluated under JIT CPDP. OBJECTIVE: To explore the effects of CPDP approaches for component-level defect prediction where JIT CPDP is adopted. METHOD: A case study was conducted through two commit dataset suites provided in past studies for JIT defect prediction. JIT defect predictions with and without 21 CPDP approaches were compared regarding the classification performance using AUC. The CPDP approaches were also compared with each other. RESULTS: Most CPDP approaches changed the prediction performance of a baseline that simply combined all CP data. A few CPDP approaches could improve the prediction performance significantly. Not a few approaches worsened the performance significantly. The results based on the two suites could specify two CPDP approaches safer than the baseline. The results were inconsistent with a previous study. CONCLUSIONS: CPDP approaches for component-level might be effective for JIT CPDP. Further evaluations were needed to bring a firm conclusion.
Sousuke Amasaki, Hirohisa Aman, Tomoyuki Yokogawa
SEAA1
2021 Searching for Bellwether Developers for Cross-Personalized Defect Prediction
Sousuke Amasaki, Hirohisa Aman, Tomoyuki Yokogawa
PROFES1
2020 A Mahalanobis Distance-Based Integration of Suspicious Scores For Bug Localization
abstract
Once a software bug is reported, it is crucial to locate the source file causing the bug and fix it as soon as possible. To this end, there have been various studies for locating bugs with the version control system and the bug reports. AmaLgam+ is one of the most promising methods for localizing bugs. This method quantifies the degree to which a source file causes the bug from different five perspectives (metrics), and combines those values (suspicious scores) into a single integrated score of the source file. However, the method has a challenge regarding the computation time because it uses the genetic algorithm (GA) to combine the above five metrics. This paper proposes an application of the Mahalanobis distance to the suspicious score integration to overcome the above challenge. The proposed method considers the above five metrics to be a five-dimensional vector. It computes the Mahalanobis distance of the vector from the origin as an alternative integrated suspicious score. The empirical study using six open source software projects proves that the proposed method has almost the same bug localization accuracy as AmaLgam+ and can reduce the computation time by up to 98%: e.g., while AmaLgam+ took about 3.5 hours, the proposed method did it about 4 minutes.
Masanao Asato, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara
APSEC3
2020 A Comparative Study of Vectorization-Based Static Test Case Prioritization Methods
abstract
To enhance the efficiency of software testing, researchers have studied various test case prioritization (TCP) methods. A topic model-based TCP is one of the promising methods, which expresses test cases by topic vectors and prioritizes them in the order such that the set of already-prioritized test cases have the maximum dispersion in the vector space. However, the topic model is not the only option available for vectorizing test cases. Moreover, the distance metric in the vector space and the scheme to prioritize test cases (the way to find the test case that is the farthest from the set of already-prioritized ones) also have some available options. Because the combinations of the above options have not been well-discussed in the past, this paper conducts a comparative study of 36 TCP methods, which are the combinations of (1) three vectorization methods, (2) three distance metrics, and (3) four prioritization schemes (36=3x3x4). The empirical results show the following findings. The choice of the vectorization method has a significant impact on the testing efficiency: a promising option is Doc2Vec (PVDBoW). The combination with the distance metric may also be impactful: a useful combination is Doc2Vec (PV-DBoW) and Euclidean distance. The third aspect, i.e., the choice of the scheme to find the farthest test case, is not always influential.
Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara
SEAA2
2020 On the Effects of File-level Information on Method-level Bug Localization
abstract
CONTEXT: Automated bug localization is expected to help developers reducing effort and has been studied for years. One of the major approaches for this topic is Information Retrieval (IR) based bug localization. It takes a bug report and gives a ranking list of suspicious software modules based on the text similarity. Method-level bug localization is more desirable than file-level to reduce developers' cost to examine the ranking list. A challenge is that methods have little information inside. OBJECTIVE: To examine whether method-level bug localization can be improved with file-level information around a method. METHOD: An empirical evaluation was conducted with public data sets. A bug localization method BLUiR was used as a testbed. RESULTS: We found the class name of a method contributed to improving method-level bug localization. The comments outside the method often improved the performance but sometimes degraded it. The other elements hardly improved the performance. CONCLUSIONS: It is worth to add the class name for improving method-level bug localization. The comments are also worth trying.
Sousuke Amasaki, Hirohisa Aman, Tomoyuki Yokogawa
SEAA1
2020 Cross-version defect prediction: use historical data, cross-project data, or both?
Sousuke Amasaki
Empir. Softw. Eng.1
2020 Empirical study of abnormality in local variables and its application to fault-prone Java method analysis
abstract
Abstract Programmers are familiar with local variables, and in many cases, they can freely define the local variables they use. Thus, the properties of these variables are widely diverse, and this may cause variations in the quality of code. Although variables are named in accordance with coding conventions, the following matters have not received much attention from an empirical viewpoint: automatically deciding whether a local variable is “abnormal” and determining the harmful effect of an abnormal variable. This study focuses on the trends in the name, type, and scope of local variables, then proposes the use of the Mahalanobis distance to evaluate their abnormality. The empirical study entailed collecting local variables from eight open‐source software projects, and the paper reports the following findings: (a) the trend in the variation of the names of variables according to their type; (b) the majority of variables have short names with narrow scopes, where a name is often a word or an abbreviation thereof; (c) methods with an abnormal variable are approximately 1.4 times more likely to be fault prone than methods that contain only normal variables; (d) the proposed abnormality metric can be useful in a random forest‐based fault‐prone method analysis model.
Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara
J. Softw. Evol. Process.2
2019 Towards Better Effort Estimation with Cross-Project Defect Prediction Approaches
abstract
This research aims to tackle a data shift problem of software effort estimation. Cross project defect prediction approaches were found to be helpful for the same problem of software defect prediction. We examined the CPDP approaches and explored its applicability and adaptability for the problem of software effort estimation.
Sousuke Amasaki, Tomoyuki Yokogawa, Hirohisa Aman
EASE1
2019 A Comparative Study of Vectorization Methods on BugLocator
abstract
CONTEXT: Debugging is a labor-intensive and time-consuming activity. Automatic bug localization techniques have been proposed for reducing this effort. Among the techniques, Information retrieval (IR) based bug localization techniques take a bug report and give a rank list of source modules which are likely to cause the bug. Those techniques use a few variants of tf-idf vectorizations for bug reports and software modules though different vectorizations may give vastly different performances. OBJECTIVE: To explore the effects of vectorization methods on IR-based bug localization. METHOD: An empirical evaluation was conducted with 46 public data sets and 6 vectorization methods. BugLocator was used as a test bed. RESULTS: We found a vectorization used in BugLocator was one of the best. However, we found a better vectorization for representing software modules. CONCLUSIONS: It is worth to examine different vectorization methods for better IR-based bug localization because a preference for the methods can change as demonstrated in this study.
Sousuke Amasaki, Hirohisa Aman, Tomoyuki Yokogawa
SEAA1
2019 Exploring Preference of Chronological and Relevancy Filtering in Effort Estimation
Sousuke Amasaki
PROFES1
2018 A Doc2Vec-Based Assessment of Comments and Its Application to Change-Prone Method Analysis
abstract
Comments in a source program can be helpful artifacts for program comprehension. While many comments are useful documents embedded in source programs, there are also poorly-informative comments in the real world. In order to quantitatively assess the value of comments, this paper proposes applying the Doc2Vec model to comment evaluation. Doc2Vec is a useful model for vectorizing the content of a document. In this paper, a Java method is regarded as a document, and its content is expressed as a vector. Then, two vectors corresponding to different versions of a method are prepared-the original version and the comment-erased version-, and the vector similarity between these two versions are computed. If the erased comments provided richer information for the source code, the corresponding vector would have a larger change through the comment elimination. A method having poorly-informative comments may be low-quality and might require more code modifications. This paper analyzes the relationship between the value of comments in a method and the change-proneness, using the data collected from five popular open source software projects. The results show that a method having poorly-informative comments is likely to be change-prone, i.e., such a method could not survive unscathed after release.
Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara
APSEC2
2018 The Effects of Vectorization Methods on Non-Functional Requirements Classification
abstract
CONTEXT: Architecture and design of systems are sensitive to non-functional requirements (NFRs). Identifying NFRs and their categories at early phase is an essential task for project success. Automatic classification methods for that purpose have been studied for supporting requirement analysis. The past studies used simple vectorization methods and might miss semantics and interactions among words in requirements. OBJECTIVE: To examine whether different vectorization methods lead to differences in the classification performance of NFRs and their categories. METHOD: Comparative experiments were conducted with open data. Five vectorization methods including document embedding methods and four supervised classification methods were supplied. RESULTS: Some advanced methods could achieve better performance than traditional ones. The preference was dependent on classification methods. CONCLUSIONS: It is beneficial to consider using advanced methods for classifying non-functional requirements categories.
Sousuke Amasaki, Pattara Leelaprute
SEAA1
2018 Fault-Prone Java Method Analysis Focusing on Pair of Local Variables with Confusing Names
abstract
Giving a name to a local variable is usually a programmer's discretion. Since it depends on the programmer's preference and experience, there is a lot of individual variation which may cause a variability in the code quality such as the readability. While there have been studies on the naming of local variables in the past, a relationship of names among local variables within a method (function) has not been well-discussed. This paper focuses on a pair of local variables with similar, confusing names, e.g., "lineIndex" vs. "lineIndent." Since such local variables are confusable with each other, the presence of such a confusing pair may be related to the fault-proneness of the method. An empirical analysis for five major open source Java projects is conducted, and the following results are reported: (1) a method having a confusing variable pair is about 1.1 - 2.6 times more fault-prone than a method having only dissimilar (non-confusing) pairs; (2) the proposed metric of how confusing the local variables are is equivalent to or better than the conventional cyclomatic complexity in predicting fault-prone methods.
Keiichiro Tashima, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara
SEAA3
2017 An Evaluation of Selection Methods for Time-Aware Effort Estimation
abstract
CONTEXT: Several studies in effort estimation havefound that it can be effective to use only recent project data for building an effort estimation model. The generality of this timeaware approach has been explored across a variety of effort estimation model approaches, organizations and definitions of recency. However, other studies have shown that it is not alwayshelpful. A question arises: how can one tell whether the approachwould be effective for a given target project? OBJECTIVE: Toinvestigate a potential method to decide between selecting recentor all project data. METHOD: Using a single-company ISBSGdata set1 studied previously in similar research, we propose andevaluate a selection method. The method utilizes a variant ofcross-validation based on recent projects to make the decision.RESULTS: There are significant differences in the estimation accuracybetween using the proposed method and using the growingportfolio (always using all available data). The method could alsoselect the better approach on average. However, the differencein estimation accuracy between using the proposed method andalways using moving windows was not statistically significant.CONCLUSIONS: The selection method could select the betterapproach on average. The results contribute to developing amethod for suggesting a better approach for practitioners.
Sousuke Amasaki, Christopher J. Lokan
APSEC1
2017 On Software Productivity Analysis with Propensity Score Matching
abstract
[Context]: Software productivity analysis is an essential activity for software process improvement. It specifies critical factors to be resolved or accepted from project data. As the nature of project data is observational, not experimental, the project data involves bias that can cause spurious relationships among analyzed factors. Analysis methods based on linear regression suffer from the spurious relationships and sometimes lead an inappropriate causal relation. The propensity score is a solution for this problem but has rarely been used. [Objective]: To investigate what differences the use of propensity score brings to software productivity analysis in comparison to a conventional method. [Method]: We revisited classical software productivity analyses on ISBSG and Finnish datasets. The differences of critical factors between the propensity score and the linear regression were investigated. [Results]: Both analysis methods specified different critical factors on the two datasets. The specified factors were both reasonable to some extent, and further considerations are needed for the propensity score results. [Conclusions]: The use of propensity score can lead new possible factors to be tackled. Although the contradiction does not necessarily indicate a flaw of the linear regression, the results by the propensity score should also be noticed for better actions.
Masateru Tsunoda, Sousuke Amasaki
ESEM2
2017 Empirical Analysis of Words in Comments Written for Java Methods
abstract
This paper focuses on comments written in source programs. While comments can work for improving the readability of code, i.e., the quality of programs, there have also been concerns thatcomments can be added for complicated source code in order to compensate for a lack of readability. That is to say, well-written comments might be associated with problematic parts to be refactored. This paper collected Java methods (programs) from six popular open source products, and performs analyses on words which appear in their comments. Then, the paper shows that a method having a longer comments (more words)tends to be more change-prone and would be required more fixes after their releases.
Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara
SEAA2
2017 A Comparative Study on Linear Combination Rules for Ensemble Effort Estimation
abstract
Context: Software effort estimation is a critical factor for project success. A new approach called ensemble effort estimation gets popular because of its performance. While many combination rules have been proposed, they were only compared in a systematic literature review. Objective: To compare linear combination rules proposed in the past studies under the same condition based on empirical approach. Method: We conducted an experiment with 9 linear combination rules, 7 datasets, and 4 effort estimation models. Results: We found 6 out of 9 linear combination rules never underperformed its base learners. No linear combination rule was superior to the others. Conclusion: No definitive rule was found while some linear combination rules can give competitive or better estimates than its base learners.
Sousuke Amasaki
SEAA1
2017 An Application of the PageRank Algorithm to Commit Evaluation on Git Repository
abstract
Many empirical studies have reported notable theories or methods for evaluating or predicting code quality through analyses of code repositories. This paper has yet another point of view: it focuses on "commits" rather than source code. That is to say, this paper proposes to evaluate commits themselves. When an aim of a commit is to fix a bug, there can be another preceding commit which made a reason of the bug fixing. Those commits are linked by a bug fixing-based causal relationship. Then, commits can be modeled as a directed graph model of causal relationships. This paper applies Google's PageRank algorithm to the graph modelin order to evaluate commits' influences on the others. Through an empirical study with Git repositories of six open source projects, the following factors are showed to be noteworthy:(1) the number of added files at the commit,(2) the length of commit message,(3) the experience of committing author, and (4) the number of developers who have been involved in the modified files at the commit.
Sho Suzuki, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara
SEAA3
2017 A Survival Analysis of Source Files Modified by New Developers
Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara
PROFES2
2017 A Virtual Study of Moving Windows for Software Effort Estimation Using Finnish Datasets
Sousuke Amasaki, Christopher J. Lokan
PROFES1
2016 A replication study on the effects of weighted moving windows for software effort estimation
abstract
Context: Recent studies have shown that estimation accuracy can be affected by only using a window of recent projects as training data for building an effort estimation model. The idea has been extended for regression-based estimation by weighting projects differently according to their order within the window. This significantly improved the accuracy of estimation in a single-company dataset from the ISBSG repository.
Sousuke Amasaki, Christopher J. Lokan
EASE1
2016 On Applicability of Fixed-Size Moving Windows for ANN-Based Effort Estimation
abstract
BACKGROUND: Several studies in software effort estimation have found that it can be effective to use a window of recent projects as training data for building an effort estimation model. The generality of the windowing approach still remains uncertain across the variety of effort estimation approaches that are based on different theory. Recent studies have focused on the use of windows with effort estimation models based on a machine learning approach, which could make better estimates than conventional linear regression. OBJECTIVE: To investigate the effect of using a window on estimation accuracy with a machine learning-based method, Artificial Neural Networks (ANN). ANN was recently found as a popular and good performance method, and is based on a different theory from other Machine Learning-based methods used in past studies. METHOD: Using a single-company ISBSG dataset studied previously in similar research, we examine the effect of using a fixed-size windowing policy on the accuracy of estimates using ANN. RESULTS: There is a difference in the estimation accuracy between using a window and not using a window. Using windows of 50 to 120 projects reduced mean absolute errors by 5-7%. The effective range of window sizes was different from previous studies. CONCLUSIONS: Windowing significantly improves estimation accuracy with ANN. The results support past studies, in that the effective window sizes were different among estimation models. The results contribute to understanding characteristics of the windowing approach.
Sousuke Amasaki, Christopher J. Lokan
IWSM-Mensura1
2016 Towards Better Selection Between Moving Windows and Growing Portfolio
Sousuke Amasaki, Christopher J. Lokan
PROFES1
2015 Empirical Analysis of Change-Proneness in Methods Having Local Variables with Long Names and Comments
abstract
This paper focuses on the local variable names and comments that are major artifacts reflecting the programmer's preference. It conducts an empirical analysis on the usefulness of those artifacts in assessing the software quality from the perspective of change-proneness in Java methods developed in six popular open source software products. The empirical results show: (1) a method having a longer named local variable is more change-prone, and (2) the presence of comments inside the method body strengthens the suspicions to be modified after the release. The above artifacts are worthy to find methods which can survive unscathed after the release.
Hirohisa Aman, Sousuke Amasaki, Takashi Sasaki, Minoru Kawahara
ESEM2
2015 The Effects of Duration-Based Moving Windows with Estimation by Analogy
Sousuke Amasaki, Christopher J. Lokan
IWSM/Mensura1
2015 On the effectiveness of weighted moving windows: Experiment on linear regression based software effort estimation
abstract
Abstract In construction of an effort estimation model, it seems effective to use a window of training data so that the model is trained with only recent projects. Considering the chronological order of projects within the window, and weighting projects according to their order within the window, may also affect estimation accuracy. In this study, we examined the effects of weighted moving windows on effort estimation accuracy. We compared weighted and non‐weighted moving windows under the same experimental settings. We confirmed that weighting methods significantly improved estimation accuracy in larger windows, although the methods also significantly worsened accuracy in smaller windows. This result contributes to understanding properties of moving windows. Copyright © 2014 John Wiley & Sons, Ltd.
Sousuke Amasaki, Christopher J. Lokan
J. Softw. Evol. Process.1
2014 Empirical Analysis of Fault-Proneness in Methods by Focusing on their Comment Lines
abstract
This paper focuses on comments described in Java programs, and conducts an empirical analysis about relationships between comments and fault-proneness in the programs. The types of comments analyzed in this paper are comments described inside a method body (inner comments), and comments followed by a method declaration (documentation comments). Although both of them play important roles in the program comprehension, they seem to be described in different purposes, The inner comments are often added to present tips about code fragments, while the documentation comments usually work as a programmer's manual. In the field of code refactoring, well-written inner comments are said to be related to "code smell" since they may cover a lack of readability in a complicated code fragment. This paper analyzes the associations of comments with the code quality from the aspect of fault-proneness, with using four popular open source products. The empirical results show that a method having inner comments tends to be 1.8 - 3.0 times likely to be faulty. The key contribution of this work is to reveal the usefulness of inner comments to point at faulty methods.
Hirohisa Aman, Sousuke Amasaki, Takashi Sasaki, Minoru Kawahara
APSEC (2)2
2014 Empirical analysis of comments and fault-proneness in methods: can comments point to faulty methods?
abstract
[Context]
Hirohisa Aman, Takashi Sasaki, Sousuke Amasaki, Minoru Kawahara
ESEM3
2014 The Effects of Gradual Weighting on Duration-Based Moving Windows for Software Effort Estimation
Sousuke Amasaki, Christopher J. Lokan
PROFES1
2013 How to treat timing information for software effort estimation?
abstract
Software development effort estimation is an essential aspect of software project management. An effort estimation model expresses relationships between effort and factors such as organizational and project features (e.g. software functional size, and the programming language used in a project). However, software development practices and tools change over time, to environmental changes. This can affect some relationships assumed in an effort estimation model. A moving windows method (a method for treating the timing information of projects), has thus been proposed for estimation models. The moving windows method uses data from a fixed number of the most recent projects data for model construction. However, it is not clear that moving windows is the best way to handle the timing information in an estimation model. The goal of our research is to determine how best to treat timing information in constructing effort estimation models. To achieve the goal, we compared six different methods (moving windows, dummy variable of moving windows, dummy variables of equal bins, dummy variables of year, year predictor, and serial number) for treating timing data, in terms of estimation accuracy. In the experiment, we use three software development project datasets. We found that moving windows is best when the number of projects included in the dataset is not small, and dummy variable of moving windows is the best when the number is small.
Masateru Tsunoda, Sousuke Amasaki, Christopher J. Lokan
ICSSP2
2013 The Effects of Variable Selection Methods on Linear Regression-Based Effort Estimation Models
abstract
Stepwise regression has often been used for variable selection of effort estimation models. However it has been criticized for inappropriate selection, and another method is recommended. We thus examined the effects of Lasso, which is one of such variable selection methods. An experiment with datasets from PROMISE repository revealed that Lasso-based selection stably selected better variables than stepwise in predictive performance. We thus concluded Lasso-based selection is preferable to stepwise regression.
Sousuke Amasaki, Tomoyuki Yokogawa
IWSM/Mensura1
2013 Consistency Verification of UML Diagrams Based on Process Bisimulation
abstract
In the development of a software system using UML, consistency between state machine diagrams and sequence diagrams is crucial. This study proposes a verification method for the consistency of a sequence diagram and state machine diagrams. The proposed method represents state machine diagrams and a sequence diagram as processes, and can verify the consistency by checking weak simulation of the processes. We confirms the method could detect inconsistency with an example.
Tomoyuki Yokogawa, Sousuke Amasaki, Keisuke Okazaki, Yoichiro Sato, Kazutami Arimoto, Hisashi Miyazaki
PRDC2
2013 The Evaluation of Weighted Moving Windows for Software Effort Estimation
Sousuke Amasaki, Christopher J. Lokan
PROFES1
2012 Handling categorical variables in effort estimation
abstract
Background: Accurate effort estimation is the basis of the software development project management. The linear regression model is one of the widely-used methods for the purpose. A dataset used to build a model often includes categorical variables denoting such as programming languages. Categorical variables are usually handled with two methods: the stratification and dummy variables. Those methods have a positive effect on accuracy but have shortcomings. The other handing method, the interaction and the hierarchical linear model (HLM), might be able to compensate for them. However, the two methods have not been examined in the research area. Aim: giving useful suggestions for handling categorical variables with the stratification, transforming dummy variables, the interaction, or HLM, when building an estimation model. Method: We built estimation models with the four handling methods on ISBSG, NASA, and Desharnais datasets, and compared accuracy of the methods with each other. Results: The most effective method was different for datasets, and the difference was statistically significant on both mean balanced relative error (MBRE) and mean magnitude of relative error (MMRE). The interaction and HLM were effective in a certain case. Conclusions: The stratification and transforming dummy variables should be tried at least, for obtaining an accurate model. In addition, we suggest that the application of the interaction and HLM should be considered when building the estimation model.
Masateru Tsunoda, Sousuke Amasaki, Akito Monden
ESEM2
2012 The Effects of Moving Windows to Software Estimation: Comparative Study on Linear Regression and Estimation by Analogy
abstract
BACKGROUND: Models for estimating software development effort are constructed using a set of historical data for training. In construction, it seems effective to use a window of training data that consists of only recently finished projects. Two previous studies evaluated the use of a window with linear regression (LR) and estimation by analogy (EbA). However, these studies were based on different datasets and thus their findings could not be compared directly. OBJECTIVE: This study investigates the effect of using a window on estimation accuracy with EbA and LR. The difference between the results with the two modeling techniques was also investigated. METHOD: We compared the effectiveness of using a window with both LR and EbA, with the same experiment settings and data. RESULTS: There is a difference in accuracy between using a window and not using a window, with both software estimation methods. However, the effect of the use of a window is weaker with EbA than with LR. CONCLUSIONS: Windowing is effective with EbA and LR. However, the degree of effectiveness is weaker with EbA than with LR. The results contribute to understand how the windowing approach interrelates with software estimation models.
Sousuke Amasaki, Christopher J. Lokan
IWSM/Mensura1
2012 A Study on Predictive Performance of Regression-Based Effort Estimation Models Using Base Functional Components
Sousuke Amasaki, Tomoyuki Yokogawa
PROFES1
2011 Performance Evaluation of Windowing Approach on Effort Estimation by Analogy
abstract
Background: In effort estimation model construction, it seems effective to window training project data so that only recently finished projects are used. This is because old projects might be less representative of an organization. The past study demonstrated windowing approach works with linear regression, which is one of global models. However, this approach has not been examined with local models. Local models use subset of historical data for model construction and thus windowing approach may influence on its performance more weakly. Aim: To investigate whether windowing approach works with local models. Method: We replicated the past study with EbA. Maxwell and CSC datasets were used for an experiment. Results: Windowing approach improved predictive performance. Although the difference was insignificant in any window size, the result indicated using windowing approach has positive effect on average. Conclusions: This result contributes to understand where windowing approach works well.
Sousuke Amasaki, Yohei Takahara, Tomoyuki Yokogawa
IWSM/Mensura1
2011 A Study on Performance Inconsistency between Estimation by Analogy and Linear Regression
Sousuke Amasaki
SEKE1
2010 Productivity Reanalysis for Unbalanced Datasets with Mixed-Effects Models
Sousuke Amasaki
PROFES1
2006 Characterization of Runaway Software Projects Using Association Rule Mining
Sousuke Amasaki, Yasuhiro Hamano, Osamu Mizuno, Tohru Kikuno
PROFES1
2005 A New Challenge for Applying Time Series Metrics Data to Software Quality Estimation
Sousuke Amasaki, Takashi Yoshitomi, Osamu Mizuno, Yasunari Takagi, Tohru Kikuno
Softw. Qual. J.1
2003 A Bayesian Belief Network for Assessing the Likelihood of Fault Content
abstract
To predict software quality, we must consider various factors because software development consists of various activities, which the software reliability growth model (SRGM) does not consider. In this paper, we propose a model to predict the final quality of a software product by using the Bayesian belief network (BBN) model. By using the BBN, we can construct a prediction model that focuses on the structure of the software development process explicitly representing complex relationships between metrics, and handling uncertain metrics, such as residual faults in the software products. In order to evaluate the constructed model, we perform an empirical experiment based on the metrics data collected from development projects in a certain company. As a result of the empirical evaluation, we confirm that the proposed model can predict the amount of residual faults that the SRGM cannot handle.
Sousuke Amasaki, Yasunari Takagi, Osamu Mizuno, Tohru Kikuno
ISSRE1
2002 Statistical Analysis of Time Series Data on the Number of Faults Detected by Statistical Analysis of Time Series Data on the Number of Faults Detected by Software Testing
abstract
According to a progress of the software process improvement, the time series data on the number of faults detected by the software testing are collected extensively. In this paper, we perform statistical analyses of relationships between the time series data and the field quality of software products. At first, we apply the rank correlation coefficient /spl tau/ to the time series data collected from actual software testing in a certain company, and classify these data into four types of trends: strict increasing, almost increasing, almost decreasing, and strict decreasing. We then investigate, for each type of trend, the field quality of software products developed by the corresponding software projects. As a result of statistical analyses, we showed that software projects having trend of almost or strict decreasing in the number of faults detected by the software testing could produce the software products with high quality.
Sousuke Amasaki, Takashi Yoshitomi, Osamu Mizuno, Tohru Kikuno, Yasunari Takagi
Asian Test Symposium1