VLDB 2026 Research / reviewers in the wild / expert
Minoru Kawahara
dblp:39/1005
· DBLP profile ↗
28ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0002-3542-5039ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 22 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Application of Program Mutations for Generating Negative Test Scripts Mimicking Human Errors on Web Applications
Tomoya Yamashita, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara |
PROFES | 5 |
| 2025 | ViFT: Visual field transformer for visual field testing via deep reinforcement learningabstractVisual field testing (perimetry) quantifies a patient's visual field sensitivity to diagnosis and follow-up on their visual impairments. Visual field testing would require the patients to concentrate on the test for a long time. However, a longer testing time makes patients more exhausted and leads to a decrease in testing accuracy. Thus, it is helpful to develop a well-designed strategy to finish the testing more quickly while maintaining high accuracy. This paper proposes the visual field transformer (ViFT) for visual field testing with deep reinforcement learning. This study contributes to the following four: (1) ViFT can fully control the visual field testing process. (2) ViFT learns the relationships of visual field locations without any pre-defined information. (3) ViFT learning process can consider the patient perception uncertainty. (4) ViFT achieves the same or higher accuracy than the other strategies, and half as test time as the other strategies. Our experiments demonstrate the ViFT efficiency on the 24-2 test pattern compared with other strategies. Shozo Saeki, Minoru Kawahara, Hirohisa Aman |
Medical Image Anal. | 2 |
| 2024 | A Quantitative Investigation of Trends in Confusing Variable Pairs Through Commits: Do Confusing Variable Pairs Survive?abstractProgrammers can make the variables easy to understand by choosing meaningful names. However, even though individual names are meaningful, a collection of them might adversely affect the code readability when their names are highly similar, such as “bottomRight” vs. “bottomHeight,” and they may cause mixing up or misreading of variables. Such a pair is referred to as a “confusing variable pair.” This paper conducts a large-scale investigation to examine the change trends of confusing variable pairs over commits, and it reports the following findings. (1) The average number of confusing variable pairs appearing in a source file is 1.4 in Java and 1.3 in Python. (2) About 67–75% of confusing variable pairs survive through commits. (3) Confusing variable pairs tend to appear in a subset of source files within a project (the median rates in Java and Python projects: 26% and 13%) and disappear from fewer files (the median rates in Java and Python projects: 6% and 2%). (4) Although the change trends do not vary among projects, some outlier projects have significantly more confusing variable pairs. Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara |
EASE | 4 |
| 2024 | Fault-Proneness of Python Programs Tested By Smelled Test CodeabstractSoftware testing is one of the most crucial quality assurance activities, and test results are of great concern to software developers. However, the quality assurance of the test code (test case) itself also becomes critical because a poor-quality test case may fail to detect latent faults and give developers false comfort regarding the test result. A code smell threatening test code quality has been studied as “test smell.” This paper conducts an investigation of test smells in 775 Python open-source programs and reports the results of a quantitative analysis regarding whether test smells impact the fault-proneness of the product code under test. The analysis results show the following two findings. (1) When a test code has one of the reported ten kinds of test smells, the production code under test is more fault- prone than the others. (2) The fault-proneness of a production code tends to get higher when the corresponding test code has two or more different kinds of test smells-over 75% of test smell combinations showed such a trend of increasing the risk of being faulty production code. Yuki Fushihara, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara |
SEAA | 5 |
| 2024 | An Application of Program Slicing and CodeBERT to Distill Variables With Inappropriate NamesabstractVariables are essential for handling objects and data in a program, and their names can provide helpful clues to understanding the program. Well-chosen names enhance code readability. On the other hand, ill-chosen names hinder the comprehension of the program or cause misunderstanding. Al-though a variable's name is worthy of attention, it is challenging to judge whether it is appropriate or not automatically. This paper proposes a method for checking variable names using the program slicing technique and CodeBERT to automate the assessment of variable names. Given a variable in a program, the proposed method extracts the program slice regarding the variable and masks the variable's name. Then, that method tries to predict the masked part using CodeBERT and assesses the original name's adequacy by comparing the predicted names with the original name. A case study shows that the proposed method may detect ill-chosen names with high accuracy (higher than 0.9). Yahiro Mori, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara |
SERA | 5 |
| 2023 | A Trend Analysis of Test Smells in Python Test Code Over Commit HistoryabstractSoftware testing is an essential activity for developing and maintaining high-quality software. Unit testing with test code (test cases) is a fundamental testing activity, and developers can test their production code whenever they create or modify the code. However, such quality assurance relies on the correctness of the test code. If a test code had a flaw, it would mislead the developers about the hidden faults and prevent early detection of the faults. This paper focuses on "test smells," which may cause test code flaws in Python programs, and analyzes their changing trends over commit history (code changes) toward better Python test code management. Through an empirical data analysis of 100 open-source projects, the paper reports the following findings: (1) a few kinds of test smells constitute the majority of smells detected in the studied projects, and (2) most kinds of smells tend to increase over commits, i.e., many test smells are likely to have remained in test code as technical debt. Yuki Fushihara, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara |
SEAA | 5 |
| 2023 | Multi proxy anchor family loss for several types of gradients
Shozo Saeki, Minoru Kawahara, Hirohisa Aman |
Comput. Vis. Image Underst. | 2 |
| 2023 | An automated detection of confusing variable pairs with highly similar compound names in Java and Python programsabstractAbstract Variable names represent a significant source of information regarding the source code, and a successful naming of variables is key to producing readable code. Programmers often use a compound variable name by concatenating two or more words to make it more informative and enhance the code readability. While each compound variable name is descriptive, a collection of them sometimes produces “confusing” variable pairs if their names are highly similar, e.g., “shippingHeight,” vs. “shippingWeight.” A confusing variable pair would adversely affect the code readability because it can cause a misreading or mix-up of variables during the programming or code review activities. Toward automated support for enhancing code readability, this paper conducts a large-scale investigation of compound variable names in Java and Python programs. The investigation collects 116,921,127 pairs of compound-named variables from 1,876 open-source Java projects and 106,943,523 pairs of such variables from 2,427 open-source Python projects. Then, this study analyzes those variable pairs from two perspectives of name similarity: string similarity and semantic similarity. Through an evaluation study with 30 human participants, the data analyses show that both string and semantic similarity can help detect confusing variable pairs in Java and Python programs. In order to distill confusing variable pairs automatically, support tools for detecting confusing variable pairs are also developed in this study. Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara |
Empir. Softw. Eng. | 4 |
| 2022 | Have Java Production Methods Co-Evolved With Test Methods Properly?: A Fine-Grained Repository-Based Co-Evolution AnalysisabstractAny source code of a software product (production code) is expected to be tested to ensure its correct behavior. Whenever a developer updates production code, the developer should also update or create the corresponding test code to check if the updated parts still work correctly. Such a desirable co-evolution relationship between production and test code forms a logical coupling. Although the logical coupling is detectable through an association analysis on the code repository such as Git, the detection granularity is coarse because the conventional repository is at the file level. For observing those logical couplings as precisely as possible, this paper utilizes the finer-grained, Java method-level repository (FinerGit). Then the paper proposes a metric measuring the extent to which a production method has co-evolved with test methods and conducts a case study using ten open-source projects. The results show that most Java methods (98% on average) have co-evolved with test methods, but some have not; The proposed metric helps detect those methods having the potential risk that the developers might not test adequately. Tenma Kitai, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara |
SEAA | 5 |
| 2021 | Significance of Emphasized Features for Good Representation on Deep Metric LearningabstractDeep metric learning (DML) learns the mapping, which maps into embedding space in which similar data is near and dissimilar data is far. Most DML frameworks apply L2 normalization to feature vectors, and these feature vectors are non-sparse [6], [10], [30]. In this paper, we propose to apply L1 regularization loss to feature vectors. Proposed regularization emphasizes important features and restraints unimportant features on L2 normalized features. L1 regularization can combine with general DML losses because L1 regularization only regularizes feature vectors. In this paper, we finally propose SparseSoft-Triple loss, which is a combination of SoftTriple loss and L1 regularization. We demonstrate the effectiveness of the proposed SparseSoftTriple loss on some data sets for image retrieval tasks and fine-grained images. Shozo Saeki, Minoru Kawahara, Hirohisa Aman |
SERA | 2 |
| 2020 | A Mahalanobis Distance-Based Integration of Suspicious Scores For Bug LocalizationabstractOnce a software bug is reported, it is crucial to locate the source file causing the bug and fix it as soon as possible. To this end, there have been various studies for locating bugs with the version control system and the bug reports. AmaLgam+ is one of the most promising methods for localizing bugs. This method quantifies the degree to which a source file causes the bug from different five perspectives (metrics), and combines those values (suspicious scores) into a single integrated score of the source file. However, the method has a challenge regarding the computation time because it uses the genetic algorithm (GA) to combine the above five metrics. This paper proposes an application of the Mahalanobis distance to the suspicious score integration to overcome the above challenge. The proposed method considers the above five metrics to be a five-dimensional vector. It computes the Mahalanobis distance of the vector from the origin as an alternative integrated suspicious score. The empirical study using six open source software projects proves that the proposed method has almost the same bug localization accuracy as AmaLgam+ and can reduce the computation time by up to 98%: e.g., while AmaLgam+ took about 3.5 hours, the proposed method did it about 4 minutes. Masanao Asato, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara |
APSEC | 5 |
| 2020 | A Comparative Study of Vectorization-Based Static Test Case Prioritization MethodsabstractTo enhance the efficiency of software testing, researchers have studied various test case prioritization (TCP) methods. A topic model-based TCP is one of the promising methods, which expresses test cases by topic vectors and prioritizes them in the order such that the set of already-prioritized test cases have the maximum dispersion in the vector space. However, the topic model is not the only option available for vectorizing test cases. Moreover, the distance metric in the vector space and the scheme to prioritize test cases (the way to find the test case that is the farthest from the set of already-prioritized ones) also have some available options. Because the combinations of the above options have not been well-discussed in the past, this paper conducts a comparative study of 36 TCP methods, which are the combinations of (1) three vectorization methods, (2) three distance metrics, and (3) four prioritization schemes (36=3x3x4). The empirical results show the following findings. The choice of the vectorization method has a significant impact on the testing efficiency: a promising option is Doc2Vec (PVDBoW). The combination with the distance metric may also be impactful: a useful combination is Doc2Vec (PV-DBoW) and Euclidean distance. The third aspect, i.e., the choice of the scheme to find the farthest test case, is not always influential. Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara |
SEAA | 4 |
| 2020 | Empirical study of abnormality in local variables and its application to fault-prone Java method analysisabstractAbstract Programmers are familiar with local variables, and in many cases, they can freely define the local variables they use. Thus, the properties of these variables are widely diverse, and this may cause variations in the quality of code. Although variables are named in accordance with coding conventions, the following matters have not received much attention from an empirical viewpoint: automatically deciding whether a local variable is “abnormal” and determining the harmful effect of an abnormal variable. This study focuses on the trends in the name, type, and scope of local variables, then proposes the use of the Mahalanobis distance to evaluate their abnormality. The empirical study entailed collecting local variables from eight open‐source software projects, and the paper reports the following findings: (a) the trend in the variation of the names of variables according to their type; (b) the majority of variables have short names with narrow scopes, where a name is often a word or an abbreviation thereof; (c) methods with an abnormal variable are approximately 1.4 times more likely to be fault prone than methods that contain only normal variables; (d) the proposed abnormality metric can be useful in a random forest‐based fault‐prone method analysis model. Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara |
J. Softw. Evol. Process. | 4 |
| 2018 | A Doc2Vec-Based Assessment of Comments and Its Application to Change-Prone Method AnalysisabstractComments in a source program can be helpful artifacts for program comprehension. While many comments are useful documents embedded in source programs, there are also poorly-informative comments in the real world. In order to quantitatively assess the value of comments, this paper proposes applying the Doc2Vec model to comment evaluation. Doc2Vec is a useful model for vectorizing the content of a document. In this paper, a Java method is regarded as a document, and its content is expressed as a vector. Then, two vectors corresponding to different versions of a method are prepared-the original version and the comment-erased version-, and the vector similarity between these two versions are computed. If the erased comments provided richer information for the source code, the corresponding vector would have a larger change through the comment elimination. A method having poorly-informative comments may be low-quality and might require more code modifications. This paper analyzes the relationship between the value of comments in a method and the change-proneness, using the data collected from five popular open source software projects. The results show that a method having poorly-informative comments is likely to be change-prone, i.e., such a method could not survive unscathed after release. Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara |
APSEC | 4 |
| 2018 | Empirical Analysis of Coding Standard Violation Focusing on Its Coverage and ImportanceabstractToward an effective utilization of static code analysis tools, this paper investigates which violation is familiar with more programmers and widely appears in source files (having a high coverage), and which violation is really related to bugfixes (having a high importance), for six popular OSS projects. The results show: 1) the familiar violations tend to differ among projects, and only 25 violations are common to all surveyed projects; 2) the trends of their importance vary from project to project. Aji Ery Burhandenny, Hirohisa Aman, Minoru Kawahara |
COMPSAC (1) | 3 |
| 2018 | Fault-Prone Java Method Analysis Focusing on Pair of Local Variables with Confusing NamesabstractGiving a name to a local variable is usually a programmer's discretion. Since it depends on the programmer's preference and experience, there is a lot of individual variation which may cause a variability in the code quality such as the readability. While there have been studies on the naming of local variables in the past, a relationship of names among local variables within a method (function) has not been well-discussed. This paper focuses on a pair of local variables with similar, confusing names, e.g., "lineIndex" vs. "lineIndent." Since such local variables are confusable with each other, the presence of such a confusing pair may be related to the fault-proneness of the method. An empirical analysis for five major open source Java projects is conducted, and the following results are reported: (1) a method having a confusing variable pair is about 1.1 - 2.6 times more fault-prone than a method having only dissimilar (non-confusing) pairs; (2) the proposed metric of how confusing the local variables are is equivalent to or better than the conventional cyclomatic complexity in predicting fault-prone methods. Keiichiro Tashima, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara |
SEAA | 5 |
| 2017 | Empirical Analysis of Words in Comments Written for Java MethodsabstractThis paper focuses on comments written in source programs. While comments can work for improving the readability of code, i.e., the quality of programs, there have also been concerns thatcomments can be added for complicated source code in order to compensate for a lack of readability. That is to say, well-written comments might be associated with problematic parts to be refactored. This paper collected Java methods (programs) from six popular open source products, and performs analyses on words which appear in their comments. Then, the paper shows that a method having a longer comments (more words)tends to be more change-prone and would be required more fixes after their releases. Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara |
SEAA | 4 |
| 2017 | An Application of the PageRank Algorithm to Commit Evaluation on Git RepositoryabstractMany empirical studies have reported notable theories or methods for evaluating or predicting code quality through analyses of code repositories. This paper has yet another point of view: it focuses on "commits" rather than source code. That is to say, this paper proposes to evaluate commits themselves. When an aim of a commit is to fix a bug, there can be another preceding commit which made a reason of the bug fixing. Those commits are linked by a bug fixing-based causal relationship. Then, commits can be modeled as a directed graph model of causal relationships. This paper applies Google's PageRank algorithm to the graph modelin order to evaluate commits' influences on the others. Through an empirical study with Git repositories of six open source projects, the following factors are showed to be noteworthy:(1) the number of added files at the commit,(2) the length of commit message,(3) the experience of committing author, and (4) the number of developers who have been involved in the modified files at the commit. Sho Suzuki, Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara |
SEAA | 5 |
| 2017 | A Survival Analysis of Source Files Modified by New Developers
Hirohisa Aman, Sousuke Amasaki, Tomoyuki Yokogawa, Minoru Kawahara |
PROFES | 4 |
| 2016 | Examination of Coding Violations Focusing on Their Change Patterns over ReleasesabstractCode review is an essential activity to ensure the quality of code being developed, and there have been static code checkers for aiding an effective code review. However, such tools have not been actively utilized in the world of programmers due to a lot of coding violations (warning) produced by tools and their false-positiveness. In order to analyze the automatically pointed violations and the actual attentions which programmers paid to those violations, this paper proposes a novel metric— the Index of Programmers' Attention (IPA)—and conducts an empirical study focusing on the change patterns of violations over the releases of popular seven open source software products, under two research questions (RQs): (RQ1) What kind of coding violations are related to the parts that many programmers tend to improve? and what kind of coding violations are likely to be disregarded?; (RQ2) How can we reduce the meaningless violations for programmers by omitting disregarded coding violations?The empirical results showed the following findings: (1) important violations (having high IPA values) may vary from project to project; (2) there are some unimportant violations common to different projects, but they are a minority of automatically detected violations (about 12%). Therefore, while many violations may be made by a code checker, most of them are likely to be worthy in improving the code quality, and it is ineffective to reduce the violations by eliminating such unimportant violations. Aji Ery Burhandenny, Hirohisa Aman, Minoru Kawahara |
APSEC | 3 |
| 2016 | Application of Mahalanobis-Taguchi Method and 0-1 Programming Method to Cost-Effective Regression TestingabstractTo enhance the cost effectiveness of regression testing, this paper proposes a method for prioritizing test cases. In general, a test case can be evaluated from various different points of view, therefore whether it is worth it to re-run should be discussed using multi criteria. This paper shows that the Mahalanobis-Taguchi (MT) method is a useful way to successfully integrate different evaluations of a test case. Moreover, this paper proposes to use the 0-1 programming method together with the MT method in order to take into account not only the priority of a test case but also its cost to run. The empirical study with 300 test cases for an industrial software system shows that the combination of the MT method and the 0-1 programming method is more cost-effective than other conventional methods. Hirohisa Aman, Yuta Tanaka, Takashi Nakano, Hideto Ogasawara, Minoru Kawahara |
SEAA | 5 |
| 2015 | Empirical Analysis of Change-Proneness in Methods Having Local Variables with Long Names and CommentsabstractThis paper focuses on the local variable names and comments that are major artifacts reflecting the programmer's preference. It conducts an empirical analysis on the usefulness of those artifacts in assessing the software quality from the perspective of change-proneness in Java methods developed in six popular open source software products. The empirical results show: (1) a method having a longer named local variable is more change-prone, and (2) the presence of comments inside the method body strengthens the suspicions to be modified after the release. The above artifacts are worthy to find methods which can survive unscathed after the release. Hirohisa Aman, Sousuke Amasaki, Takashi Sasaki, Minoru Kawahara |
ESEM | 4 |
| 2014 | Empirical Analysis of Fault-Proneness in Methods by Focusing on their Comment LinesabstractThis paper focuses on comments described in Java programs, and conducts an empirical analysis about relationships between comments and fault-proneness in the programs. The types of comments analyzed in this paper are comments described inside a method body (inner comments), and comments followed by a method declaration (documentation comments). Although both of them play important roles in the program comprehension, they seem to be described in different purposes, The inner comments are often added to present tips about code fragments, while the documentation comments usually work as a programmer's manual. In the field of code refactoring, well-written inner comments are said to be related to "code smell" since they may cover a lack of readability in a complicated code fragment. This paper analyzes the associations of comments with the code quality from the aspect of fault-proneness, with using four popular open source products. The empirical results show that a method having inner comments tends to be 1.8 - 3.0 times likely to be faulty. The key contribution of this work is to reveal the usefulness of inner comments to point at faulty methods. Hirohisa Aman, Sousuke Amasaki, Takashi Sasaki, Minoru Kawahara |
APSEC (2) | 4 |
| 2014 | Empirical analysis of comments and fault-proneness in methods: can comments point to faulty methods?abstract[Context] Hirohisa Aman, Takashi Sasaki, Sousuke Amasaki, Minoru Kawahara |
ESEM | 4 |
| 2008 | Encoding for secure computations in distributed interactive real-time applications
Keiichi Endo, Minoru Kawahara, Yutaka Takahashi 0001 |
Comput. Commun. | 2 |
| 2003 | Parallel Vector Computing Technique for Discovering Communities on the Very Large Scale Web Graph
Kikuko Kawase, Minoru Kawahara, Takeshi Iwashita, Hiroyuki Kawano, Masanori Kawazawa |
DaWaK | 2 |
| 2000 | Mondou: Information Navigator with Visual Interface
Hiroyuki Kawano, Minoru Kawahara |
DaWaK | 2 |
| 1999 | Mining Association Algorithm Based on ROC Convex Hull Method in Bibliographic Navigation System
Minoru Kawahara, Hiroyuki Kawano |
Discovery Science | 1 |