Katsuro Inoue

dblp:33/779 · DBLP profile ↗
← Back
131ranked-venue papers
12as first author
13since 2021 · last 2025
0000-0001-5424-0614ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 123 · 12 first-author · 13 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-authorHuman-computer interaction and ubiquitous computing · 5Artificial intelligence and machine learning · 3Systems, architecture and hardware · 1Computer networks · 1Security and privacy · 1
YearPublicationVenuePosition
2025 BiFuzz: A Two-Stage Fuzzing Tool for Open-World Video Games
abstract
Open-world video games present a broader search space than other video games, posing challenges for test automation. Fuzzing, which generates new inputs by mutating an initial input, is commonly used to uncover issues. In this study, we proposed BiFuzz, a two-stage fuzzer designed for automated testing of open-world video games, and investigated its effectiveness. The results revealed that BiFuzz mutated the overall strategy of gameplay and test cases, including actual movement paths, step by step. Consequently, BiFuzz can detect character stuck issues. The tool and its video are at https://github.com/ Yusaku-Kato/BiFuzz and https://www.youtube.com/watch? $\mathbf{v}=$ VOrHfnLJSbk. Index Terms-open-world video game, fuzzing
Yusaku Kato, Norihiro Yoshida, Erina Makihara, Katsuro Inoue
APSEC4
2025 A Dataset of Software Bill of Materials for Evaluating SBOM Consumption Tools
abstract
A Software Bill of Materials (SBOM) is becoming an essential tool for effective software dependency management. An SBOM is a list of components used in software, including details such as component names, versions, and licenses. Using SBOMs, developers can quickly identify software components and assess whether their software depends on vulnerable libraries. Numerous tools support software dependency management through SBOMs, which can be broadly categorized into two types: tools that generate SBOMs and tools that utilize SBOMs. A substantial collection of accurate SBOMs is required to evaluate tools that utilize SBOMs. However, there is no publicly available dataset specifically designed for this purpose, and research on SBOM consumption tools remains limited. In this paper, we present a dataset of SBOMs to address this gap. The dataset we constructed comprises 46 SBOMs generated from real-world Java projects, with plans to expand it to include a broader range of projects across various programming languages. Accurate and well-structured SBOMs enable researchers to evaluate the functionality of SBOM consumption tools and identify potential issues. We collected 3,271 Java projects from GitHub and generated SBOMs for 798 of them using Maven with an open-source SBOM generation tool. These SBOMs were refined through both automatic and manual corrections to ensure accuracy, currently resulting in 46 SBOMs that comply with the SPDX Lite profile, which defines minimal requirements tailored to practical workflows in industries. This process also revealed issues with the SBOM generation tools themselves. The dataset is publicly available on Zenodo (DOI: 10.5281/zenodo.14233414).
Rio Kishimoto, Tetsuya Kanda 0001, Yuki Manabe 0001, Katsuro Inoue, Yoshiki Higo
MSR4
2025 A Retrospective on Developing Code Clone Detector CCFinder and Its Impact
abstract
In this retrospective article of our TSE paper “CCFinder: A Multilinguistic Token-Based Code Clone Detection System for Large Scale Source Code” (Kamiya et al., 2002), we revisit the reasons why we became deeply involved in code clone research, and explore what has driven its frequent citation in many studies. Furthermore, we reflect on why not only our own lab, but also numerous researchers and tool developers have pursued code clone research and the development of related tools.
Toshihiro Kamiya, Shinji Kusumoto, Katsuro Inoue
IEEE Trans. Software Eng.3
2024 An Empirical Analysis of Git Commit Logs for Potential Inconsistency in Code Clones
abstract
Code clones are code snippets that are identical or similar to other snippets within the same or different files. They are often created through copy-and-paste practices and modified during development and maintenance activities. Since a pair of code clones, known as a clone pair, has a possible logical coupling between them, it is expected that changes to each snippet are made simultaneously (co-changed) and consistently. There is extensive research on code clones, including studies related to the co-change of clones; however, detailed analysis of commit logs for code clone pairs has been limited. In this paper, we investigate the commit logs of code snippets from clone pairs, using the git-log command to extract changes to cloned code snippets. We analyzed 45 repositories owned by the Apache Software Foundation on GitHub and addressed three research questions regarding commit frequency, co-change ratio, and commit patterns. Our findings indicate that (1) on average, clone snippets are changed infrequently, typically only two or three times throughout their lifetime, (2) the ratio of co-changes is about half of all clone changes, with 10-20% of co-changed commits being concerning (potentially inconsistent), and (3) 35–65 % of all clone pairs being classified as concerning clone pairs (potentially inconsistent clone pairs). These results suggest the need for a consistent management system through the commit timeline of clones.
Reishi Yokomori, Katsuro Inoue
SCAM2
2024 SBOM Challenges for Developers: From Analysis of Stack Overflow Questions
abstract
Current software development takes advantage of many external libraries, but it entails security and copyright risks. While the use of the Software Bill of Materials (SBOM) has been encouraged to cope with this problem, its adoption is still insufficient. In this research, we analyzed the challenges that developers faced in practicing SBOM use by examining questions about SBOM utilization on Stack Overflow, a Q&A site for developers. As a result, we found that (1) the proportion of resolved questions about SBOM use is 15.0% which is extremely low, (2) the number of new questions has increased steadily from 2020 to 2023, and (3) SBOM users have three major challenges on SBOM tools.
Wataru Otoda, Tetsuya Kanda 0001, Yuki Manabe 0001, Katsuro Inoue, Yoshiki Higo
SERA4
2024 Osmy: A Tool for Periodic Software Vulnerability Assessment and File Integrity Verification using SPDX Documents
abstract
Libraries have become integral to modern software development, yet their management often falls short, resulting in issues such as delayed responses to vulnerabilities. To address these issues, the use of a Software Bill of Materials (SBOM) is recommended. Despite the recommendation, there is a lack of tools supporting software management using SBOM. In this paper, we present “Osmy”, a tool designed to facilitate effective software management using SBOM in the SPDX format-one of the major SBOM formats. Osmy is designed to simplify and streamline SBOM-based software management for end users. It automates vulnerability assessment and file integrity verification, operating periodically to ensure continuous protection. Users receive timely notification of any identified issues, ensuring a proactive approach to software security. Osmy is available at https://github.com/higolab/Osmy.
Rio Kishimoto, Tetsuya Kanda 0001, Yuki Manabe 0001, Katsuro Inoue, Yoshiki Higo
SANER4
2024 Evaluating the effectiveness of size-limited execution trace with near-omniscient debugging
Kazumasa Shimari, Takashi Ishio, Tetsuya Kanda 0001, Katsuro Inoue
Sci. Comput. Program.4
2023 Investigating the Generalizability of Deep Learning-based Clone Detectors
abstract
The generalizability of Deep Learning (DL) models is a significant challenge, as poor generalizability indicates that the model has overfitted to the training data and is not able to generalize to new data. Despite numerous DL-based clone detectors emerging in recent years, their generalizability has not been thoroughly assessed. This study investigates the generalizability of three DL-based clone detectors (CCLearner, ASTNN, and CodeBERT) by comparing their detection accuracy on different training and testing clone benchmarks. The results show that all three clone detectors do not generalize well to new data and there is a strong relationship between clone types and generalizability for CCLearner and ASTNN.
Eunjong Choi, Norihiro Fuke, Yuji Fujiwara, Norihiro Yoshida, Katsuro Inoue
ICPC5
2022 Comparison of Developer's Work Efficiency between Different Editors
abstract
In this paper, as an example of comparing developer’s work efficiency between different editors, we propose a method to collect and compare self-evaluations and quantitative evaluations of developer’s work efficiency for each editor. We also practiced the proposed method on Visual Studio Code and Eclipse, and confirmed the applicability of the method.
Sentaro Onizuka, Tetsuya Kanda 0001, Katsuro Inoue
APSEC3
2022 Selecting Test Cases based on Similarity of Runtime Information: A Case Study of an Industrial Simulator
abstract
Regression testing is required to check the changes in behavior whenever developers make any changes to a software system. The cost of regression testing is a major problem because developers have to frequently update dependent components to minimize security risks and potential bugs. In this paper, we report a current practice in a company that maintains an industrial simulator as a critical component of their business. The simulator automatically records all the users’ requests and the simulation results in storage. The feature provides a huge number of test cases for regression testing to developers; however, their time budget for testing is limited (i.e., at most one night). Hence, the developers need to select a small number of test cases to confirm both the simulation result and execution performance are unaffected by an update of a dependent component. In other words, the test cases should achieve high coverage while keeping diversity of execution time. To solve the problem, we have developed a clustering-based method to select test cases, using the similarity of execution traces produced by them. The developers have used the method for a half year; they recognize that the method is better than the previous rule-based method used in the company.
Kazumasa Shimari, Masahiro Tanaka, Takashi Ishio, Makoto Matsushita, Katsuro Inoue, Satoru Takanezawa
ICSME5
2022 didiffff: a viewer for comparing changes in both code and execution traces
abstract
One of the important purposes of code review is to find potential defects caused by other developers' code changes. When reviewing bug fixes, it is important to check the program behavior is properly changed to remove the bug. On the other hand, it is also important to check the program behavior that is not related to the bug is not changed. To investigate the program behavior, omniscient debugging which records all the runtime events is proposed. With omniscient debugging techniques, existing tools visualize multiple execution paths and the states of local variables of a method, but they are not focusing on code changes. In this paper, we implemented a prototype tool that compares and visualizes the difference between two execution traces caused by code changes. Each variable has a maximum of two lists of values, before and after the code changes, so we proposed their categorization based on their difference of length and contents. We also developed a viewer to show both code changes and the difference of execution traces at a glance by extending our previous viewer for omniscient debugging.
Tetsuya Kanda 0001, Kazumasa Shimari, Katsuro Inoue
ICPC3
2021 Finding repeated strings in code repositories and its applications to code-clone detection
abstract
Although researchers have created many advanced code-clone detection techniques, more effort is required to realize wide adaptation of these techniques in the industry. One of the reasons behind this is the reliance of these advanced techniques on lexing and parsing programs. Modern programming languages have complex lexical conventions and grammar, which evolve constantly. Therefore, using advanced code-clone detection techniques requires substantial and continuous effort. This paper proposes a lightweight language-independent method to detect code clones by simply finding repeated strings in a code repository, relying on neither lexing nor parsing. The proposed method is based on an efficient technique developed in a bio-informatics context to find repeated strings. We refer to the repeated strings in the source-code as weak Type-1 clones. Because the proposed technique normalizes newlines, tabs, and white spaces into a single white space, it can find clones in which newline positions or indentations are changed, as often in the case when copy-pasting occurs. Although the proposed method only finds verbatim copies, it also makes interesting observations regarding repository structures. Many developers may prefer the proposed simple approach because it is easier to understand than other advanced techniques that use heuristics, approximation, and machine learning.
Yoriyuki Yamagata, Fabien Hervé, Yuji Fujiwara, Katsuro Inoue
APSEC4
2021 NOD4J: Near-omniscient debugging tool for Java using size-limited execution trace
abstract
Logging is an important feature of a software system to record run-time information. Detailed logging allows developers to collect run-time information in situations where they cannot use an interactive debugger, such as continuous integration and web application server cases. However, extensive logging leads to larger execution traces because few instructions can be repeated many times. This paper presents our tool NOD4J, which monitors a Java program's execution within limited storage space constraints and annotates the source code with observed values in an HTML format. Developers can easily investigate the execution and share the report on a web server. We show two examples that our tool can debug defects using incomplete execution traces.
Kazumasa Shimari, Takashi Ishio, Tetsuya Kanda 0001, Naoto Ishida, Katsuro Inoue
Sci. Comput. Program.5
2020 Identifying Compiler and Optimization Options from Binary Code using Deep Learning Approaches
abstract
When compiling a source file, several flags can be passed to the compiler. These flags, however, can vary between debug and release compilation. In the release compilation, in fact, smaller or faster executables are usually preferred, whereas for a debug one, ease-of-debug is preferred over speed and no optimization is involved. After the compilation, however, most of the flags used cannot be inferred from the compiled file. These flags could be useful in case we want to classify if an older build was made for release or debug purposes, or to check if the file was compiled with flags that could expose vulnerabilities. In this paper we present a deep learning network capable of automatically detecting, with function granularity, the compiler used and the presence of optimization with 99% accuracy. We also analyze the change in accuracy when submitting increasingly shorter amounts of data, from 2048 up to a single byte, obtaining competitive results with less than 100 bytes. We also present our process in the huge dataset creation and manipulation, along with a comparison with other less successful networks using functions of varying size.
Davide Pizzolotto, Katsuro Inoue
ICSME2
2020 Clone Notifier: Developing and Improving the System to Notify Changes of Code Clones
abstract
A code clone is a code fragment that is identical or similar to it in the source code. It has been identified as one of the main problems in software maintenance. When a developer fixes a defect, they need to find the code clones corresponding to the code fragments. In this paper, we present Clone Notifier, a system that alerts on creations and changes of code clones to software developers. First, Clone Notifier identifies creations and changes of code clones. Subsequently, it groups them into four categories (new, deleted, changed, stable) and assigns labels (e.g., consistent, inconsistent) to them. Finally, it notifies on creations and changes of code clones along with the corresponding categories and labels. Clone Notifier and its video are available at: https://github.com/s-tokui/CloneNotifier.
Shogo Tokui, Norihiro Yoshida, Eunjong Choi, Katsuro Inoue
SANER4
2019 Near-Omniscient Debugging for Java Using Size-Limited Execution Trace
abstract
Logging is an important feature for a software system to record its run-time information. Detailed logging allows developers to collect information in situations where they cannot use an interactive debugger, such as continuous integration and web application server cases. However, extensive logging leads to larger execution traces because few instructions could be repeated many times. To record detailed program behavior within limited storage space constraints, we propose Near-Omniscient Debugging, a methodology that records an execution trace using fixed size buffers for each observed instruction. Our tool monitors a Java program's execution and annotates source code with observed values in an HTML format. Developers can easily investigate the execution and share the report on a web server. In case of DaCapo benchmark applications, our tool requires fewer than 1% of the complete execution traces to visualize all runtime values used by 66% of instructions that are executed less than 64 times. Developers also can obtain data dependencies with precision 91.8% and recall 79.0% using this tool.
Kazumasa Shimari, Takashi Ishio, Tetsuya Kanda 0001, Katsuro Inoue
ICSME4
2019 CCEvovis: a clone evolution visualization system for software maintenance
abstract
Understanding the evolution of code clones is important in software maintenance. With the information about how code clones evolve, both developers and researchers can understand the impacts of code clones and build a more robust code clone management system. So far, many studies have investigated the evolution of code clones to better understand the effects of code clones. However, only a few systems have been presented to support managing code clones based on the information about how code clone evolves. To mitigate this problem, in this paper, we present CCEvovis, a system that visualizes the evolved code clones across multiple versions of a program. CCEvovis highlights and visualizes the clone change to support software maintenance. CCEvovis is available at: https://github.com/hirotaka0616/CCEvovis.
Hirotaka Honda, Shogo Tokui, Kazuki Yokoi, Eunjong Choi, Norihiro Yoshida, Katsuro Inoue
ICPC6
2019 PADLA: a dynamic log level adapter using online phase detection
abstract
Logging is an important feature for a software system to record its run-time information. Although detailed logs are helpful to identify the cause of a failure in a program execution, constantly recording detailed logs of a long-running system is challenging because of its performance overhead and storage cost. To solve the problem, we propose PADLA (Phase-Aware Dynamic Log Level Adapter) that dynamically adjusts the log level of a running system so that the system can record irregular events such as performance anomalies in detail while recording regular events concisely. PADLA is an extension of Apache Log4j, one of the most popular logging framework for Java. It employs an online phase detection algorithm to recognize irregular events. It monitors run-time performance of a system and learns regular execution phases of a program. If it recognizes a performance anomalies, it automatically changes the log level of a system to record the detailed behavior. In the case study, PADLA successfully recorded a detailed log for performance analysis of a server system under high load while suppressing the amount of log data and performance overhead.
Tsuyoshi Mizouchi, Kazumasa Shimari, Takashi Ishio, Katsuro Inoue
ICPC4
2019 How do developers utilize source code from stack overflow?
abstract
Technical question and answer Q&A platforms, such as Stack Overflow, provide a platform for users to ask and answer questions about a wide variety of programming topics. These platforms accumulate a large amount of knowledge, including hundreds of thousands lines of source code. Developers can benefit from the source code that is attached to the questions and answers on Q&A platforms by copying or learning from (parts of) it. By understanding how developers utilize source code from Q&A platforms, we can provide insights for researchers which can be used to improve next-generation Q&A platforms to help developers reuse source code fast and easily. In this paper, we first conduct an exploratory study on 289 files from 182 open-source projects, which contain source code that has an explicit reference to a Stack Overflow post. Our goal is to understand how developers utilize code from Q&A platforms and to reveal barriers that may make code reuse more difficult. In 31.5% of the studied files, developers needed to modify source code from Stack Overflow to make it work in their own projects. The degree of required modification varied from simply renaming variables to rewriting the whole algorithm. Developers sometimes chose to implement an algorithm from scratch based on the descriptions from Stack Overflow answers, even if there was an implementation readily available in the post. In 35.5% of the studied files, developers used Stack Overflow posts as an information source for later reference. To further understand the barriers of reusing code and to obtain suggestions for improving the code reuse process on Q&A platforms, we conducted a survey with 453 open-source developers who are also on Stack Overflow. We found that the top 3 barriers that make it difficult for developers to reuse code from Stack Overflow are: (1) too much code modification required to fit in their projects, (2) incomprehensive code, and (3) low code quality. We summarized and analyzed all survey responses and we identified that developers suggest improvements for future Q&A platforms along the following dimensions: code quality, information enhancement & management, data organization, license, and the human factor. For instance, developers suggest to improve the code quality by adding an integrated validator that can test source code online, and an outdated code detection mechanism. Our findings can be used as a roadmap for researchers and developers to improve code reuse.
Shaowei Wang 0002, Cor-Paul Bezemer, Katsuro Inoue
Empir. Softw. Eng.4
2019 A Hybrid Approach for Improving the Design Quality of Web Service Interfaces
abstract
A key success of a Web service is to appropriately design its interface to make it easy to consume and understand. In the context of service-oriented computing (SOC), the service’s interface is the main source of interaction with the consumers to reuse the service functionality in real-world applications. The SOC paradigm provides a collection of principles and guidelines to properly design services to provide best practice of third-party reuse. However, recent studies showed that service designers tend to pay little care to the design of their service interfaces, which often lead to several side effects known as antipatterns . One of the most common Web service interface antipatterns is to expose a large number of semantically unrelated operations, implementing different abstractions, in one single interface. Such bad design practices may have a significant impact on the service reusability, understandability, as well as the development and run-time characteristics. To address this problem, in this article, we propose a hybrid approach to improve the design quality of Web service interfaces and fix antipatterns as a combination of both deterministic and heuristic-based approaches. The first step consists of a deterministic approach using a graph partitioning-based technique to split the operations of a large service interface into more cohesive interfaces, each one representing a distinct abstraction. Then, the produced interfaces will be checked using a heuristic-based approach based on the non-dominated sorting genetic algorithm (NSGA-II) to correct potential antipatterns while reducing the interface design deviation to avoid taking the service away from its original design. To evaluate our approach, we conduct an empirical study on a benchmark of 26 real-world Web services provided by Amazon and Yahoo. Our experiments consist of a quantitative evaluation based on design quality metrics, as well as a qualitative evaluation with developers to assess its usefulness in practice. The results show that our approach significantly outperforms existing approaches and provides more meaningful results from a developer’s perspective.
Ali Ouni 0001, Marouane Kessentini, Salah Bouktif, Katsuro Inoue
ACM Trans. Internet Techn.5
2018 Multilingual Detection of Code Clones Using ANTLR Grammar Definitions
abstract
So far, many tools have been developed for the detection of code clones in source code. The existing clone detection tools support only a limited number of programming languages and do not provide any easy extension mechanism to handle additional language. However, from our experience in industry/university collaboration, we found that many practitioners need to analyze source code written in various languages. In this paper, we propose an approach for the multilingual detection of code clones using grammar files for a parser generator ANTLR. We extended a clone detection tool CCFinderSW with the proposed approach and then apply the extended CCFinderSW to ANTLR grammar files for 43 languages. As a result, the files for 39 out of the 43 languages can be analyzed correctly by the extended CCFinderSW.
Yuichi Semura, Norihiro Yoshida, Eunjong Choi, Katsuro Inoue
APSEC4
2018 Investigating Vector-Based Detection of Code Clones Using BigCloneBench
abstract
In a vector-based approach to detecting code clones from source code, all code fragments in the source are mapped to a vector space and then code fragments are detected as code clones if they are neighbors in the vector space. So far, our research group has developed a vector-based approach using TF-IDF and cosine similarity. For the improvement of the vector-based approach, we preliminary investigated what kind of vectorization algorithms and similarity measurements are effective in terms of recall and detection time. In this paper, we present preliminary investigation results using BigCloneBench, a large-scale code clone benchmark.
Kazuki Yokoi, Eunjong Choi, Norihiro Yoshida, Katsuro Inoue
APSEC4
2018 "Was my contribution fairly reviewed?": a framework to study the perception of fairness in modern code reviews
abstract
Modern code reviews improve the quality of software products. Although modern code reviews rely heavily on human interactions, little is known regarding whether they are performed fairly. Fairness plays a role in any process where decisions that affect others are made. When a system is perceived to be unfair, it affects negatively the productivity and motivation of its participants. In this paper, using fairness theory we create a framework that describes how fairness affects modern code reviews. To demonstrate its applicability, and the importance of fairness in code reviews, we conducted an empirical study that asked developers of a large industrial open source ecosystem (OpenStack) what their perceptions are regarding fairness in their code reviewing process. Our study shows that, in general, the code review process in OpenStack is perceived as fair; however, a significant portion of respondents perceive it as unfair. We also show that the variability in the way they prioritize code reviews signals a lack of consistency and the existence of bias (potentially increasing the perception of unfairness). The contributions of this paper are: (1) we propose a framework---based on fairness theory---for studying and managing social behaviour in modern code reviews, (2) we provide support for the framework through the results of a case study on a large industrial-backed open source project, (3) we present evidence that fairness is an issue in the code review process of a large open source ecosystem, and, (4) we present a set of guidelines for practitioners to address unfairness in modern code reviews.
Daniel M. Germán, Gregorio Robles, Germán Poo-Caamaño, Xin Yang 0018, Hajimu Iida, Katsuro Inoue
ICSE6
2018 Cloned Buggy Code Detection in Practice Using Normalized Compression Distance
abstract
Software developers often write similar source code fragments in a software product. Since such code fragments may include the same mistake, developers have to inspect code clones if they found a bug in their code. In this study, we developed a tool to detect clones of a faulty code fragment for a software company, since existing code clone detection tools do not fit the requirements of the company. The tool employs Normalized Compression Distance for source code comparison, because its definition is understandable for developers, and also it is easy to support multiple programming languages. We conducted two experiments using an existing research dataset and actual examples. Based on the evidence, the tool has been deployed in several projects in the company.
Takashi Ishio, Naoto Maeda, Kensuke Shibuya, Katsuro Inoue
ICSME4
2018 How slim will my system be?: estimating refactored code size by merging clones
abstract
We have been doing code clone analysis with industry collaborators for a long time, and have been always asked a question, "OK, I understand my system contains a lot of code clones, but how slim will it be after merging redundant code clones?" As a software system evolves for long period, it would increasingly contain many code clones due to quick bug fix and new feature addition. Industry collaborators would recognize decay of initial design simplicity, and try to evaluate current system from the view point of maintenance effort and cost. As one of resources for the evaluation, the estimated code size by merging code clone is very important for them. In this paper, we formulate this issue as "slimming" problem, and present three different slimming methods, Basic, Complete, and Heuristic Methods, each of which gives a lower bound, upper bound, and modest reduction rates, respectively. Application of these methods to OSS systems written in C/C++ showed that the reduction rate is at most 5.7% of the total size, and to a commercial COBOL system, it is at most 15.4%. For this approach, we have gotten initial but very positive feedback from industry collaborators.
Norihiro Yoshida, Takuya Ishizu, Bufurod Edwards, Katsuro Inoue
ICPC4
2018 A generalized model for visualizing library popularity, adoption, and diffusion within a software ecosystem
abstract
The popularity of super repositories such as Maven Central and the CRAN is a testament to software reuse activities in both open-source and commercial projects alike. However, several studies have highlighted the risks and dangers brought about by application developers keeping dependencies on outdated library versions. Intelligent mining of super repositories could reveal hidden trends within the corresponding software ecosystem and thereby provide valuable insights for such dependency-related decisions. In this paper, we propose the Software Universe Graph (SUG) Model as a structured abstraction of the evolution of software systems and their library dependencies over time. To demonstrate the SUG's usefulness, we conduct an empirical study using 6,374 Maven artifacts and over 6,509 CRAN packages mined from their real-world ecosystems. Visualizations of the SUG model such as `library coexistence pairings' and `dependents diffusion' uncover popularity, adoption and diffusion patterns within each software ecosystem. Results show the Maven ecosystem as having a more conservative approach to dependency updating than the CRAN ecosystem.
Raula Gaikovina Kula, Coen De Roover, Daniel M. Germán, Takashi Ishio, Katsuro Inoue
SANER5
2018 Do developers update their library dependencies? - An empirical study on the impact of security advisories on library migration
Raula Gaikovina Kula, Daniel M. Germán, Ali Ouni 0001, Takashi Ishio, Katsuro Inoue
Empir. Softw. Eng.5
2018 An empirical study on the impact of refactoring activities on evolving client-used APIs
Raula Gaikovina Kula, Ali Ouni 0001, Daniel M. Germán, Katsuro Inoue
Inf. Softw. Technol.4
2018 Improving reusability of software libraries through usage pattern mining
Mohamed Aymen Saied, Ali Ouni 0001, Houari Sahraoui, Raula Gaikovina Kula, Katsuro Inoue, David Lo 0001
J. Syst. Softw.5
2017 CCFinderSW: Clone Detection Tool with Flexible Multilingual Tokenization
abstract
So far, many tools have been developed for the detection of code clones in source code. The existing clone detection tools support only a limited number of programming languages and do not provide any easy extension mechanism to handle additional language. However, from our experience in industry/university collaboration, we found that many practitioners need to analyze source code written in various languages. In this paper, we propose a clone detection tool CCFinderSW that has extension mechanism to handle addition language on demand from practitioners.
Yuichi Semura, Norihiro Yoshida, Eunjong Choi, Katsuro Inoue
APSEC4
2017 Source file set search for clone-and-own reuse analysis
abstract
Clone-and-own approach is a natural way of source code reuse for software developers. To assess how known bugs and security vulnerabilities of a cloned component affect an application, developers and security analysts need to identify an original version of the component and understand how the cloned component is different from the original one. Although developers may record the original version information in a version control system and/or directory names, such information is often either unavailable or incomplete. In this research, we propose a code search method that takes as input a set of source files and extracts all the components including similar files from a software ecosystem (i.e., a collection of existing versions of software packages). Our method employs an efficient file similarity computation using b-bit minwise hashing technique. We use an aggregated file similarity for ranking components. To evaluate the effectiveness of this tool, we analyzed 75 cloned components in Firefox and Android source code. The tool took about two hours to report the original components from 10 million files in Debian GNU/Linux packages. Recall of the top-five components in the extracted lists is 0.907, while recall of a baseline using SHA-1 file hash is 0.773, according to the ground truth recorded in the source code repositories.
Takashi Ishio, Yusuke Sakaguchi, Kaoru Ito, Katsuro Inoue
MSR4
2017 SoL Mantra: Visualizing Update Opportunities Based on Library Coexistence
abstract
In software development, software reuse has become a pivotal factor in creating and providing high-quality software at a reduced cost. The reuse of a code creates dependencies, which as they increase over time become difficult to manage and avoid compatibility issues or bugs. With newer version releases, come various quality improvements, new features and issue fixes, but deciding whether or not to adopt those is a difficult task for large software with a lot of dependencies. To address those difficulties, we propose SoL Mantra which is a tool that shows update opportunities by leveraging the Wisdom of the Crowd in a software ecosystem. Using this combined knowledge, our tool displays information about the complexity of each update opportunity. The orbital layout provides the means to visualize the update opportunities and demonstrate its merits by showcasing two examples from the JavaScript ecosystem. Through these examples, we demonstrate how maintainers can benefit from SoL Mantra's visual cues.
Boris Todorov, Raula Gaikovina Kula, Takashi Ishio, Katsuro Inoue
VISSOFT4
2017 An exploratory study on library aging by monitoring client usage in a software ecosystem
abstract
In recent times, use of third-party libraries has become prevalent practice in contemporary software development. Much like other code components, unmaintained libraries are a cause for concern, especially when it risks code degradation over time. Therefore, awareness of when a library should be updated is important. With the emergence of large libraries hosting repositories such as Maven Central, we can leverage the dynamics of these ecosystems to understand and estimate when a library is due for an update. In this paper, based on the concepts of software aging, we empirically explore library usage as a means to describe its age. The study covers about 1,500 libraries belonging to the Maven software ecosystem. Results show that library usage changes are not random, with 81.7% of the popular libraries fitting typical polynomial models. Further analysis show that ecosystem factors such as emerging rivals has an effect on aging characteristics. Our preliminary findings demonstrate that awareness of library aging and its characteristics is a promising step towards aiding client systems in the maintenance of their libraries.
Raula Gaikovina Kula, Daniel M. Germán, Takashi Ishio, Ali Ouni 0001, Katsuro Inoue
SANER5
2017 c-JRefRec: Change-based identification of Move Method refactoring opportunities
abstract
We propose, in this paper, a lightweight refactoring recommendation tool, namely c-JRefRec, to identify Move Method refactoring opportunities based on four heuristics using static and semantic program analysis. Our tool aims at identiying refactoring opportunities before a code change is committed to the codebase based on current code changes whenever the developer saves/compiles his code. We evaluate the efficiency of our approach in detecting Feature Envy smells and recommending Move Method refactorings to fix them on three Java open-source systems and 30 code changes. Results show that our approach achieves an average precision of 0.48 and 0.73 of recall and outperforms a state-of-the-art approach namely JDeodorant.
Naoya Ujihara, Ali Ouni 0001, Takashi Ishio, Katsuro Inoue
SANER4
2017 Analysis of license inconsistency in large collections of open source projects
Yuki Manabe 0001, Tetsuya Kanda 0001, Daniel M. Germán, Katsuro Inoue
Empir. Softw. Eng.5
2017 Search-based software library recommendation using multi-objective optimization
Ali Ouni 0001, Raula Gaikovina Kula, Marouane Kessentini, Takashi Ishio, Daniel M. Germán, Katsuro Inoue
Inf. Softw. Technol.6
2017 MORE: A multi-objective refactoring recommendation approach to introducing design patterns and fixing code smells
abstract
Refactoring is widely recognized as a crucial technique applied when evolving object‐oriented software systems. If applied well, refactoring can improve different aspects of software quality including readability, maintainability, and extendibility. However, despite its importance and benefits, recent studies report that automated refactoring tools are underused much of the time by software developers. This paper introduces an automated approach for refactoring recommendation, called MORE, driven by 3 objectives: (1) to improve design quality (as defined by software quality metrics), (2) to fix code smells, and (3) to introduce design patterns. To this end, we adopt the recent nondominated sorting genetic algorithm, NSGA‐III, to find the best trade‐off between these 3 objectives. We evaluated the efficacy of our approach using a benchmark of 7 medium and large open‐source systems, 7 commonly occurring code smells (god class, feature envy, data class, spaghetti code, shotgun surgery, lazy class, and long parameter list), and 4 common design pattern types (visitor, factory method, singleton, and strategy). Our approach is empirically evaluated through a quantitative and qualitative study to compare it against 3 different state‐of‐the art approaches, 2 popular multiobjective search algorithms, and random search. The statistical analysis of the results confirms the efficacy of our approach in improving the quality of the studied systems while successfully fixing 84% of code smells and introducing an average of 6 design patterns. In addition, the qualitative evaluation shows that most of the suggested refactorings (an average of 69%) are considered by developers to be relevant and meaningful.
Ali Ouni 0001, Marouane Kessentini, Mel Ó Cinnéide, Houari Sahraoui, Kalyanmoy Deb, Katsuro Inoue
J. Softw. Evol. Process.6
2017 Search-Based Web Service Antipatterns Detection
abstract
Service Oriented Architecture (SOA) is widely used in industry and is regarded as one of the preferred architectural design technologies. As with any other software system, service-based systems (SBSs) may suffer from poor design, i.e., antipatterns, for many reasons such as poorly planned changes, time pressure or bad design choices. Consequently, this may lead to an SBS product that is difficult to evolve and that exhibits poor quality of service (QoS). Detecting web service antipatterns is a manual, time-consuming and error-prone process for software developers. In this paper, we propose an automated approach for detection of web service antipatterns using a cooperative parallel evolutionary algorithm (P-EA). The idea is that several detection methods are combined and executed in parallel during an optimization process to find a consensus regarding the identification of web service antipatterns. We report the results of an empirical study using eight types of common web service antipatterns. We compare the implementation of our cooperative P-EA approach with random search, two single population-based approaches and one state-of-the-art detection technique not based on heuristic search. Statistical analysis of the obtained results demonstrates that our approach is efficient in antipattern detection, with a precision score of 89 percent and a recall score of 93 percent.
Ali Ouni 0001, Marouane Kessentini, Katsuro Inoue, Mel Ó Cinnéide
IEEE Trans. Serv. Comput.3
2016 Search-Based Peer Reviewers Recommendation in Modern Code Review
abstract
Code review is of primary importance in modern software development. It is widely recognized that peer review is an efficient and effective practice for improving software quality and reducing defect proneness. For successful review process, peer reviewers should have a deep experience and knowledge with the code being reviewed, and familiar to work and collaborate together. However, one of the main challenging tasks in modern code review is to find the most appropriate reviewers for submitted code changes. So far, reviewers assignment is still a manual, costly and time-consuming task. In this paper, we introduce a search-based approach, namely RevRec, to provide decision-making support for code change submitters and/or reviewers assigners to identify most appropriate peer reviewers for their code changes. RevRec aims at finding reviewers to be assigned for a code change based on their expertise and collaboration in past reviews using genetic algorithm (GA). We evaluated our approach on a benchmark of three open-source software systems, Android, OpenStack, and Qt. Results indicate that RevRec accurately recommends code reviewers with up to 59% of precision and 74% of recall. Our experiments provide evidence that leveraging reviewers expertise from their prior reviews and the socio-technical aspects of the team work and collaboration is relevant in improving the performance of peer reviewers recommendation in modern code review.
Ali Ouni 0001, Raula Gaikovina Kula, Katsuro Inoue
ICSME3
2016 SIM: An Automated Approach to Improve Web Service Interface Modularization
abstract
Service interface structure is of primary importance in SOA to ensure best practice of third-party reuse. One of the key factors for deploying successful services is assuring an adequate interface structure. However, a common bad service design practice is to place semantically unrelated operations in a single interface. This poor design practice typically result in a system which is difficult to comprehend, maintain and evolve providing low performance and reusability. To address this problem, we present an automated approach, SIM, to support service developers improve the quality of their interface modularization. Our approach analyzes structural and semantic relationships among the operations exposed in a service interface to identify chains of strongly related operations. The identified operation chains are used to define new interfaces with higher cohesion and better usability. We empirically evaluate our approach on a benchmark of 22 realworld Web services, provided by Amazon and Yahoo. The obtained results show that the produced interfaces are (i) able to improve the service design quality, and (ii) recognized as 'useful' from developers point of view in improving their service design. Additionally, we found that SIM significantly outperforms a recent state-of-the-art approach.
Ali Ouni 0001, Zouhour Salem, Katsuro Inoue, Makram Soui
ICWS3
2016 Revisiting the relationship between code smells and refactoring
abstract
Refactoring is a critical technique in evolving software systems. Martin Fowler presented a catalogue of refactoring patterns that defines a list of code smells and their corresponding refactoring patterns. This list aimed at supporting programmers in finding suitable refactoring patterns that remove code smells from their systems. However, a recent empirical study by Bavota et al. shows that refactoring rarely removes code smells which do not align with Fowler's catalog. To bridge the gap between them, we revisit the relationship between code smells and refactorings. In this study, we investigate whether developers apply appropriate refactoring patterns to fix code smells in three open source software systems.
Norihiro Yoshida, Tsubasa Saika, Eunjong Choi, Ali Ouni 0001, Katsuro Inoue
ICPC5
2016 Software ingredients: detection of third-party component reuse in Java software release
abstract
A software product is often dependent on a large number of third-party components. To assess potential risks, such as security vulnerabilities and license violations, a list of components and their versions in a product is important for release engineers and security analysts. Since such a list is not always available, a code comparison technique named Software Bertillonage has been proposed to test whether a product likely includes a copy of a particular component or not. Although the technique can extract candidates of reused components, a user still has to manually identify the original components among the candidates. In this paper, we propose a method to automatically select the most likely origin of components reused in a product, based on an assumption that a product tends to include an entire copy of a component rather than a partial copy. More concretely, given a Java product and a repository of jar files of existing components, our method selects jar files that can provide Java classes to the product in a greedy manner. To compare the method with the existing technique, we have conducted an evaluation using randomly created jar files including up to 1,000 components. The Software Bertillonage technique reports many candidates; the precision and recall are 0.357 and 0.993, respectively. Our method reports a list of original components whose precision and recall are 0.998 and 0.997.
Takashi Ishio, Raula Gaikovina Kula, Tetsuya Kanda 0001, Daniel M. Germán, Katsuro Inoue
MSR5
2016 On the Effectiveness of Vector-Based Approach for Supporting Simultaneous Editing of Software Clones
Seiya Numata, Norihiro Yoshida, Eunjong Choi, Katsuro Inoue
PROFES4
2016 Multi-Criteria Code Refactoring Using Search-Based Software Engineering: An Industrial Case Study
abstract
One of the most widely used techniques to improve the quality of existing software systems is refactoring—the process of improving the design of existing code by changing its internal structure without altering its external behavior. While it is important to suggest refactorings that improve the quality and structure of the system, many other criteria are also important to consider, such as reducing the number of code changes, preserving the semantics of the software design and not only its behavior, and maintaining consistency with the previously applied refactorings. In this article, we propose a multi-objective search-based approach for automating the recommendation of refactorings. The process aims at finding the optimal sequence of refactorings that (i) improves the quality by minimizing the number of design defects, (ii) minimizes code changes required to fix those defects, (iii) preserves design semantics, and (iv) maximizes the consistency with the previously code changes. We evaluated the efficiency of our approach using a benchmark of six open-source systems, 11 different types of refactorings (move method, move field, pull up method, pull up field, push down method, push down field, inline class, move class, extract class, extract method, and extract interface) and six commonly occurring design defect types (blob, spaghetti code, functional decomposition, data class, shotgun surgery, and feature envy) through an empirical study conducted with experts. In addition, we performed an industrial validation of our technique, with 10 software engineers, on a large project provided by our industrial partner. We found that the proposed refactorings succeed in preserving the design coherence of the code, with an acceptable level of code change score while reusing knowledge from recorded refactorings applied in the past to similar contexts.
Ali Ouni 0001, Marouane Kessentini, Houari Sahraoui, Katsuro Inoue, Kalyanmoy Deb
ACM Trans. Softw. Eng. Methodol.4
2015 Web Service Antipatterns Detection Using Genetic Programming
abstract
Service-Oriented Architecture (SOA) is an emerging paradigm that has radically changed the way software applications are architected, designed and implemented. SOA allows developers to structure their systems as a set of ready-made, reusable and compostable services. The leading technology used today for implementing SOA is Web Services. Indeed, like all software, Web services are prone to change constantly to add new user requirements or to adapt to environment changes. Poorly planned changes may risk introducing antipatterns into the system. Consequently, this may ultimately leads to a degradation of software quality, evident by poor quality of service (QoS). In this paper, we introduce an automated approach to detect Web service antipatterns using genetic programming. Our approach consists of using knowledge from real-world examples of Web service antipatterns to generate detection rules based on combinations of metrics and threshold values. We evaluate our approach on a benchmark of 310 Web services and a variety of five types of Web service antipatterns. The statistical analysis of the obtained results provides evidence that our approach is efficient to detect most of the existing antipatterns with a score of 85% of precision and 87% of recall.
Ali Ouni 0001, Raula Gaikovina Kula, Marouane Kessentini, Katsuro Inoue
GECCO4
2015 VerXCombo: an interactive data visualization of popular library version combinations
abstract
In large software systems, it is common practice to adopt third-party libraries. Decisions by system maintainers to either update or introduce new third-party libraries can range from trivial to complex. For instance, incompatibility between internal library dependencies may complicate adoption. Therefore, system maintainers especially need adequate assurance of any candidate library release. Using the 'wisdom of the crowd', VerXCombo aims to assist system maintainers by mining popular library dependency patterns of similar systems. Through data interactions, VerXCombo leverages parallel sets to break-down large and complex dataset into distinguishable patterns of 1.) popular and 2.) latest library dependency release combinations. Populating our tool with a maven library dependency dataset from over 4,000 Java Open Source projects, we demonstrate through a case scenario navigation and best fit combinations of the VerXCombo tool. A video highlighting the main features of the tool can be found at: http://goo.gl/wWPylL.
Yuki Yano, Raula Gaikovina Kula, Takashi Ishio, Katsuro Inoue
ICPC4
2015 Quick Trigger on Stack Overflow: A Study of Gamification-Influenced Member Tendencies
abstract
In recent times, gamification has become a popular technique to aid online communities stimulate active member participation. Gamification promotes a reward-driven approach, usually measured by response-time. Possible concerns of gamification could a trade-off between speedy over quality responses. Conversely, bias toward easier question selection for maximum reward may exist. In this study, we analyze the distribution gamification-influenced tendencies on the Q&A Stack Overflow online community. In addition, we define some gamification-influenced metrics related to response time to a question post. We carried experiments of a four-month period analyzing 101,291 members posts. Over this period, we determined a Rapid Response time of 327 seconds (5.45 minutes). Key findings suggest that around 92% of SO members have fewer rapid responses that non-rapid responses. Accepted answers have no clear relationship with rapid responses. However, we did find that rapid responses significantly contain tags that did not follow their usual tagging tendencies.
Xin Yang 0018, Raula Gaikovina Kula, Eunjong Choi, Katsuro Inoue, Hajimu Iida
MSR5
2015 A Method to Detect License Inconsistencies in Large-Scale Open Source Projects
abstract
The reuse of free and open source software (FOSS) components is becoming more and more popular. They usually contain one or more software licenses describing the requirements and conditions which should be followed when been reused. Licenses are usually written in the header of source code files as program comments. Removing or modifying the license header by re-distributors will result in the inconsistency of license with its ancestor, and may potentially cause license infringement. But to the best of our knowledge, no research has been devoted to investigate such kind of license infringements nor license inconsistencies. In this paper, we describe and categorize different types of license inconsistencies and propose a feasible method to detect them. Then we apply this method to Debian 7.5 and present the license inconsistencies found in it. With a manual analysis, we summarized various reasons behind these license inconsistencies, some of which imply license infringement and require the attention from the developers. This analysis also exposes the difficulty to discover license infringements, highlighting the usefulness of finding and maintaining source code provenance.
Yuki Manabe 0001, Tetsuya Kanda 0001, Daniel M. Germán, Katsuro Inoue
MSR5
2015 Extracting a unified directory tree to compare similar software products
abstract
Source code is often reused in software development. Although developers can avoid re-implementing features in existing products, doing so may result in a large number of similar software products. To understand the commonalities and variabilities of similar products, comparing their source code is critical. However, a product may change its own directory structure, even if the products share the same source code with other products. Hence, comparing source code among products in a systematic manner is difficult. In this paper, we propose a technique to extract and visualize a unified directory tree to compare the source code of similar products. This tree includes all directories of given products and merges corresponding directories into a single node. Since a node in a tree corresponds to multiple directories in products, developers can easily compare the contents of products. In our study, we implemented the visualization as a GUI tool. In addition, we conducted a case study using four Android products to demonstrate the tool's ability to assist developers in accessing the source code of multiple products.
Yusuke Sakaguchi, Takashi Ishio, Tetsuya Kanda 0001, Katsuro Inoue
VISSOFT4
2015 Evolution analysis for Accessibility Excessiveness in Java
abstract
In Java programs, access modifiers are used to control the accessibility of fields and methods from other objects. Choosing appropriate access modifiers is one of the key factors to improve program quality and to reduce potential vulnerability. In our previous work, we presented a static analysis method named Accessibility Excessiveness (AE) detection for each field and method in Java program. We have also developed an AE analysis tool named ModiChecker that analyzes each field and method of the input Java programs, and reports their excessiveness. In this paper, we have applied ModiChecker to several OSS repositories to investigate the evolution of AE over versions, and identified transition of AE status and the difference in the amount of AE change between major version releases and minor ones. Also we propose when to evaluate source code with AE analysis.
Kazuo Kobori, Makoto Matsushita, Katsuro Inoue
SANER3
2015 Trusting a library: A study of the latency to adopt the latest Maven release
abstract
With the popularity of open source library (re)use in both industrial and open source settings, `trust' plays vital role in third-party library adoption. Trust involves the assumption of both functional and non-functional correctness. Even with the aid of dependency management build tools such as Maven and Gradle, research have still found a latency to trust the latest release of a library. In this paper, we investigate the trust of OSS libraries. Our study of 6,374 systems in Maven Super Repository suggests that 82% of systems are more trusting of adopting the latest library release to existing systems. We uncover the impact of maven on latent and trusted library adoptions.
Raula Gaikovina Kula, Daniel M. Germán, Takashi Ishio, Katsuro Inoue
SANER4
2015 Improving multi-objective code-smells correction using development history
Ali Ouni 0001, Marouane Kessentini, Houari Sahraoui, Katsuro Inoue, Mohamed Salah Hamdi
J. Syst. Softw.4
2014 Repeatedly-executed-method viewer for efficient visualization of execution paths and states in Java
abstract
The state of a program at runtime is useful information for developers to understand a program. Omniscient debugging and logging-based tools enable developers to investigate the state of a program at an arbitrary point of time in an execution. While these tools are effective to analyze the state at a single point of time, they might be insufficient to understand the generic behavior of a method which includes various control-flow paths. In this paper, we propose REMViewer (Repeatedly-Executed-Method Viewer), or a tool that visualizes multiple execution paths of a Java method. The tool shows each execution path in a separated view so that developers can firstly select actual execution paths of interest and then compare the state of local variables in the paths.
Toshinori Matsumura, Takashi Ishio, Yu Kashima, Katsuro Inoue
ICPC4
2014 Identifying Source Code Reuse across Repositories Using LCS-Based Source Code Similarity
abstract
Developers often reuse source files developed for another project. In order to update a reused file to a newer version released by the original project, developers have to track which revision of a file was reused and how its content was modified. However, such tracking is tedious for developers. Many projects keep older versions of files whose bugs are already fixed in the original project. In this paper, we propose a technique to automatically identify source code reuse relationships between two repositories. Using a similarity metric based on longest common subsequence, we identify pairs of similar revisions of files across the repositories. To evaluate our approach, we have analyzed eight project pairs of open source software projects and compared the result with the recorded information in the repositories. As a result, we have identified 1394 file revisions as instances of source code reuse. While 75.3% of the instances are recorded in the repositories, 20.1% of the instances are unrecorded but recovered by our approach.
Naohiro Kawamitsu, Takashi Ishio, Tetsuya Kanda 0001, Raula Gaikovina Kula, Coen De Roover, Katsuro Inoue
SCAM6
2014 The Ekeko/X Program Transformation Tool
abstract
Developers often need to perform repetitive changes to source code. For instance, to repair several instances of a bug or to update all clients of a library to a newer version. Manually performing such changes is laborious and error-prone. Program transformation tools enable automating changes, but specifying changes as a program transformation requires significant expertise. Code templates are often touted as a remedy, yet have never been endorsed wholeheartedly. Their use is mostly limited to expressing the syntactic characteristics of the intended change subjects. Less familiar means have to be resorted to for expressing their structural, control flow, and data flow characteristics. In this tool paper, we introduce a decidedly template-driven program transformation tool called Ekeko/X. Its specifications feature templates for specifying all of the aforementioned characteristics of its subjects. To this end, developers can associate different directives with individual components of a template. Each matching directive imposes particular constraints on the matches for the component it is associated with. Rewriting directives, on the other hand, determine how each match should be changed. We develop Ekeko/X from the ground up, starting from its applicative logic meta-programming foundation. We highlight the key choices in this implementation and demonstrate its use through two example program transformations.
Coen De Roover, Katsuro Inoue
SCAM2
2014 Visualizing the Evolution of Systems and Their Library Dependencies
abstract
System maintainers face several challenges stemming from a system and its library dependencies evolving separately. Novice maintainers may lack the historical knowledge required to efficiently manage an inherited system. While some libraries are regularly updated, some systems keep a dependency on older versions. On the other hand, maintainers may be unaware that other systems have settled on a different version of a library. In this paper, we visualize how the dependency relation between a system and its dependencies evolves from two perspectives. Our system-centric dependency plots (SDP) visualize the successive library versions a system depends on over time. The radial layout and heat-map metaphor provide visual clues about the change in dependencies along the system's release history. From this perspective, maintainers can navigate to a library-centric dependants diffusion plot (LDP). The LDP is a time-series visualization that shows the diffusion of users across the different versions of a library. We demonstrate on real-world systems how maintainers can benefit from our visualizations through four case scenarios.
Raula Gaikovina Kula, Coen De Roover, Daniel M. Germán, Takashi Ishio, Katsuro Inoue
VISSOFT5
2014 Special issue on software clones (IWSC'12)
Katsuro Inoue, Rainer Koschke, Jens Krinke
Sci. Comput. Program.1
2013 Applying clone change notification system into an industrial development process
abstract
Programmers tend to write code clones unintentionally even in the case that they can easily avoid them. Clone change management is one of crucial issues in open source software (OSS) development as well as in industrial software development (e.g., development of social infrastructure, financial system, and medical equipment). When an industrial developer fixes a defect, he/she has to find the code clones corresponding to the code fragment including it. So far, several studies performed on the analysis of clone evolution in OSS. However, to our knowledge, a few researches have been reported on an application of a clone change notification system to industrial development process. In this paper, we introduce a system for notifying creation and change of code clones, and then report on the experience with 40-days application of it into a development process in NEC Corporation. In the industrial application, a developer successfully identified ten unintentionally-developed clones that should be refactored.
Yuki Yamanaka, Eunjong Choi, Norihiro Yoshida, Katsuro Inoue, Tateki Sano
ICPC4
2013 Extraction of product evolution tree from source code of product variants
abstract
A large number of software products may be derived from an original single product. Although software product line engineering is advocated as an effective approach to maintaining such a family of products, re-engineering existing products requires developers to understand the evolution history of the products. This can be challenging because developers typically only have access to product source code. In this research, we propose to extract a Product Evolution Tree that approximates the evolution history from source code of products. Our key idea is that two successive products are the most similar to one another in the evolution history. We construct a Product Evolution Tree as a minimum spanning tree whose cost function is defined by the number of similar files between products. As an experiment, we extracted Product Evolution Trees from 6 datasets of open-source projects. The result showed that 53% to 92% of edges in the extracted trees were consistent with the actual evolution history of the projects.
Tetsuya Kanda 0001, Takashi Ishio, Katsuro Inoue
SPLC3
2012 An Experience Report on Analyzing Industrial Software Systems Using Code Clone Detection Techniques
abstract
A variety of application results of code clone detection and analysis has been reported. There are many reports of code clone detection and analysis on open source software whereas few reports on industrial systems are open to the public. This paper reports an experience of code clone analysis on a governmental project. In the project, a software system was developed by multiple Japanese vendors. We detected and analyzed code clones in the system, and found that there were many code clones in the project, however we concluded that the presence of the code clones did not have negative impacts on the maintenance of the system because of the following reasons: (1) when different modules are similar to each other in the design document, they also share many code clones in the source code, (2) code clones located in trusted modules, which are libraries maintained by one of the companies.
Norihiro Yoshida, Yoshiki Higo, Shinji Kusumoto, Katsuro Inoue
APSEC4
2012 Where does this code come from and where does it go? - Integrated code history tracker for open source systems
abstract
When we reuse a code fragment in an open source system, it is very important to know the history of the code, such as the code origin and evolution. In this paper, we propose an integrated approach to code history tracking for open source repositories. This approach takes a query code fragment as its input, and returns the code fragments containing the code clones with the query code. It utilizes publicly available code search engines as external resources. Based on this model, we have designed and implemented a prototype system named Ichi Tracker. Using Ichi Tracker, we have conducted three case studies. These case studies show the ancestors and descendents of the code, and we can recognize their evolution history.
Katsuro Inoue, Yusuke Sasaki, Pei Xia, Yuki Manabe 0001
ICSE1
2012 A lightweight visualization of interprocedural data-flow paths for source code reading
abstract
To understand the behavior of a program, developers must read source code fragments in various modules. For developers investigating data-flow paths among modules, a call graph is too abstract since it does not visualize how parameters of method calls are related to each other. On the other hand, a system dependence graph is too fine-grained to investigate interprocedural data-flow paths. In this research, we propose an intermediate-level of visualization; we visualize interprocedural data-flow paths among method parameters and fields with summarized intraprocedural data-flow paths. We have implemented our visualization as an Eclipse plug-in for Java. The tool comprises a lightweight data-flow analysis and an interactive graph viewer using fractal value to extract a small subgraph of data-flow related to variables specified by a developer. A case study has shown our visualization enabled developers to investigate more data-flow paths in a fixed time slot. In addition, we report our lightweight data-flow analysis can generate precise data-flow paths for 98% of Java methods.
Takashi Ishio, Shogo Etsuda, Katsuro Inoue
ICPC3
2011 Fifth international workshop on software clones: (IWSC 2011)
abstract
Software clones are identical or similar pieces of code, design or other artifacts. Clones are known to be closely related to various issues in software engineering, such as software quality, complexity, architecture, refactoring, evolution, licensing, plagiarism, and so on. Various characteristics of software systems can be uncovered through clone analysis, and system restructuring can be performed by merging clones.
James R. Cordy, Katsuro Inoue, Stan Jarzabek, Rainer Koschke
ICSE2
2011 ImpactScale: Quantifying change impact to predict faults in large software systems
abstract
In software maintenance, both product metrics and process metrics are required to predict faults effectively. However, process metrics cannot be always collected in practical situations. To enable accurate fault prediction without process metrics, we define a new metric, ImpactScale. ImpactScale is the quantified value of change impact, and the change propagation model for ImpactScale is characterized by probabilistic propagation and relation-sensitive propagation. To evaluate ImpactScale, we predicted faults in two large enterprise systems using the effort-aware models and Poisson regression. The results showed that adding ImpactScale to existing product metrics increased the number of detected faults at 10% effort (LOC) by over 50%. ImpactScale also improved the predicting model using existing product metrics and dependency network measures.
Kenichi Kobayashi 0001, Akihiko Matsuo, Katsuro Inoue, Yasuhiro Hayase, Manabu Kamimura, Toshiaki Yoshino
ICSM3
2011 A Pluggable Tool for Measuring Software Metrics from Source Code
abstract
This paper proposes a new mechanism to measure a variety of source code metrics at low cost. The proposed mechanism is very promising because it realizes to add new metrics as necessary. Users do not need to use multiple measurement tools for measuring multiple metrics. The proposed mechanism has been implemented as an actual software tool MASU. This paper shows how using MASU makes it easy and less costly to develop plugins of the CK metrics suite.
Yoshiki Higo, Akira Saitoh, Goro Yamada, Tatsuya Miyake, Shinji Kusumoto, Katsuro Inoue
IWSM/Mensura6
2010 Fourth International Workshop on Software Clones (IWSC)
abstract
Software clones are identical or similar pieces of code. They are often the result of copy--and--paste activities as ad-hoc code reuse by programmers. Software clones research is of high relevance for the industry. Many researchers have reported high rates of code cloning in both industrial and open-source systems.
Katsuro Inoue, Stan Jarzabek, James R. Cordy, Rainer Koschke
ICSE (2)1
2010 A sentence-matching method for automatic license identification of source code files
abstract
The reuse of free and open source software (FOSS) components is becoming more prevalent. One of the major challenges in finding the right component is finding one that has a license that is e for its intended use. The license of a FOSS component is determined by the licenses of its source code files. In this paper, we describe the challenges of identifying the license under which source code is made available, and propose a sentence-based matching algorithm to automatically do it. We demonstrate the feasibility of our approach by implementing a tool named Ninka. We performed an evaluation that shows that Ninka outperforms other methods of license identification in precision and speed. We also performed an empirical study on 0.8 million source code files of Debian that highlight interesting facts about the manner in which licenses are used by FOSS
Daniel M. Germán, Yuki Manabe 0001, Katsuro Inoue
ASE3
2010 Finding file clones in FreeBSD Ports Collection
abstract
In Open Source System (OSS) development, software components are often imported and reused; for this reason we might expect that files are copied in multiple projects (file clones). In this paper, we propose a file clone detection tool called FCFinder and show the analysis performed with it on the FreeBSD Ports Collection, a large OSS project collection. We found many file clones among similar or related projects, which are systematically introduced from base projects.
Yusuke Sasaki, Tetsuo Yamamoto, Yasuhiro Hayase, Katsuro Inoue
MSR4
2010 Standardizing the Software Tag in Japan for Transparency of Development
Masateru Tsunoda, Tomoko Matsumura, Hajimu Iida, Kozo Kubo, Shinji Kusumoto, Katsuro Inoue, Ken-ichi Matsumoto
PROFES6
2009 IT SPIRAL: A Case Study in Scalable Software Engineering Education
abstract
IT SPIRAL is a collaborative project by nine universities and four industries to develop a common curriculum for teaching software engineering. It combines existing foundation educational practices at the individual universities, a shared DVD library on advanced software engineering topics, and common intensive sessions led by industry participants. It aims to develop advanced IT skills in top-level students, shared educational skills and materials among the universities, and practical cooperation with industry to focus and advance masters level software engineering education. IT SPIRAL combines fundamentals, advanced topics, and a practical focus in a scalable approach to developing world-class software engineers.
Michael Barker, Katsuro Inoue
CSEE&T2
2009 Assessing the impact of framework changes using component ranking
abstract
Most of today's software applications are built on top of libraries or frameworks. Just as applications evolve, libraries and frameworks also evolve. Upgrading is straightforward when the framework changes preserve the API and behavior of the offered services. However, in most cases, major changes are introduced with the new framework release, which can have a significant impact on the application. Hence, a common question a framework user might ask is, ldquoIs it worth upgrading to the new framework version?rdquo In this paper, we study the evolution of an application and its underlying framework to understand the information we can get through a multi-version use relation analysis. We use component rank changes to measure this impact. Component rank measurement is a way of quantifying the importance of a component by its usage. As framework components are used by applications, the rankings of the components are changed. We use component ranking to identify the core components in each framework version. We also confirm that upgrading to the new framework version has an impact to a component rank of the entire system and the framework, and this impact not only involves components which use the framework directly, but also other indirectly-related components. Finally, we also confirm that there is a difference in the growth of use relations between application and framework.
Reishi Yokomori, Harvey P. Siy, Masami Noro, Katsuro Inoue
ICSM4
2009 OGAN: Visualizing object interaction scenarios based on dynamic interaction context
abstract
Visualizing an execution trace of an object-oriented system as sequence diagrams is effective to understand the behavior of the system. To support developers to understand concrete interaction among classes, our tool named OGAN extracts sequence diagrams representing interaction scenarios for a pair of classes specified by a user. OGAN classifies objects into groups based on their neighbor classes that directly interact with the objects, and visualizes interaction scenarios for each pair of object groups.
Munakata Satoshi, Takashi Ishio, Katsuro Inoue
ICPC3
2009 An Empirical Study of the Feedback of the In-process Measurement in a Japanese Consortium-type Software Project
Yoshiki Mitani, Tomoko Matsumura, Katsuro Inoue, Mike Barker, Akito Monden, Ken-ichi Matsumoto
SEKE3
2008 Software tag for traceability and transparency of maintenance
abstract
We have proposed the notion of software tag, which is a complex of various characteristic elements of software development project. Empirical data for the project is collected and abstracted into the tag, and the resulting tag is given back to the software purchaser who can evaluate the quality of software product and development process. In this paper, we will discuss an extension of the application area of the software tag into the software maintenance activities.
Katsuro Inoue
ICSM1
2008 A metric-based approach to identifying refactoring opportunities for merging code clones in a Java software system
abstract
Abstract A code clone is a code fragment that has other code fragments identical or similar to it in the source code. The presence of code clones is generally regarded as one factor that makes software maintenance more difficult. For example, if a code fragment with code clones is modified, it is necessary to consider whether each of the other code clones has to be modified as well. Removing code clones is one way of avoiding problems that arise due to the presence of code clones. This makes the source code more maintainable and more comprehensible. This paper proposes a set of metrics that suggest how code clones can be refactored. As well, the tool Aries, which automatically computes these metrics, is presented. The tool gives metrics that are indicators for certain refactoring methods rather than suggesting the refactoring methods themselves. The tool performs only lightweight source code analysis; hence, it can be applied to a large number of code lines. This paper also describes a case study that illustrates how this tool can be used. Based on the results of this case study, it can be concluded that this method can efficiently merge code clones. Copyright © 2008 John Wiley & Sons, Ltd.
Yoshiki Higo, Shinji Kusumoto, Katsuro Inoue
J. Softw. Maintenance Res. Pract.3
2007 Simultaneous Modification Support based on Code Clone Analysis
abstract
Maintaining software systems becomes more difficult as their size and complexity increase. One factor that makes software maintenance more difficult is the presence of code clones. A code clone is defined as a code fragment occurring more than once in identical or similar form into a software system. For example, the presence of code clones is a big factor of overlooking some places that should be modified simultaneously. One technique that helps the number of code clones is refactoring. There are several research efforts that provide support to refactor code clones, but unfortunately some code clones cannot or should not be refactored (ex. stereotyped process, absence of abstraction functionality, performance enhancement). In order to support maintaining the consistency among code clones, we propose a simultaneous modification support method. Given a software system, firstly, a maintainer identifies a code fragment that must be modified. Then, only the code clones between the identified code fragment and the source files of the software system are detected. We developed a simultaneous modification support tool, Libra, and applied it to open source software systems. The results showed that Libra was a good searching tool as much as grep, which is a useful tool of UNIX.
Yoshiki Higo, Yasushi Ueda, Shinji Kusumoto, Katsuro Inoue
APSEC4
2007 Accountability and Traceability in Global Software Engineering (ATGSE2007)
abstract
An Overview of the new workshop on Accountability and Traceability in Global Software Engineering (ATGSE2007) will be introduced here.
Katsuro Inoue, Mike Barker
APSEC1
2007 Proposal of a Complete Life Cycle In-Process Measurement Model Based on Evaluation of an In-Process Measurement Experiment Using a Standardized Requirement Definition Process
abstract
This paper focuses on in-process measurements during requirements definition where measurements of processes and products are relatively difficult. However, development processes in Japan based on the enterprise architecture method provide standardized formats for such upstream processes and products, allowing in-process measurements. Based on previous work and on this examination of in-process measurements of requirements definition with the enterprise architecture method and previous results of empirical studies of in-process measurements and empirically validates of later development processes, this paper proposes a new measurement model, the "full in-process process and product (I-PAP) measurement model," which includes the complete software development process from requirements to maintenance. Standardization of the requirements definition phase using the enterprise architecture method in Japan allows in-process measurement across the complete development lifecycle. Combining this with collaborative filtering and a project benchmark database will support project evaluation, estimation, and prediction.
Yoshiki Mitani, Tomoko Matsumura, Mike Barker, Seishiro Tsuruho, Katsuro Inoue, Ken-ichi Matsumoto
ESEM5
2007 Very-Large Scale Code Clone Analysis and Visualization of Open Source Programs Using Distributed CCFinder: D-CCFinder
abstract
The increasing performance-price ratio of computer hardware makes possible to explore a distributed approach at code clone analysis. This paper presents D-CCFinder, a distributed approach at large-scale code clone analysis. D-CCFinder has been implemented with 80 PC workstations in our student laboratory, and a vast collection of open source software with about 400 million lines in total has been analyzed with it in about 2 days. The result has been visualized as a scatter plot, which showed the presence of frequently used code as easy recognizable patterns. Also, D-CCFinder has been used to analyze a single software system against the whole collection in order to explore the presence of code imported from open source software.
Simone Livieri, Yoshiki Higo, Makoto Matsushita, Katsuro Inoue
ICSE4
2007 An Empirical Study of Process and Product Metrics Based on In-process Measurements of a Standardized Requirements Definition Phase
Yoshiki Mitani, Tomoko Matsumura, Mike Barker, Seishiro Tsuruho, Katsuro Inoue, Ken-ichi Matsumoto
IWSM/Mensura5
2007 Method and implementation for investigating code clones in a software system
Yoshiki Higo, Toshihiro Kamiya, Shinji Kusumoto, Katsuro Inoue
Inf. Softw. Technol.4
2006 Evaluation of Source Code Updates in Software Development Based on Component Rank
abstract
Essential activities for the achievement of trouble-free software development are monitoring a software product and management of a software project. Monitoring changes that have major impact, however, is usually a very hard to complete because every engineer usually could not know entire source code in detail. In this paper, we propose our metric for source code updates based on component rank. Software components and their use-relation alter as development goes. The component rank also changes as a response to the changes of use-relation among components. We use the degree representing change of component rank as a metric of impact for a source code update in a development. We applied the metric to open source projects, and demonstrated that the metric is useful to know refactoring activities or important updates. We also discuss how the metric can contribute for process management.
Reishi Yokomori, Masami Noro, Katsuro Inoue
APSEC3
2006 Effects of software industry structure on a research framework for empirical software engineering
abstract
The authors describe a new research framework for applying empirical software engineering methods in industrial practice and accomplishments in using it. The selected target for applying the framework is a governmentally funded software development project involving multiple vendors. This project involved in-process project data measurement in real time, data sharing with industry and academia, data analysis, and feedback to the project members. Today the project is in the system integration process. This paper shows the value of this research framework and describes issues of empirical data sharing between industry and academia which have emerged while using it. This experiment raised two major issues. One is the necessity of a new research framework for project measurement called the "Macro Measurement Tool". The other is effects of the software industry structure on this framework.
Yoshiki Mitani, Nahomi Kikuchi, Tomoko Matsumura, Satoshi Iwamura, Yoshiki Higo, Katsuro Inoue, Mike Barker, Ken-ichi Matsumoto
ICSE6
2006 JAAT: Java Alias Analysis Tool for Program Maintenance Activities
abstract
Alias analysis is a method for extracting sets of expressions which may possibly refer to the same memory locations during program execution. Although many researchers have already proposed analysis methods for the purpose of program optimization, difficulties still remain in applying such methods to practical software engineering tools in the sense of precision, extensibility and scalability. Focusing mainly on a practical use for program maintenance activities such as program debugging and understanding, we propose an alias analysis method for object-oriented programs and discuss our implementation. Using this method, we have developed a tool named JAAT. Our proposed method employs a two-phase, on-demand, and instance-based algorithm, in which intra-class analysis is done in phase 1 for whole programs and libraries, and inter-class analysis is done in phase 2 only for a user-demanded target. JAAT can analyze large programs or libraries such as JDK class library. Also, JAAT includes various features for program maintenance activities, such as GUI for displaying aliases, and an XML database for storing analysis information
Fumiaki Ohata, Katsuro Inoue
ISORC2
2006 MUDABlue: An automatic categorization system for Open Source repositories
Shinji Kawaguchi, Pankaj K. Garg, Makoto Matsushita, Katsuro Inoue
J. Syst. Softw.4
2005 Aspect-Oriented Modularization of Assertion Crosscutting Objects
abstract
Assertion checking is a powerful tool to detect software faults during debugging, testing and maintenance. Although assertion documents the behavior of one component, it is hard to document relations and interactions among several objects since such assertion statements are spread across the modules. Therefore, we propose to modularize such assertion as an aspect in order to improve software maintainability. In this paper, taking Observer pattern as an example, we point out that some assertions tend to be crosscutting, and propose a modularization of such assertion with aspect-oriented language. We show a limitation of traditional assertion and effectiveness of assertion aspect through the case study, and discuss various situations to which assertion aspects are applicable.
Takashi Ishio, Shinji Kusumoto, Katsuro Inoue, Toshihiro Kamiya
APSEC3
2005 CoxR: Open Source Development History Search System
abstract
In typical open source software development, developers use revision control systems for product management, mailing list systems for human communications, and bug tracking systems for process management. All of these systems store development histories of the products that show significant information of problems during the development. However, it would be a hard job to retrieve useful information related to a current problem faced by developers. In this paper, we describe a software development supporting system CoxR that is capable of crawling the development histories. CoxR creates software development information Web which consists of developers, emails, and program deltas, and provides an interface to search, navigate, browse, and retrieve past development results. Through a case study, we confirmed that CoxR helps developers to solve their problems by making it easier to search development history.
Makoto Matsushita, Kei Sasaki, Katsuro Inoue
APSEC3
2005 Mega Software Engineering
Katsuro Inoue, Pankaj K. Garg, Hajimu Iida, Ken-ichi Matsumoto, Koji Torii
PROFES1
2005 Measuring Similarity of Large Software Systems Based on Source Code Correspondence
Tetsuo Yamamoto, Makoto Matsushita, Toshihiro Kamiya, Katsuro Inoue
PROFES4
2005 Ranking Significance of Software Components Based on Use Relations
abstract
Collections of already developed programs are important resources for efficient development of reliable software systems. In this paper, we propose a novel graph-representation model of a software component library (repository), called component rank model. This is based on analyzing actual usage relations of the components and propagating the significance through the usage relations. Using the component rank model, we have developed a Java class retrieval system named SPARS-J and applied SPARS-J to various collections of Java files. The result shows that SPARS-J gives a higher rank to components that are used more frequently. As a result, software engineers looking for a component have a better chance of finding it quickly. SPARS-J has been used by two companies, and has produced promising results.
Katsuro Inoue, Reishi Yokomori, Tetsuo Yamamoto, Makoto Matsushita, Shinji Kusumoto
IEEE Trans. Software Eng.1
2004 MUDABlue: An Automatic Categorization System for Open Source Repositories
abstract
Open source communities typically use a software repository to archive various software projects with their source code, mailing list discussions, documentation, bug reports, and so forth. For example, SourceForge currently hosts over seventy thousand open source software systems. Because of the size of the rich information content, such repositories offer numerous opportunities for sharing information among projects. For example, one would like to know a set of projects that are related or similar to each other, so that the project groups can collaborate and share their work. With thousands of projects in typical repositories, however, manually locating related projects can be difficult. Hence, we propose MUDABlue, a tool that automatically categorizes software systems. MUDABlue has three major aspects: 1) it relies on no other information than the source code, 2) it determines category sets automatically, and 3) it allows a software system to be a member of multiple categories. MUDABlue has a Web interface to visualize determined categories, which eases browsing a software repository. We show the effectiveness of MUDABlue's categorization capability by comparing its generated categories with that of some other existing research tools.
Shinji Kawaguchi, Pankaj K. Garg, Makoto Matsushita, Katsuro Inoue
APSEC4
2004 Debugging Support for Aspect-Oriented Program Based on Program Slicing and Call Graph
abstract
Aspect-oriented programming (AOP) introduces a new software module unit named aspect to encapsulate crosscutting concerns. While AOP modularizes crosscutting concerns to improve maintainability and reusability, AOP introduces a new factor of complexity. It is difficult to find defects caused by an aspect modifying or preventing the behavior of other objects and aspects. We examine a method to support a debugging task in aspect-oriented software development. We propose an application of a call graph generation and program slicing to assist in debugging. A call graph visualizes control dependence relations between objects and aspects and supports the detection of an infinite loop. On the other hand, program slicing shows the user changes of dependence relations caused by aspects. We implement a program-slicing tool for AspectJ and apply it to certain programs. The experiment illustrates how our approach effectively helps developers understand the influence of aspects in a program.
Takashi Ishio, Shinji Kusumoto, Katsuro Inoue
ICSM3
2004 Refactoring Support Based on Code Clone Analysis
Yoshiki Higo, Toshihiro Kamiya, Shinji Kusumoto, Katsuro Inoue
PROFES4
2004 Assessing defect detection performance of interacting teams in object-oriented design inspection
Giedre Sabaliauskaite, Shinji Kusumoto, Katsuro Inoue
Inf. Softw. Technol.3
2003 Component Rank: Relative Significance Rank for Software Component Search
abstract
Collections of already developed programs are important resources for efficient development of reliable software systems. In this paper, we propose a novel method of ranking software components, called Component Rank, based on analyzing actual use relations among the components and propagating the significance through the use relations. We have developed a component-rank computation system, and applied it to various Java programs. The result is promising such that non-specific and generic components are ranked high. Using the Component Rank system as a core part, we are currently developing Software Product Archiving, analyzing, and Retrieving System named SPARS.
Katsuro Inoue, Reishi Yokomori, Hikaru Fujiwara, Tetsuo Yamamoto, Makoto Matsushita, Shinji Kusumoto
ICSE1
2003 Java Program Analysis Projects in Osaka University: Aspect-Based Slicing System ADAS and Ranked-Component Search System SPARS-J
abstract
In our research demonstration, we show two development support systems for Java programs. One is an Aspect-oriented Dynamic Analysis and Slice calculation system named ADAS, and another is a Software Product archiving, Analyzing, and Retrieving System for Java named SPARS-J.
Reishi Yokomori, Takashi Ishio, Tetsuo Yamamoto, Makoto Matsushita, Shinji Kusumoto, Katsuro Inoue
ICSE6
2003 Case studies to evaluate a domain specific application framework based on complexity and functionality metrics
Hikaru Fujiwara, Shinji Kusumoto, Katsuro Inoue, Ayane Suzuki, Toshifusa Ootsubo, Katsuhiko Yuura
Inf. Softw. Technol.3
2003 Further investigations of reading techniques for object-oriented design inspection
Giedre Sabaliauskaite, Fumikazu Matsukawa, Shinji Kusumoto, Katsuro Inoue
Inf. Softw. Technol.4
2002 On Detection of Gapped Code Clones using Gap Locations
abstract
It is generally accepted that a code clone is one factor making software maintenance difficult. A code clone is a code portion in source files that is identical or similar to another. Clones are introduced because of various reasons such as reusing code by 'copy-and-paste' and so on. Since developers usually modify the copied-and-pasted code portions, there are gaps between them and the original code portion. We call such code portions including gaps gapped code clones. Several code clone detection methods, which consider such gaps, have been proposed. However, it is costly to detect all gapped code clones. This paper proposes a new method to visualize gapped code clones as if they were actually detected, based on the detection results of conventional code clones. Using the proposed method, the developer can specify target clones efficiently. Moreover, we implement the proposed method in the maintenance support environment and conduct an experimental evaluation.
Yasushi Ueda, Toshihiro Kamiya, Shinji Kusumoto, Katsuro Inoue
APSEC4
2002 Principles of software evolution: 5th international workshop on principles of software evolution (IWPSE 2002)
abstract
We present an overview of the 5th International Workshop on Principles of Software Evolution (IWPSE 2002).
Mikio Aoyama, Katsuro Inoue, Václav Rajlich
ICSE2
2002 Function point measurement from Java programs
abstract
Function point analysis (FPA) was proposed to help measure the functionality of software systems. It is used to estimate the effort required for the software development. However, it has been reported that since function point measurement involves judgment on the part of the measurer, differences for the same product may occur even in the same organization. Also, if an organization tries to introduce FPA, FP will have to be measured from the past software developed there, and this measurement is cost-consuming. In this paper, we intend to examine the possibility to measure FP from source code automatically. At first, we propose measurement rules to count data and transactional functions for object-oriented program based on IFPUG method and develop the function point measurement tool. Then, we have applied the tool to practical Java programs in a computer company and examined the difference between the FP values obtained by the tool and those of an FP measurement specialist. As the results, the number of data and transactional functions extracted by the tool is similar to ones by the specialist though for the classification of each function there is some difference between them.
Shinji Kusumoto, Masahiro Imagawa, Katsuro Inoue, Shuuma Morimoto, Kouji Matsusita, Michio Tsuda
ICSE3
2002 On Software Maintenance Process Improvement Based on Code Clone Analysis
Yoshiki Higo, Yasushi Ueda, Toshihiro Kamiya, Shinji Kusumoto, Katsuro Inoue
PROFES5
2002 Experimental Evaluation of Program Slicing for Fault Localization
Shinji Kusumoto, Akira Nishimatsu, Keisuke Nishie, Katsuro Inoue
Empir. Softw. Eng.4
2002 Architectural styles for distributed processing systems and practical selection method
Yoshitomi Morisawa, Katsuro Inoue, Koji Torii
Inf. Softw. Technol.2
2002 An information-leak analysis system based on program slicing
Reishi Yokomori, Fumiaki Ohata, Yoshiaki Takata, Hiroyuki Seki, Katsuro Inoue
Inf. Softw. Technol.5
2002 CCFinder: A Multilinguistic Token-Based Code Clone Detection System for Large Scale Source Code
abstract
A code clone is a code portion in source files that is identical or similar to another. Since code clones are believed to reduce the maintainability of software, several code clone detection techniques and tools have been proposed. This paper proposes a new clone detection technique, which consists of the transformation of input source text and a token-by-token comparison. For its implementation with several useful optimization techniques, we have developed a tool, named CCFinder (Code Clone Finder), which extracts code clones in C, C++, Java, COBOL and other source files. In addition, metrics for the code clones have been developed. In order to evaluate the usefulness of CCFinder and metrics, we conducted several case studies where we applied the new tool to the source code of JDK, FreeBSD, NetBSD, Linux, and many other systems. As a result, CCFinder has effectively found clones and the metrics have been able to effectively identify the characteristics of the systems. In addition, we have compared the proposed technique with other clone detection techniques.
Toshihiro Kamiya, Shinji Kusumoto, Katsuro Inoue
IEEE Trans. Software Eng.3
2001 A Slicing Method for Object-Oriented Programs Using Lightweight Dynamic Information
abstract
Program slicing has been used for efficient program debugging activities. A program slice is computed by analyzing dependence relations between program statements. We can divide dependence analyses into two categories, static and dynamic; the former requires small analysis costs, but the resulting slices are large, and in the latter the cost is high but the slices are small. In this paper, we propose a program slicing method for object-oriented programs and evaluate its effectiveness with Java programs. Since object-oriented languages have many dynamically determined elements, static analysis could not compute practical analysis results. Our method uses static and dynamic analyses appropriately and computes accurate slices with small costs.
Fumiaki Ohata, Kouya Hirose, Masato Fujii, Katsuro Inoue
APSEC4
2001 An Efficient Information Flow Analysis of Recursive Programs Based on a Lattice Model of Security Classes
Shigeta Kuninobu, Yoshiaki Takata, Hiroyuki Seki, Katsuro Inoue
ICICS4
2001 Maintenance Support Tools for JAVA Programs: CCFinder and JAAT
abstract
This paper describes two software maintenance support tools, CCFinder (Code Clone Finder) and JAAT (Java Alias Analysis Tool), for Java programs. CCFinder identifies code clones in Java programs, while JAAT executes alias analysis for Java programs.
Toshihiro Kamiya, Fumiaki Ohata, Kazuhiro Kondou, Shinji Kusumoto, Katsuro Inoue
ICSE5
2001 Evaluation of a Business Application Framework Using Complexity and Functionality Metrics
Hikaru Fujiwara, Shinji Kusumoto, Katsuro Inoue, Toshifusa Ootsubo, Katsuhiko Yuura
PROFES3
2001 Function-point analysis using design specifications based on the Unified Modelling Language
abstract
Abstract Function‐point analysis was introduced to help measure the functionality of software systems. For more than a decade, function points have been widely used to measure the size of information systems, often as a part of estimating the effort required for software development and maintenance processes. Limiting the use of function point measurement have been concerns about variable judgements on the part of the personnel doing the measurement, yielding differences in function‐point measures for the same software product even in the same organization. Also, if an organization tries to introduce function‐point analysis, the process normally starts with measurements from the organization's own past software products—a time consuming task with start‐up costs. In this paper, we propose detailed function‐point analysis measurement rules using design specifications based on the Unified Modelling Language and describe a function‐point measurement tool, whose inputs are design specifications developed on Rational Rose®. Then in this paper, we report tool validation work on software involved in software evolution at an organization where we have applied the tool to actual design specifications and examined the differences between the function point values obtained by the tool and those of an experienced function point measurement specialist at the organization. Copyright © 2001 John Wiley & Sons, Ltd.
Takuya Uemura, Shinji Kusumoto, Katsuro Inoue
J. Softw. Maintenance Res. Pract.3
2000 Button Selection for General GUIs Using Eye and Hand Together
abstract
This paper proposes an efficient technique for eye gaze interface suitable for the general GUI environments such as Microsoft Windows. Our technique uses an eye and a hand together: the eye for moving cursors onto the GUI button (move operation), and the hand for pushing the GUI button (push operation). We also propose the following two techniques to assist the move operation: (1) Automatic adjustment and (2) Manual adjustment. In the automatic adjustment, the cursor automatically moves to the closest GUI button when we push a mouse button. In the manual adjustment, we can move the cursor roughly by an eye, then move it a little more by the mouse onto the GUI button. In the experiment to evaluate our method, GUI button selection by manual adjustment showed better performance than the selection by a mouse even in the situation that has many small GUI buttons placed very closely each other on the GUI.
Masatake Yamamoto, Akito Monden, Ken-ichi Matsumoto, Katsuro Inoue, Koji Torii
Advanced Visual Interfaces4
2000 Function Point Measurement for Object-Oriented Requirements Specification
abstract
Function point analysis (FPA) was proposed to help measure the size of software systems and has been widely used in actual software development. However, it has been reported that since function point counting involves judgment on the part of the counter, some difference for the same product would be caused even in the same organization. The paper describes an actual experience of applying FPA to requirements specification at Hitachi Ltd. The authors propose a detailed FPA measurement method for the requirements specification analyzed using the requirements analysis system REQUARIO developed by Hitachi Ltd., and develop a measurement tool based on the method. They have also applied the tool to the actual requirements specification and show the applicability of the tool.
Shinji Kusumoto, Katsuro Inoue, Takashi Kasimoto, Ayane Suzuki, Katsuhiko Yuura, Michio Tsuda
COMPSAC2
2000 A Practical Method for Watermarking Java Programs
abstract
Java programs distributed through the Internet are now suffering from program theft. This is because Java programs can be easily decomposed into reusable class files and even decompiled into source code by program users. We propose a practical method that discourages program theft by embedding Java programs with a digital watermark. Embedding a program developer's copyright notation as a watermark in Java class files will ensure the legal ownership of class files. Our embedding method is indiscernible by program users, yet enables us to identify an illegal program that contains stolen class files. The result of the experiment to evaluate our method showed most of the watermarks (20 out of 23) embedded in class files survived two kinds of attacks that attempt to erase watermarks: an obfuscactor attack, and a decompile-recompile attack.
Akito Monden, Hajimu Iida, Ken-ichi Matsumoto, Koji Torii, Katsuro Inoue
COMPSAC5
2000 Modeling and Analysis of Software Aging Process
Akito Monden, Shin-ichi Sato, Ken-ichi Matsumoto, Katsuro Inoue
PROFES4
2000 Accumulative versioning file system Moraine and its application to metrics environment MAME
abstract
It is essential to manage versions of software products created during software development. There are various versioning tools actually used in these days, although most of them require the developers to issue management commands for consistent versioning. In this paper, we present a novel versioning file system Moraine, which accumulatively and automatically collects all files created or modified. Those files are versioned and stored as compressed forms. The older versions are easily retrieved from Moraine by the time-stamps or tags if required.
Tetsuo Yamamoto, Makoto Matsushita, Katsuro Inoue
SIGSOFT FSE3
2000 Analyzing dependence locality for efficient construction of program dependence graph
Fumiaki Ohata, Akira Nishimatsu, Katsuro Inoue
Inf. Softw. Technol.3
1999 Slicing Methods Using Static and Dynamic Analysis Information
abstract
In this paper, we propose four slicing methods using both static and dynamic analysis information. (1) Statement-mark slice removes the unnecessary statements using an execution history of the statements. (2) Partial program analysis reduces the static analysis cost using invocation history of procedures. (3) Dynamic data dependence analysis extracts precise data dependence relations using dynamic data dependence analysis. (4) Array and pointer analysis improves the efficiency of (3) by dynamically analyzing pointer and array variables only. Using both dynamic and static information, we show that the precision of the slicing is improved with smaller run-time overhead.
Yoshiyuki Ashida, Fumiaki Ohata, Katsuro Inoue
APSEC3
1999 Factor Analysis of Comprehension States in the Learning Phases of a Programming Language
abstract
Presents an experiment in understanding how learners of the Java programming language comprehend its concepts, such as classes, inheritance, interfaces, etc., in lectures and exercises. The authors used an empirical technique to test conjectures about how we learn the programming language. Usually, observations about how we learn a programming language are treated anecdotally. In this experiment, learners received lectures and did an exercise. The comprehension states of the learners were measured by tests in three learning phases. The first phase was before the lecture. In this phase, the learners had no knowledge of the programming language. The second phase was after the lecture and before the exercise. Here, the learners acquired some basic knowledge. The third phase was after the exercise. In this phase, the learners put the acquired knowledge into practice. Factor analysis was used to obtain factors affecting the test result of each learning phase. Changes in comprehension states are explained as a result of tracing the factors between the learning phases.
Yasuhiro Takemura, Kazuyuki Shima, Ken-ichi Matsumoto, Katsuro Inoue, Koji Torii
APSEC4
1999 Call-Mark Slicing: An Efficient and Economical Way of Reducing Slice
abstract
Article Call-mark slicing: an efficient and economical way of reducing slice Share on Authors: Akira Nishimatsu Graduate School of Engineering Science, Osaka University, 1-3 Machikaneyama, Toyonaka, Osaka 560-8531, Japan Graduate School of Engineering Science, Osaka University, 1-3 Machikaneyama, Toyonaka, Osaka 560-8531, JapanView Profile , Minoru Jihira Graduate School of Information Science, Nara Institute of Science and Technology, 8916-5, Takayama, Ikoma, Nara 630-0101, Japan Graduate School of Information Science, Nara Institute of Science and Technology, 8916-5, Takayama, Ikoma, Nara 630-0101, JapanView Profile , Shinji Kusumoto Graduate School of Engineering Science, Osaka University, 1-3 Machikaneyama, Toyonaka, Osaka 560-8531, Japan Graduate School of Engineering Science, Osaka University, 1-3 Machikaneyama, Toyonaka, Osaka 560-8531, JapanView Profile , Katsuro Inoue Graduate School of Engineering Science, Osaka University, 1-3 Machikaneyama, Toyonaka, Osaka 560-8531, Japan Graduate School of Engineering Science, Osaka University, 1-3 Machikaneyama, Toyonaka, Osaka 560-8531, JapanView Profile Authors Info & Claims ICSE '99: Proceedings of the 21st international conference on Software engineeringMay 1999 Pages 422–431https://doi.org/10.1145/302405.302674Online:16 May 1999Publication History 26citation284DownloadsMetricsTotal Citations26Total Downloads284Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Akira Nishimatsu, Minoru Jihira, Shinji Kusumoto, Katsuro Inoue
ICSE4
1999 Prediction of Fault-proneness at Early Phase in Object-Oriented Development
abstract
To analyse the complexity of object-oriented software, several metrics have been proposed. Among them, Chidamber and Kemerer's (1994) metrics are well-known object-oriented metrics. Also, their effectiveness has been empirically evaluated from the viewpoint of estimating the fault-proneness of object-oriented software. In the evaluations, these metrics were applied, not to the design specification but to the source code, because some of them measure the inner complexity of a class, and such information cannot be obtained until the algorithm and the class structure are determined at the end of the design phase. However, the estimation of the fault-proneness should be done in the early phase so as to effectively allocate effort for fixing the faults. This paper proposes a new method to estimate the fault-proneness of an object class in the early phase, using several complexity metrics for object-oriented software. In the proposed method, we introduce four checkpoints into the analysis/design/implementation phase, and we estimate the fault-prone classes using applicable metrics at each checkpoint.
Toshihiro Kamiya, Shinji Kusumoto, Katsuro Inoue
ISORC3
1999 Empirical evaluation of reuse sensitiveness of complexity metrics
Toshihiro Kamiya, Shinji Kusumoto, Katsuro Inoue, Yukio Mohri
Inf. Softw. Technol.3
1997 Conceptual Issues of an Object-Centered Process Model
abstract
We propose an object-centered software process description model. We also present the idea of a software development management environment based on the model. To use this model and environment, we illustrate the software development environment as it is, and provide a framework for software process description, management and improvement.
Makoto Matsushita, Makoto Oshita, Hajimu Iida, Katsuro Inoue
APSEC4
1996 An Interaction Support Mechanism in Software Development
abstract
The paper proposes a new modeling method of interactions in the software development process, which focuses on the interactions among the elements of the process, and a new software development environment based on the model. In this method, interactions in the software process are modeled as a set of agents and communication channels. An agent interacts with other agents with channels. Channels are classified according to their content and type of interaction. A prototype of the supporting environment for software development which is based on the model is also developed. The environment consists of a proxy program for the agent and integrated communication server, which provides mechanisms for interaction, process execution, and user navigation.
Makoto Matsushita, Katsuro Inoue, Hajimu Iida
APSEC2
1993 Process-centered project management system by stepwise particularizing software process
abstract
This article proposes a process-centered project management system to facilitate controlling a wide variety of project activities. The essential idea is to employ software process descriptions as vehicles of exchanging activity plans between a project manager and a large number of staff. The system allows a manager to plan and monitor overall project processes which are stepwise particularized according to actual processes dynamically determined by individual staff. By the cooperative process description and stepwise particularization of software processes, a wide variety of software project activities can be coordinated in advance to lead the project to success. This article also proposes a graphical process representation method which maps software processes into a directed graph.>
Kagetomo Genji, Michitoshi Ishiwaka, Takeshi Ogihara, Katsuro Inoue
COMPSAC4
1991 Generating software development environments from the description of product relations
abstract
A method is described for constructing software development support system from the description of software product relations. Logical structures of the products have ben expressed by a tree structure. Each product appearing through the software development corresponds to a leaf in a tree and a set of those products and other sets correspond to an internal node. Various kinds of software development are easily expressed by changing the products at the leaves. The description of product relation is translated into a script of process description language (PDL). A support system is obtained by executing the translated PDL script with the PDL interpreter which the authors have developed.>
Hajimu Iida, Yoshihiro Nishimura, Katsuro Inoue, Koji Torii
COMPSAC3
1991 Functional language for enacting software processes
abstract
In order to define software processes formally and to use the defined processes in actual development situations, the authors have designed the process description language (PDL) and an associated PDL System. PDL is a functional programming language based on an algebraic specification language, and development processes may be defined in PDL at various levels of abstraction. Abstract PDL scripts (programs) specify the general course of process execution. Essential features of the language for process description are discussed. The characteristics of the PDL System are also presented.>
Katsuro Inoue, Takeshi Ogihara, Hajimu Iida, Minoru Nitta
COMPSAC1
1989 A Formal Adaptation Method for Process Descriptions
abstract
Requirement to describing software development processes in formal manners has been increased, and demand for altering and tailoring the process descriptions has been emerged.In this paper, we propose a functional language PDL (Process Description Language), designed to describe various development processes under a certain environment.To create and modify the PDL scripts easily and correctly, we propose a method of stepwise refinement from abstract scripts into concrete scripts.By this method, the abstract definitions of software process flow and product flow initially given as function definitions in PDL, are transformed into the concrete definitions of the tool activations, message displays, and so on.We also discuss an architecture design of a software development environment (Adaptable Software Development Environment), which can be adapted in many ways for designer's requirements.
Katsuro Inoue, Takeshi Ogihara, Tohru Kikuno, Koji Torii
ICSE1
1989 TERM: a parallel executable graph reduction machine for equational language
Kazuhito Ohmaki, Satoru Tomura, Katsuro Inoue, Toshio Ito, Keiichi Ito, Koji Torii
Parallel Comput.3
1988 Analysis of Functional Programs to Detect Run-Time Garbage Cells
abstract
We propose a method for detecting the generation of garbage cells by analyzing a source text written in a functional programming language which uses ordinary linked lists to implement list-type values. For a subexpression such as F ( G ( . . . )) in a program where the function values of F and G are of list type, if a cell c is created during the computation of G and if c does not appear in a list-type value of F , then c becomes a garbage cell at the end of the computation of F . We discuss this problem on the basis of formal languages derived from the functional program text and show some sufficient conditions that predict the generation of garbage cells. Also, we give an efficient algorithm to detect at compile time the generation of garbage cells which are linearly linked. We have implemented these algorithms in an experimental LISP system. By executing several sample programs on the system, we conclude that our method is effective in detecting the generation of garbage cells.
Katsuro Inoue, Hiroyuki Seki, Hikaru Yagi
ACM Trans. Program. Lang. Syst.1
1986 Compiling and Optimizing Methods for the Functional Language ASL/F
Katsuro Inoue, Hiroyuki Seki, Kenichi Taniguchi, Tadao Kasami
Sci. Comput. Program.1