Norihiro Yoshida

dblp:85/877 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-4910-1729ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 29 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021
YearPublicationVenuePosition
2026 Why Are Agentic Pull Requests Merged or Rejected? An Empirical Study
Sien Reeve Ordonez Peralta, Fumika Hoshi, Hironori Washizaki, Naoyasu Ubayashi, Inase Kondo, Yoshiki Higo, Hiroki Mukai, Norihiro Yoshida, Kazuki Kusama, Hidetake Tanaka, Youmei Fan
MSR8
2025 BiFuzz: A Two-Stage Fuzzing Tool for Open-World Video Games
abstract
Open-world video games present a broader search space than other video games, posing challenges for test automation. Fuzzing, which generates new inputs by mutating an initial input, is commonly used to uncover issues. In this study, we proposed BiFuzz, a two-stage fuzzer designed for automated testing of open-world video games, and investigated its effectiveness. The results revealed that BiFuzz mutated the overall strategy of gameplay and test cases, including actual movement paths, step by step. Consequently, BiFuzz can detect character stuck issues. The tool and its video are at https://github.com/ Yusaku-Kato/BiFuzz and https://www.youtube.com/watch? $\mathbf{v}=$ VOrHfnLJSbk. Index Terms-open-world video game, fuzzing
Yusaku Kato, Norihiro Yoshida, Erina Makihara, Katsuro Inoue
APSEC2
2025 Development and benchmarking of multilingual code clone detector
Norihiro Yoshida, Toshihiro Kamiya, Eunjong Choi, Hiroaki Takada
J. Syst. Softw.2
2023 Cost-Benefit Analysis for Modernizing a Large-Scale Industrial System
abstract
Legacy systems pose significant challenges to companies. Software modernization approaches have been proposed to address this issue. However, a lack of standardization and reliance on ad hoc processes often lead to software modernization failures. Incremental modernization, a strategy that improves software systems in a step-by-step manner rather than attempting to simultaneously overhaul the entire system, aims to mitigate the risk of failure. However, this approach can increase costs owing to the complexity of integrating legacy and modernized products. In this paper, we present a case study that employs a cost-benefit estimation analysis in a large-scale industrial project that underwent incremental modernization in the past. We compare the actual and estimated cost-benefit values in the context of incremental modernization. As a result, we confirmed that the cost estimates were valid, but we could not judge whether the benefit estimates were valid.
Kazuki Yokoi, Eunjong Choi, Norihiro Yoshida, Joji Okada, Yoshiki Higo
APSEC3
2023 Investigating the Generalizability of Deep Learning-based Clone Detectors
abstract
The generalizability of Deep Learning (DL) models is a significant challenge, as poor generalizability indicates that the model has overfitted to the training data and is not able to generalize to new data. Despite numerous DL-based clone detectors emerging in recent years, their generalizability has not been thoroughly assessed. This study investigates the generalizability of three DL-based clone detectors (CCLearner, ASTNN, and CodeBERT) by comparing their detection accuracy on different training and testing clone benchmarks. The results show that all three clone detectors do not generalize well to new data and there is a strong relationship between clone types and generalizability for CCLearner and ASTNN.
Eunjong Choi, Norihiro Fuke, Yuji Fujiwara, Norihiro Yoshida, Katsuro Inoue
ICPC4
2022 MSCCD: grammar pluggable clone detection based on ANTLR parser generation
abstract
For various reasons, programming languages continue to multiply and evolve. It has become necessary to have a multilingual clone detection tool that can easily expand supported programming languages and detect various code clones is needed. However, research on multilingual code clone detection has not received sufficient attention. In this study, we propose MSCCD (Multilingual Syntactic Code Clone Detector), a grammar pluggable code clone detection tool that uses a parser generator to generate a code block extractor for the target language. The extractor then extracts the semantic code blocks from a parse tree. MSCCD can detect Type-3 clones at various granularities. We evaluated MSCCD's language extensibility by applying MSCCD to 20 modern languages. Sixteen languages were perfectly supported, and the remaining four were provided with the same detection capabilities at the expense of execution time. We evaluated MSCCD's recall by using BigCloneEval and conducted a manual experiment to evaluate precision. MSCCD achieved equivalent detection performance equivalent to state-of-the-art tools.
Norihiro Yoshida, Toshihiro Kamiya, Eunjong Choi, Hiroaki Takada
ICPC2
2021 Extracting a Micro State Transition Table Using the KLEE Symbolic Execution Engine
abstract
In this paper, we suggest an approach for extracting fine-grained state transition tables using the KLEE symbolic execution engine to assist developers in understanding the behavior of C source code for embedded systems.
Norihiro Yoshida, Takahiro Shimizu, Ryota Yamamoto, Hiroaki Takada
APSEC1
2020 Understanding Build Errors in Agile Software Development Project-Based Learning
abstract
Recently, various institutions have been conducting advanced programming education aimed at experiencing agile software development in the form of project-based learning (PBL). In the agile software development model, an essential part is the build process. In this study, we investigated students' build behaviors in agile software development PBL (SDPBL) by monitoring and collecting logs of the build process from 2013 to 2016. In our investigation, we collected two types of logs, the local build logs collected by each student's build in their own local programming environment, and remote build logs collected by any team member's commit in team repository. Based on our analysis of the build logs from 2013 to 2015, we found that the causes of remote build errors are related to both technical factors and communications among students in a team. In 2016, the instructors tried to educate to students the reason why the remote build error occur and how it can be resolved. As a result, in 2016, the number of remote build errors and the time required to solve the build errors decreased compared to previous years. It indicates a possibility that the student's comprehension of build error is effective on improving quality of software product and team development.
Erina Makihara, Hiroshi Igaki, Norihiro Yoshida, Kenji Fujiwara, Hajimu Iida
APSEC3
2020 Clone Notifier: Developing and Improving the System to Notify Changes of Code Clones
abstract
A code clone is a code fragment that is identical or similar to it in the source code. It has been identified as one of the main problems in software maintenance. When a developer fixes a defect, they need to find the code clones corresponding to the code fragments. In this paper, we present Clone Notifier, a system that alerts on creations and changes of code clones to software developers. First, Clone Notifier identifies creations and changes of code clones. Subsequently, it groups them into four categories (new, deleted, changed, stable) and assigns labels (e.g., consistent, inconsistent) to them. Finally, it notifies on creations and changes of code clones along with the corresponding categories and labels. Clone Notifier and its video are available at: https://github.com/s-tokui/CloneNotifier.
Shogo Tokui, Norihiro Yoshida, Eunjong Choi, Katsuro Inoue
SANER2
2019 CCEvovis: a clone evolution visualization system for software maintenance
abstract
Understanding the evolution of code clones is important in software maintenance. With the information about how code clones evolve, both developers and researchers can understand the impacts of code clones and build a more robust code clone management system. So far, many studies have investigated the evolution of code clones to better understand the effects of code clones. However, only a few systems have been presented to support managing code clones based on the information about how code clone evolves. To mitigate this problem, in this paper, we present CCEvovis, a system that visualizes the evolved code clones across multiple versions of a program. CCEvovis highlights and visualizes the clone change to support software maintenance. CCEvovis is available at: https://github.com/hirotaka0616/CCEvovis.
Hirotaka Honda, Shogo Tokui, Kazuki Yokoi, Eunjong Choi, Norihiro Yoshida, Katsuro Inoue
ICPC5
2018 An Investigation of the Relationship between Extract Method and Change Metrics: A Case Study of JEdit
abstract
Extract Method is one of the most widely used refactoring patterns. So far, low quality of source code has been regarded as an indicator for Extract Method opportunities. However, recent studies showed that there is no clear relationship between source code quality and Extract Method. Change metrics can be indicators for Extract Method because the characteristics of software evolution strongly affect software quality. However, there has been no study that investigated the relationship between change metrics and Extract Method. In this study, we conducted two studies investigating the relationship between Extract Method and change metrics. As a result, we found that (1) change metrics have a clear relationship with Extract Method and (2) both product and change metrics are necessary to recommend candidates for Extract Method with high accuracy.
Eunjong Choi, Daiki Tanaka, Norihiro Yoshida, Kenji Fujiwara, Daniel Port, Hajimu Iida
APSEC3
2018 Multilingual Detection of Code Clones Using ANTLR Grammar Definitions
abstract
So far, many tools have been developed for the detection of code clones in source code. The existing clone detection tools support only a limited number of programming languages and do not provide any easy extension mechanism to handle additional language. However, from our experience in industry/university collaboration, we found that many practitioners need to analyze source code written in various languages. In this paper, we propose an approach for the multilingual detection of code clones using grammar files for a parser generator ANTLR. We extended a clone detection tool CCFinderSW with the proposed approach and then apply the extended CCFinderSW to ANTLR grammar files for 43 languages. As a result, the files for 39 out of the 43 languages can be analyzed correctly by the extended CCFinderSW.
Yuichi Semura, Norihiro Yoshida, Eunjong Choi, Katsuro Inoue
APSEC2
2018 Investigating Vector-Based Detection of Code Clones Using BigCloneBench
abstract
In a vector-based approach to detecting code clones from source code, all code fragments in the source are mapped to a vector space and then code fragments are detected as code clones if they are neighbors in the vector space. So far, our research group has developed a vector-based approach using TF-IDF and cosine similarity. For the improvement of the vector-based approach, we preliminary investigated what kind of vectorization algorithms and similarity measurements are effective in terms of recall and detection time. In this paper, we present preliminary investigation results using BigCloneBench, a large-scale code clone benchmark.
Kazuki Yokoi, Eunjong Choi, Norihiro Yoshida, Katsuro Inoue
APSEC3
2018 How slim will my system be?: estimating refactored code size by merging clones
abstract
We have been doing code clone analysis with industry collaborators for a long time, and have been always asked a question, "OK, I understand my system contains a lot of code clones, but how slim will it be after merging redundant code clones?" As a software system evolves for long period, it would increasingly contain many code clones due to quick bug fix and new feature addition. Industry collaborators would recognize decay of initial design simplicity, and try to evaluate current system from the view point of maintenance effort and cost. As one of resources for the evaluation, the estimated code size by merging code clone is very important for them. In this paper, we formulate this issue as "slimming" problem, and present three different slimming methods, Basic, Complete, and Heuristic Methods, each of which gives a lower bound, upper bound, and modest reduction rates, respectively. Application of these methods to OSS systems written in C/C++ showed that the reduction rate is at most 5.7% of the total size, and to a commercial COBOL system, it is at most 15.4%. For this approach, we have gotten initial but very positive feedback from industry collaborators.
Norihiro Yoshida, Takuya Ishizu, Bufurod Edwards, Katsuro Inoue
ICPC1
2017 CCFinderSW: Clone Detection Tool with Flexible Multilingual Tokenization
abstract
So far, many tools have been developed for the detection of code clones in source code. The existing clone detection tools support only a limited number of programming languages and do not provide any easy extension mechanism to handle additional language. However, from our experience in industry/university collaboration, we found that many practitioners need to analyze source code written in various languages. In this paper, we propose a clone detection tool CCFinderSW that has extension mechanism to handle addition language on demand from practitioners.
Yuichi Semura, Norihiro Yoshida, Eunjong Choi, Katsuro Inoue
APSEC2
2017 Frame-based behavior preservation in refactoring
abstract
Behavior preservation often bothers programmers in refactoring. This poster paper proposes a new approach that tames the behavior preservation by introducing the concept of a frame. A frame in refactoring defines stakeholder's individual concerns about the refactored code. Frame-based refactoring preserves the observable behavior within a particular frame. Therefore, it helps programmers distinguish the behavioral changes that they should observe from those that they can ignore.
Katsuhisa Maruyama, Shinpei Hayashi, Norihiro Yoshida, Eunjong Choi
SANER3
2016 Detecting exploratory programming behaviors for introductory programming exercises
abstract
Developers often perform the repeating cycle of implementation and evaluation when they need to deal with the unfamiliar portion of the source code. This cycle is named as exploratory programming. We regard exploratory programming as an effective way not only to improve novice's programming skill but also to support educators in programming exercise in University. Because when novices often use the exploratory programming, it means novices struggle to solve their assignments. Therefore, educators should grasp which elements, APIs or blocks novices often used exploratory programming for. In this paper, firstly we propose the definition of novice's exploratory programming to collect logs of exploratory based on various granularity by novices. Secondly, we propose an algorithm based on our proposed definition to automatically detect exploratory programming behaviors. We also conducted a small case study. As a result of automatic detection, our proposed algorithm allows us to know what elements of program novices often feel difficult and struggle for.
Erina Makihara, Hiroshi Igaki, Norihiro Yoshida, Kenji Fujiwara, Hajimu Iida
ICPC3
2016 Revisiting the relationship between code smells and refactoring
abstract
Refactoring is a critical technique in evolving software systems. Martin Fowler presented a catalogue of refactoring patterns that defines a list of code smells and their corresponding refactoring patterns. This list aimed at supporting programmers in finding suitable refactoring patterns that remove code smells from their systems. However, a recent empirical study by Bavota et al. shows that refactoring rarely removes code smells which do not align with Fowler's catalog. To bridge the gap between them, we revisit the relationship between code smells and refactorings. In this study, we investigate whether developers apply appropriate refactoring patterns to fix code smells in three open source software systems.
Norihiro Yoshida, Tsubasa Saika, Eunjong Choi, Ali Ouni 0001, Katsuro Inoue
ICPC1
2016 Mining the modern code review repositories: a dataset of people, process and product
abstract
In this paper, we present a collection of Modern Code Review data for five open source projects. The data showcases mined data from both an integrated peer review system and source code repositories. We present an easy-to-use and richer data structure to retrieve the (a) People, (b) Process, and (c) Product aspects of the peer review. This paper presents the extraction methodology, the dataset structure, and a collection of database dumps.
Xin Yang 0018, Raula Gaikovina Kula, Norihiro Yoshida, Hajimu Iida
MSR3
2016 On the Effectiveness of Vector-Based Approach for Supporting Simultaneous Editing of Software Clones
Seiya Numata, Norihiro Yoshida, Eunjong Choi, Katsuro Inoue
PROFES2
2015 Who should review my code? A file location-based code-reviewer recommendation approach for Modern Code Review
abstract
Software code review is an inspection of a code change by an independent third-party developer in order to identify and fix defects before an integration. Effectively performing code review can improve the overall software quality. In recent years, Modern Code Review (MCR), a lightweight and tool-based code inspection, has been widely adopted in both proprietary and open-source software systems. Finding appropriate code-reviewers in MCR is a necessary step of reviewing a code change. However, little research is known the difficulty of finding code-reviewers in a distributed software development and its impact on reviewing time. In this paper, we investigate the impact of reviews with code-reviewer assignment problem has on reviewing time. We find that reviews with code-reviewer assignment problem take 12 days longer to approve a code change. To help developers find appropriate code-reviewers, we propose RevFinder, a file location-based code-reviewer recommendation approach. We leverage a similarity of previously reviewed file path to recommend an appropriate code-reviewer. The intuition is that files that are located in similar file paths would be managed and reviewed by similar experienced code-reviewers. Through an empirical evaluation on a case study of 42,045 reviews of Android Open Source Project (AOSP), OpenStack, Qt and LibreOffice projects, we find that RevFinder accurately recommended 79% of reviews with a top 10 recommendation. RevFinder also correctly recommended the code-reviewers with a median rank of 4. The overall ranking of RevFinder is 3 times better than that of a baseline approach. We believe that RevFinder could be applied to MCR in order to help developers find appropriate code-reviewers and speed up the overall code review process.
Patanamon Thongtanunam, Chakkrit Tantithamthavorn, Raula Gaikovina Kula, Norihiro Yoshida, Hajimu Iida, Ken-ichi Matsumoto
SANER4
2014 ReDA: A Web-Based Visualization Tool for Analyzing Modern Code Review Dataset
abstract
ReDA (http://reda.naist.jp/) is a web-based visualization tool for analyzing Modern Code Review (MCR) datasets for large Open Source Software (OSS) projects. MCR is a commonly practiced and lightweight inspection of source code using a support tool such as Gerrit system. Recently, mining code review history of such systems has received attention as a potentially effective method of ensuring software quality. However, due to increasing size and complexity of softwares being developed, these datasets are becoming unmanageable. ReDA aims to assist researchers of mining code review data by enabling better understand of dataset context and identifying abnormalities. Through real-time data interaction, users can quickly gain insight into the data and hone in on interesting areas to investigate. A video highlighting the main features can be found at: http://youtu.be/ fEoTRRas0U.
Patanamon Thongtanunam, Xin Yang 0018, Norihiro Yoshida, Raula Gaikovina Kula, Ana Erika Camargo Cruz, Kenji Fujiwara, Hajimu Iida
ICSME3
2013 Seamless Code Reuse with Source Code Corpus
abstract
Code reuse is attracting much attention as a promising technique for efficient software development. However, code reuse itself requires human resources: for example, searching and opening source files including code fragments that users would like to reuse, or considering keywords in using code search systems. The present paper proposes a novel technique that hardly requires such reuse cost. In the proposed technique, what programmers have to do for obtaining reusable code is just inputting a trigger key for code reuse on their development environments. Also, this paper describes some applications on OSS with a prototype tool working on Eclipse.
Tetsuo Yamamoto, Norihiro Yoshida, Yoshiki Higo
APSEC (2)2
2013 Applying clone change notification system into an industrial development process
abstract
Programmers tend to write code clones unintentionally even in the case that they can easily avoid them. Clone change management is one of crucial issues in open source software (OSS) development as well as in industrial software development (e.g., development of social infrastructure, financial system, and medical equipment). When an industrial developer fixes a defect, he/she has to find the code clones corresponding to the code fragment including it. So far, several studies performed on the analysis of clone evolution in OSS. However, to our knowledge, a few researches have been reported on an application of a clone change notification system to industrial development process. In this paper, we introduce a system for notifying creation and change of code clones, and then report on the experience with 40-days application of it into a development process in NEC Corporation. In the industrial application, a developer successfully identified ten unintentionally-developed clones that should be refactored.
Yuki Yamanaka, Eunjong Choi, Norihiro Yoshida, Katsuro Inoue, Tateki Sano
ICPC3
2013 Who does what during a code review? datasets of OSS peer review repositories
abstract
We present four datasets that are focused on the general roles of OSS peer review members. With data mined from both an integrated peer review system and code source repositories, our rich datasets comprise of peer review data that was automatically recorded. Using the Android project as a case study, we describe our extraction methodology, the datasets and their application used for three separate studies. Our datasets are available online at http://sdlab.naist.jp/reviewmining/.
Kazuki Hamasaki, Raula Gaikovina Kula, Norihiro Yoshida, Ana Erika Camargo Cruz, Kenji Fujiwara, Hajimu Iida
MSR3
2013 Assessing Refactoring Instances and the Maintainability Benefits of Them from Version Archives
Kenji Fujiwara, Kyohei Fushida, Norihiro Yoshida, Hajimu Iida
PROFES3
2013 Micro process analysis of maintenance effort: an open source software case study using metrics based on program slicing
abstract
SUMMARY For any software project, most experts regard the maintenance phase as the most effort and cost intensive of all phases in the software development life cycle. This is due to the highmaintenance effort, time, and resources needed to effectively address issues during software maintenance (maintenance activities). Mismanagement of these efforts can lead to the degradation of software maintainability. Understanding the assessment of the related software processes can help sustain or improve maintainability during these maintenance activities. Recent studies have shown that current software process assessments are expensive, generic, and complex, especially for smaller organizations. In this paper, we investigate an alternative software process assessment approach performed by analyzing fine‐grained processes (micro processes) of maintenance activities. This approach assesses maintenance efforts based on micro processes in relation to their impact on source code. The approach derives maintenance effort from the complexity and duration of micro processes and uses proposed metrics based on program slicing to measure change impact. In this paper, we investigate an alternative software process assessment approach by analysing fine‐grained processes (micro processes) of maintenance activities. At statistically significant levels, results suggest that the level of the maintenance efforts correlates with its impact on source code. Copyright © 2012 John Wiley & Sons, Ltd.
Raula Gaikovina Kula, Kyohei Fushida, Norihiro Yoshida, Hajimu Iida
J. Softw. Evol. Process.3
2012 Understanding OSS Peer Review Roles in Peer Review Social Network (PeRSoN)
abstract
Due to the distributed collaborations and the volunteering nature of Open Source Software (OSS), OSS peer review processes differs from traditional approaches. Despite the latest research efforts to understand OSS peer review processes, very little is known. Unlike related work, this study investigates OSS peer review processes from a different perspective. We investigate the importance of OSS peer review contributor roles and their review activities by using social network analysis (SNA), proposed as PeRSoN (Peer Review Social Network). As a case study, we extracted and analyzed the review process of Android Open Source Project (AOSP). To the best of our knowledge, this is the first research constructing social networks from mining a peer review repository. Our preliminary results provided hints on relationships among the OSS peer review contributor roles, their activities, and the network structure. The results raised issues that will be used to refine our approach in the future.
Xin Yang 0018, Raula Gaikovina Kula, Ana Erika Camargo Cruz, Norihiro Yoshida, Kazuki Hamasaki, Kenji Fujiwara, Hajimu Iida
APSEC4
2012 An Experience Report on Analyzing Industrial Software Systems Using Code Clone Detection Techniques
abstract
A variety of application results of code clone detection and analysis has been reported. There are many reports of code clone detection and analysis on open source software whereas few reports on industrial systems are open to the public. This paper reports an experience of code clone analysis on a governmental project. In the project, a software system was developed by multiple Japanese vendors. We detected and analyzed code clones in the system, and found that there were many code clones in the project, however we concluded that the presence of the code clones did not have negative impacts on the maintenance of the system because of the following reasons: (1) when different modules are similar to each other in the design document, they also share many code clones in the source code, (2) code clones located in trusted modules, which are libraries maintained by one of the companies.
Norihiro Yoshida, Yoshiki Higo, Shinji Kusumoto, Katsuro Inoue
APSEC1