Hajimu Iida

dblp:44/6147 · DBLP profile ↗
← Back
70ranked-venue papers
3as first author
24since 2021 · last 2026
0000-0002-2919-6620ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 58 · 3 first-author · 20 since 2021Databases, data management, data science and information retrieval · 10 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 2 since 2021Systems, architecture and hardware · 4 · 1 since 2021Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Do AI Agents Really Improve Code Readability?
abstract
Code readability is fundamental to software quality and maintainability. Poor readability extends development time, increases bug-inducing risks, and contributes to technical debt. With the rapid advancement of Large Language Models, AI agent-based approaches have emerged as a promising paradigm for automated refactoring, capable of decomposing complex tasks through autonomous planning and execution. While prior studies have examined refactoring by AI agents, these analyses cover all forms of refactoring, including performance optimization and structural improvement. As a result, the extent to which AI agent-based refactoring specifically improves code readability remains unclear.
Kyogo Horikawa, Kosei Horikawa, Yutaro Kashiwa, Hidetake Uwano, Hajimu Iida
MSR5
2026 Does Programming Language Matter? An Empirical Study of Fuzzing Bug Detection
abstract
Fuzzing has become a popular technique for automatically detecting vulnerabilities and bugs by generating unexpected inputs. In recent years, the fuzzing process has been integrated into continuous integration workflows (i.e., continuous fuzzing), enabling short and frequent testing cycles. Despite its widespread adoption, prior research has not examined whether the effectiveness of continuous fuzzing varies across programming languages.
Tatsuya Shirai, Olivier Nourry, Yutaro Kashiwa, Kenji Fujiwara, Hajimu Iida
MSR5
2026 What to Cut? Predicting Unnecessary Methods in Agentic Code Generation
abstract
Agentic Coding, powered by autonomous agents such as GitHub Copilot and Cursor, enables developers to generate code, tests, and pull requests from natural language instructions alone. While this accelerates implementation, it produces larger volumes of code per pull request, shifting the burden from implementers to reviewers. In practice, a notable portion of AI-generated code is eventually deleted during review, yet reviewers must still examine such code before deciding to remove it. No prior work has explored methods to help reviewers efficiently identify code that will be removed.
Kan Watanabe, Tatsuya Shirai, Yutaro Kashiwa, Hajimu Iida
MSR4
2026 Testing with AI Agents: An Empirical Study of Test Generation Frequency, Quality, and Coverage
abstract
Agent-based coding tools have transformed software development practices. Unlike prompt-based approaches that require developers to manually integrate generated code, these agent-based tools autonomously interact with repositories to create, modify, and execute code, including test generation. While many developers have adopted agent-based coding tools, little is known about how these tools generate tests in real-world development scenarios or how AI-generated tests compare to human-written ones.
Suzuka Yoshimoto, Shun Fujita, Kosei Horikawa, Daniel Feitosa, Yutaro Kashiwa, Hajimu Iida
MSR6
2026 Evaluating Cross-Language Transfer for Refactoring Detection with Large Language Models
Nabhan Suwanachote, Yutaro Kashiwa, Brittany Reid, Hajimu Iida
SANER5
2026 Large-Scale Empirical Analysis of Continuous Fuzzing: Insights From 1 Million Fuzzing Sessions
abstract
Software vulnerabilities are constantly being reported and exploited in software products, causing significant impacts on society. In recent years, the main approach to vulnerability detection, fuzzing, has been integrated into the continuous integration process to run in short and frequent cycles. This continuous fuzzing allows for fast identification and remediation of vulnerabilities during the development process. Despite adoption by thousands of projects, however, it is unclear how continuous fuzzing contributes to vulnerability detection.This study aims to elucidate the role of continuous fuzzing in vulnerability detection. Specifically, we investigate the coverage and the total number of fuzzing sessions when fuzzing bugs are discovered. We collect issue reports, coverage reports, and fuzzing logs from OSS-Fuzz, an online service provided by Google that performs fuzzing during continuous integration. Through an empirical study of a total of approximately 1.12 million fuzzing sessions from 878 projects participating in OSS-Fuzz, we reveal that (i) a substantial number of fuzzing bugs exist prior to the integration of continuous fuzzing, leading to a high detection rate in the early stages; (ii) code coverage continues to increase as continuous fuzzing progresses; (iii) changes in coverage contribute to the detection of fuzzing bugs; and (iv) developer-provided seed corpus exhibit long-term effectiveness. This study provides empirical insights into how continuous fuzzing contributes to fuzzing bug detection, offering practical implications for future strategies and tool development in continuous fuzzing.
Tatsuya Shirai, Olivier Nourry, Yutaro Kashiwa, Kenji Fujiwara, Yasutaka Kamei, Hajimu Iida
IEEE Trans. Software Eng.6
2025 An Empirical Investigation into Maintenance of Load Testing Scripts
abstract
Background: Modern software systems are expected to deliver high performance under a variety of different workloads. In order to automatically verify whether a system operates correctly under specific load conditions, load testing has become a widely adopted technique. As software systems evolve, their load requirements-such as performance thresholds and usage patterns-also change, necessitating updates to load tests. Aims: This study investigates the maintenance of load testing scripts to better understand how load requirements evolve and how these changes are reflected in the tests themselves. Method: We analyzed 35 open-source software (OSS) repositories that incorporate load testing. We examined the frequency and nature of load test updates. Results: Our analysis reveals that 45.7% of the studied projects do not update their load testing scripts after initial creation. However, a small subset of projects demonstrates extensive and ongoing maintenance of these scripts. Furthermore, we identified 20 distinct update types across 5 major categories of purposes for load testing script modifications. The most frequent update type is related to “Test Maintenance”, followed by “Test Scenario Modification.” Conclusions: Our findings suggest that load testing scripts are often left unmaintained over time in many projects. The updates, when performed, serve a wide range of purposes, with test maintenance being the most frequent.
Ibuki Nakamura, Kosei Horikawa, Brittany Reid, Yutaro Kashiwa, Hajimu Iida
ESEM5
2025 How Does Test Code Differ from Production Code in Terms of Refactoring? An Empirical Study
abstract
Refactoring is a widely applied practice for improving the internal structure of source code without altering its external behavior. Researchers have proposed approaches to detect refactoring operations and investigated their impact on the code quality. However, these studies often focus on production code, paying little attention to test code. It is still unclear whether developers perform refactoring on test code in the same way or for the same purpose. To fill this gap, we first investigate the types and prevalence of refactoring applied in production and test code, and then examine whether these refactorings impact the code quality in a different way. Our results show that certain refactorings are less common in the test code. Besides, while refactoring-related changes in production and/or test code improved readability, they had limited impact on most design smells. We also find that some specific refactoring types do impact certain design smells. These findings indicate the special attention needed for test code when analyzing refactorings.
Kosei Horikawa, Yutaro Kashiwa, Bin Lin 0008, Kenji Fujiwara, Hajimu Iida
ICSME5
2025 Leveraging Context Information for Self-Admitted Technical Debt Detection
abstract
Self-Admitted Technical Debt (SATD) refers to nonoptimal software design or implementation that is acknowledged and explicitly documented in the code by developers. Detecting SATD and understanding its evolution can help developers better manage their development activities and monitor the software quality. In recent years, numerous approaches have been proposed to automatically identify SATD. However, these approaches still suffer from a high number of false positives (i.e., non-SATD comments being detected as SATD). To further advance this field, in this paper, we conduct an empirical study to evaluate the performance of the state-of-theart SATD detection tools and investigate the causes behind the false positives. By manually analyzing 135 false positive cases, we identify the main types of comments that are easily misclassified. To address this issue, we propose a new approach, CASTI, which integrates context information into CodeBERT, a pre-trained model for programming languages. Our evaluation demonstrates that CASTI can significantly reduce the false positives and that the context information does help improve the performance.
Miki Yonekura, Yutaro Kashiwa, Bin Lin 0008, Kenji Fujiwara, Hajimu Iida
ICPC5
2025 On the Evolution of Unused Dependencies in Java Project Releases: An Empirical Study
abstract
Modern software development heavily relies on third-party dependencies to reduce workload and improve developer productivity. Given the vast number of dependencies available and the ease of including them in projects, some introduced dependencies are never used, leading to bloated software, longer build times, and increased network bandwidth usage. While several previous studies have examined the prevalence of unused dependencies and their impact on security, it remains unclear how these dependencies are introduced and removed in software projects. This study aims to answer this question through an empirical study involving 3,020 release versions of 417 Java projects. Our analysis shows that unused packages are common in most projects ($52 \%$ of projects), but few releases (9%) introduce new unused dependencies. Among those resolved unused dependencies, $59 \%$ of them were removed and $41 \%$ were used in later versions. Our findings highlight that not all unused dependencies should be removed in practice.
Nabhan Suwanachote, Yagut Shakizada, Yutaro Kashiwa, Bin Lin 0008, Hajimu Iida
MSR5
2025 On the Use of Agentic Coding Manifests: An Empirical Study of Claude Code
Worawalan Chatlatanagulchai, Kundjanasith Thonglek, Brittany Reid, Yutaro Kashiwa, Pattara Leelaprute, Arnon Rungsawang, Bundit Manaskasemsak, Hajimu Iida
PROFES8
2025 An Empirical Study of Security-Policy Related Issues in Open Source Projects
Rintaro Kanaji, Brittany Reid, Yutaro Kashiwa, Raula Gaikovina Kula, Hajimu Iida
PROFES5
2025 Detecting and Characterizing Low and No Functionality Packages in the NPM Ecosystem
Napasorn Tevarut, Brittany Reid, Yutaro Kashiwa, Pattara Leelaprute, Arnon Rungsawang, Bundit Manaskasemsak, Hajimu Iida
PROFES7
2025 Load-Aware Multi-Objective Optimization of Controller and Datastore Placement in Distributed Sdns
abstract
ABSTRACT In distributed Software Defined Networking (SDN), multiple controllers need to maintain a consistent view of the network state among themselves using consensus algorithms, introducing additional communication overhead and network delay, especially in large‐scale networks. Therefore, optimizing controller placement presents significant challenges, as it must account not only for the delay between switches and controllers but also for the delay introduced by consensus algorithms. Additionally, SDN controllers have limited capacity in terms of the number of switches they can manage and the network events they can process. Improper placement of controllers can lead to longer message processing times, increased queuing delays, or even controller failures. Thus, achieving balanced workloads among controllers is essential. This study introduces and validates a practical Flow Setup Time (FST) model to measure controller response times. We proposed an advanced multi‐objective optimization approach that incorporates the Variance of Load Balancing (VOLB), to determine the optimal placements of controllers and datastore nodes involved in processing consensus algorithms. Furthermore, we applied this optimization method to different types of real networks from the Internet Topology Zoo dataset. Based on experimental findings, we identified key factors to consider when selecting optimal placement strategies, including the trade‐offs between the number of controllers, the number of datastore nodes, FST, and VOLB.
Xingyuan Kang, Keichi Takahashi, Chawanat Nakasan, Kohei Ichikawa, Hajimu Iida
Concurr. Comput. Pract. Exp.5
2024 An Evaluation of Time-Sliced GPU Sharing with KubeRay for Machine Learning Workloads
abstract
The increasing complexity of machine learning (ML) and artificial intelligence (AI) applications necessitates efficient GPU resource management in distributed environments such as Kubernetes. Conventional one-to-one GPU mapping, which allocates a single G PU to a single container, often results in the underutilization of these critical resources. Our study introduces an approach that leverages KubeRay and time slicing to enable dynamic GPU sharing among multiple concurrent workloads, significantly improving memory utilization and overall response times. Our findings reveal that while memory efficiency is notably enhanced, the proposed method incurs longer task completion times due to the overhead associated with managing distributed tasks. Specifically, we observed an average increase in task completion times of approximately 74.43 % with two parallel workloads. For three parallel workloads, the average increase in completion times was approximately 158.4 %. This study reveals the trade-offs between improved resource utilization and execution time, highlighting the need for future research to optimize these mechanisms in Kubernetes-based ML operations.
Papon Choonhaklai, Kohei Ichikawa, Hajimu Iida
COMPSAC3
2024 On the Use of ChatGPT for Code Review: Do Developers Like Reviews By ChatGPT?
abstract
Code review is a critical but time-consuming process for ensuring code quality in modern software engineering. To alleviate the effort of reviewing source code, recent studies have investigated the possibility of automating the review process. Moreover, tools based on large language models such as ChatGPT are playing an increasingly important role in this vision. Understanding how these tools are used during code review can provide valuable insights for code review automation.
Miku Watanabe, Yutaro Kashiwa, Bin Lin 0008, Toshiki Hirao, Ken-ichi Yamaguchi, Hajimu Iida
EASE6
2024 RevToken: A Token-Level Review Recommendation: How Far Are We?
abstract
Code review plays an important role in quality assurance, which improves readability, maintainability, etc. On the other hand, code review is notorious for being very time-consuming work because reviewers need to carefully inspect numerous lines for each change. To alleviate the efforts, many studies have proposed approaches to highlighting the code that reviewers need to review. However, even using the finest-grained approaches (i.e., line-level), the granularity of the recommendations is coarse-grained so it is difficult for developers to figure out what needs to be reviewed. For example, lines in the QtBase projects have a median of 7 tokens and sometimes have more than 12 tokens (10% of the lines). This study proposes a token level approach to recommend where developers review, which is finer-grained than the previous studies. Specifically, we fine-tune CodeBERT so that it can return the tokens that are most likely to be commented on and revised. Our empirical evaluation using the OpenStack and QtBase datasets demonstrated that the proposed approach outperforms the state-of-the-art model to predict lines to be revised. Also, we find that 86% of the predicted tokens are accurately distinguished as either needing modification or not when the lines to be revised are correctly identified.
Yasuhito Morikawa, Yutaro Kashiwa, Kenji Fujiwara, Hajimu Iida
ICSME4
2024 An Empirical Investigation into the use of Dockerfile Preprocessors for Docker Image Management
abstract
Docker plays a crucial role in providing uniform software development. Many Docker development projects deliver multiple images in order to support various users who need different base images, versions, and architectures. To do so, the projects need to develop different contents of Dockerfiles for each support. For example, if developers provide their product on different Linux OSs, Dockerfiles need to contain package installing commands with an appropriate package manager for each Linux OS. To reduce the development tasks, many projects often develop their own tool to generate multiple Docker images automatically (hereafter, Dockerfile Preprocessors). However, it is still not clear how the projects adopt Dockerfile Preprocessors and what the benefits are. This study explores the characteristics of projects using Dockerfile Preprocessors, the timing, impact, and purpose, and the maintenance effort of using Dockerfile Preprocessors. Our empirical results show that (i) there is “Container build” pattern that does not generate multiple Dockerfiles; (ii) Projects using DPPs have more tags, supported Docker images, and architecture supports than projects without DPPs; (iii) 66% of projects develop DPPs in the middle of development; (iv) the common reasons for adopting DPP is to reduce the effort of creating Dock-erfiles, and to ease updating versions/variations/architectures; (v) the adoption of DPPs does not increase releasing activities.
Wataru Mabuchi, Yutaro Kashiwa, Kenji Fujiwara, Hajimu Iida
SCAM4
2023 A Pilot Study of Testing Infrastructure as Code for Cloud Systems
abstract
Infrastructure as Code (IaC) has become the de-facto standard method for managing cloud resources. Just like general source code (e.g., Java, etc.), infrastructure code also has numerous bugs so it needs to be tested. While several testing frameworks for IaC for cloud systems have been developed in practice, researchers have paid little attention to their testing. This study presents an empirical investigation of the use of tests for IaC for cloud systems. Our empirical results show that (i) 55.2% of the repositories using Terratest have at least one server infrastructure test; (ii) developers often maintain server infrastructure tests (1.7%-11.3% commits out of all the commits); (iii) many repositories have tests for system functionality (28%), deployment (20%), and configuration (17%).
Nabhan Suwanachote, Soratouch Pornmaneerattanatri, Yutaro Kashiwa, Kohei Ichikawa, Pattara Leelaprute, Arnon Rungsawang, Bundit Manaskasemsak, Hajimu Iida
APSEC8
2023 An Empirical Study on the Use of Snapshot Testing
abstract
Testing is one of the most critical processes in software quality assurance. Developers spend a large portion of their time writing test code to avoid potential software failures. In recent years, snapshot testing, which compares snapshots of UI components to detect unexpected changes, has gained popularity in front-end development due to the need to reduce testing efforts. However, it is still unclear how software developers adopt snapshot testing and maintain them. To facilitate future work to reveal the potentials of snapshot testing, this paper presents a preliminary study which examines how developers use snapshot tests. More specifically, this study investigates 1) the characteristics of projects adopting snapshot testing, and 2) when snapshot tests were introduced and how they evolve. Our study is among the first to understand snapshot testing, providing valuable insights on its adoption. We also highlight the future directions to work on.
Shun Fujita, Yutaro Kashiwa, Bin Lin 0008, Hajimu Iida
ICSME4
2023 An Empirical Investigation on the Performance of Domain Adaptation for T5 Code Completion
abstract
Code completion has the benefit of improving coding speed and reducing the chance of inducing bugs. In recent years, DL-based code completion techniques have been proposed. In particular, pre-trained models have shown outstanding performance because they can complete code by considering the context before and after it is completed. While the model can generate the set of candidate codes, some of those might need to be modified by developers because projects can have different coding rules.In this study, to complete code that fits a specific project appropriately, we train the CodeT5 model with additional data from the target project. This fine-tuning approach is called do-main adaptation, and is often used in neural machine translation. Our preliminary experiment observes that our domain-adapted model improves 5.3% of the perfect prediction rate and, 3.4% of the edit distance rate, compared to the fine-tuned model with the out-of-domain dataset. Furthermore, we discover that the improvement is greater with a larger repository size. The model that is trained with a small dataset, however, hardly improves or performs worse.
Daisuke Fukumoto, Yutaro Kashiwa, Toshiki Hirao, Kenji Fujiwara, Hajimu Iida
SANER5
2022 Acar: An application-aware network routing system using SRv6
abstract
The optimal path varies depending on the communication characteristics of each application. However, existing routing protocols such as BGP and OSPF do not take this fact into account. Although Software Defined Networking (SDN) has been considered as a possible solution to this problem, SDN technologies relying on a centralized controller has scalability issues. SRv6, which is a source routing protocol that enables SDN, handles routing decisions in a decentralized manner and is expected to scale better than previous SDN technologies.This paper proposes Acar, an adaptive routing system using SRv6 that adaptively controls routing by considering the bandwidth requirements of applications and the link utilization of the network. We conducted experiments on a virtual network and demonstrated that Acar achieves better load balancing between links and higher throughput compared to ECMP.
Tomoki Sugiura, Keichi Takahashi, Kohei Ichikawa, Hajimu Iida
CCNC4
2022 Sparse Communication for Federated Learning
abstract
Federated learning trains a model on a centralized server using datasets distributed over a massive amount of edge devices. Since federated learning does not send local data from edge devices to the server, it preserves data privacy. It transfers the local models from edge devices instead of the local data. However, communication costs are frequently a problem in federated learning. This paper proposes a novel method to reduce the required communication cost for federated learning by transferring only top updated parameters in neural network models. The proposed method allows adjusting the criteria of updated parameters to trade-off the reduction of communication costs and the loss of model accuracy. We evaluated the proposed method using diverse models and datasets and found that it can achieve comparable performance to transfer original models for federated learning. As a result, the proposed method has achieved a reduction of the required communication costs around 90% when compared to the conventional method for VGG16. Furthermore, we found out that the proposed method is able to reduce the communication cost of a large model more than of a small model due to the different threshold of updated parameters in each model architecture.
Kundjanasith Thonglek, Keichi Takahashi, Kohei Ichikawa, Chawanat Nakasan, Pattara Leelaprute, Hajimu Iida
ICFEC6
2021 Comparative Performance Study of Lightweight Hypervisors Used in Container Environment
Keichi Takahashi, Kohei Ichikawa, Hajimu Iida, Pree Thiengburanathum, Passakorn Phannachitta
CLOSER4
2020 Understanding Build Errors in Agile Software Development Project-Based Learning
abstract
Recently, various institutions have been conducting advanced programming education aimed at experiencing agile software development in the form of project-based learning (PBL). In the agile software development model, an essential part is the build process. In this study, we investigated students' build behaviors in agile software development PBL (SDPBL) by monitoring and collecting logs of the build process from 2013 to 2016. In our investigation, we collected two types of logs, the local build logs collected by each student's build in their own local programming environment, and remote build logs collected by any team member's commit in team repository. Based on our analysis of the build logs from 2013 to 2015, we found that the causes of remote build errors are related to both technical factors and communications among students in a team. In 2016, the instructors tried to educate to students the reason why the remote build error occur and how it can be resolved. As a result, in 2016, the number of remote build errors and the time required to solve the build errors decreased compared to previous years. It indicates a possibility that the student's comprehension of build error is effective on improving quality of software product and team development.
Erina Makihara, Hiroshi Igaki, Norihiro Yoshida, Kenji Fujiwara, Hajimu Iida
APSEC5
2020 Federated Learning of Neural Network Models with Heterogeneous Structures
abstract
Federated learning trains a model on a centralized server using datasets distributed over a large number of edge devices. Applying federated learning ensures data privacy because it does not transfer local data from edge devices to the server. Existing federated learning algorithms assume that all deployed models share the same structure. However, it is often infeasible to distribute the same model to every edge device because of hardware limitations such as computing performance and storage space. This paper proposes a novel federated learning algorithm to aggregate information from multiple heterogeneous models. The proposed method uses weighted average ensemble to combine the outputs from each model. The weight for the ensemble is optimized using black box optimization methods. We evaluated the proposed method using diverse models and datasets and found that it can achieve comparable performance to conventional training using centralized datasets. Furthermore, we compared six different optimization methods to tune the weights for the weighted average ensemble and found that tree parzen estimator achieves the highest accuracy among the alternatives.
Kundjanasith Thonglek, Keichi Takahashi, Kohei Ichikawa, Hajimu Iida, Chawanat Nakasan
ICMLA4
2020 Retraining Quantized Neural Network Models with Unlabeled Data
abstract
Running neural network models on edge devices is attracting much attention by neural network researchers since edge computing technology is becoming more powerful than ever. However, deploying large neural network models on edge devices is challenging due to the limitation in available computing resources and storage space. Therefore, model compression techniques have been recently studied to reduce the model size and fit models on resource-limited edge devices. Compressing neural network models reduces the size of a model, but also degrades the accuracy of the model since it reduces the precision of weights in the model. Consequently, a retraining method is required to recover the accuracy of compressed models. Most existing retraining methods require the original labeled training datasets to retrain the models, but labeling is a time-consuming process. In particular, we cannot always access the original labeled datasets because of privacy policies and license limitations. In this paper, we propose a method to retrain a compressed neural network model with an unlabeled dataset that is different from the original labeled dataset. We compress the neural network model using quantization to decrease the size of the model. Subsequently, the compressed model is retrained by our proposed retraining method without using a labeled dataset to recover the accuracy of the model. We compared the proposed retraining method against the conventional retraining. The proposed method reduced the size of VGG-16 and ResNet-50 by 81.10% and 52.45%, respectively without significant accuracy loss. In addition, our proposed retraining method is clearly faster than the conventional retraining method.
Kundjanasith Thonglek, Keichi Takahashi, Kohei Ichikawa, Chawanat Nakasan, Hidemoto Nakada, Ryousei Takano, Hajimu Iida
IJCNN7
2019 Improving Resource Utilization in Data Centers using an LSTM-based Prediction Model
abstract
Data centers are centralized facilities where computing and networking hardware are aggregated to handle large amounts of data and computation. In a data center, computing resources such as CPU and memory are usually managed by a resource manager. The resource manager accepts resource requests from users and allocates resources to their applications. A commonly known problem in resource management is that users often request more resources than their applications actually use. This leads to the degradation of overall resource utilization in a data center. This paper aims to improve resource utilization in data centers by predicting the required resource for each application. We designed and implemented a neural network model based on Long Short-Term Memory (LSTM) to predict more efficient resource allocation for a job based on historical data. Our model has two LSTM layers each of which learns the relationship between: (1) allocation and usage, and (2) CPU and memory. We used Googles cluster-usage trace, which contains a trace of resource allocation and usage for each job executed on a Google data center, to train our neural network. Googles cluster scheduler simulator was used to evaluate our proposed method. Our simulation indicated that the proposed method improved the CPU utilization and memory utilization by 10.71% and 47.36%, respectively, compared to a conventional resource manager. Moreover, we discovered that increasing the memory cell size of our LSTM model improves the accuracy of the prediction in return for longer training time.
Kundjanasith Thonglek, Kohei Ichikawa, Keichi Takahashi, Hajimu Iida, Chawanat Nakasan
CLUSTER4
2019 Identifying and predicting key features to support bug reporting
abstract
Abstract Bug reports are the primary means through which developers triage and fix bugs. To achieve this effectively, bug reports need to clearly describe those features that are important for the developers. However, previous studies have found that reporters do not always provide such features. Therefore, we first perform an exploratory study to identify the key features that reporters frequently miss in their initial bug report submissions. Then, we propose an approach that predicts whether reporters should provide certain key features to ensure a good bug report. A case study of the bug reports for Camel, Derby, Wicket, Firefox, and Thunderbird projects shows that Steps to Reproduce, Test Case, Code Example, Stack Trace, and Expected Behavior are the additional features that reporters most often omit from their initial bug report submissions. We also find that these features significantly affect the bug‐fixing process. On the basis of our findings, we build and evaluate classification models using four different text‐classification techniques to predict key features by leveraging historical bug‐fixing knowledge. The evaluation results show that our models can effectively predict the key features. Our comparative study of different text‐classification techniques shows that naïve Bayes multinomial (NBM) outperforms other techniques. Our findings can benefit reporters to improve the contents of bug reports.
Md. Rejaul Karim, Akinori Ihara, Eunjong Choi, Hajimu Iida
J. Softw. Evol. Process.4
2018 An Investigation of the Relationship between Extract Method and Change Metrics: A Case Study of JEdit
abstract
Extract Method is one of the most widely used refactoring patterns. So far, low quality of source code has been regarded as an indicator for Extract Method opportunities. However, recent studies showed that there is no clear relationship between source code quality and Extract Method. Change metrics can be indicators for Extract Method because the characteristics of software evolution strongly affect software quality. However, there has been no study that investigated the relationship between change metrics and Extract Method. In this study, we conducted two studies investigating the relationship between Extract Method and change metrics. As a result, we found that (1) change metrics have a clear relationship with Extract Method and (2) both product and change metrics are necessary to recommend candidates for Extract Method with high accuracy.
Eunjong Choi, Daiki Tanaka, Norihiro Yoshida, Kenji Fujiwara, Daniel Port, Hajimu Iida
APSEC6
2018 "Was my contribution fairly reviewed?": a framework to study the perception of fairness in modern code reviews
abstract
Modern code reviews improve the quality of software products. Although modern code reviews rely heavily on human interactions, little is known regarding whether they are performed fairly. Fairness plays a role in any process where decisions that affect others are made. When a system is perceived to be unfair, it affects negatively the productivity and motivation of its participants. In this paper, using fairness theory we create a framework that describes how fairness affects modern code reviews. To demonstrate its applicability, and the importance of fairness in code reviews, we conducted an empirical study that asked developers of a large industrial open source ecosystem (OpenStack) what their perceptions are regarding fairness in their code reviewing process. Our study shows that, in general, the code review process in OpenStack is perceived as fair; however, a significant portion of respondents perceive it as unfair. We also show that the variability in the way they prioritize code reviews signals a lack of consistency and the existence of bias (potentially increasing the perception of unfairness). The contributions of this paper are: (1) we propose a framework---based on fairness theory---for studying and managing social behaviour in modern code reviews, (2) we provide support for the framework through the results of a case study on a large industrial-backed open source project, (3) we present evidence that fairness is an issue in the code review process of a large open source ecosystem, and, (4) we present a set of guidelines for practitioners to address unfairness in modern code reviews.
Daniel M. Germán, Gregorio Robles, Germán Poo-Caamaño, Xin Yang 0018, Hajimu Iida, Katsuro Inoue
ICSE5
2018 Review participation in modern code review: An empirical study of the Android, Qt, and OpenStack projects (journal-first abstract)
abstract
This paper empirically investigates the factors influence review participation in the MCR process. Through a case study of the Android, Qt, and OpenStack open source projects, we find that the amount of review participation in the past is a significant indicator of patches that will suffer from poor review participation. Moreover, the description length of a patch and the purpose of introducing new features also share a relationship with the likelihood of receiving poor review participation. This paper is an extended abstract of a paper published in the Empirical Software Engineering journal. The original paper is communicated by Jeffrey C. Carver.
Patanamon Thongtanunam, Shane McIntosh, Ahmed E. Hassan, Hajimu Iida
SANER4
2017 Container Rebalancing: Towards Proactive Linux Containers Placement Optimization in a Data Center
abstract
Similar to Virtualization, Linux Containers (LXC) provides high-performance, lightweight computing resource allocation and isolation. Each LXC container has a resource overhead smaller than that of a virtual machine, leading to significantly lower container migration time and making frequent container placement modification a viable optimization technique. Traditional container scheduling mechanisms do not leverage this property of LXC. Generally, a scheduler tries to find the most optimal placement for a new container, the allocated host then executes the scheduled container until the end of the container's life cycle. This strategy works fine for short-lived containers. With a long-lived container such as a server process becoming more and more common, and the container placement calculated at the beginning of the execution may not remain optimal during the container's lifetime, since the other containers are moving in and out of the cluster. This research proposes container rebalancing, a novel scheduling mechanism with a rebalancing process working alongside a scheduling process. The container rebalancing method increases LXC cluster utilization while maintaining minimal interference with the scheduling process. This is done by continuously modifying container placement, by using the rebalancing process, in order to load-balance utilization of each host in the LXC cluster. LXC cluster simulation driven by Google's cluster data is used to verify the feasibility of container rebalancing. Simulation results show an observable increase in container scheduled rate and cluster utilization with no drawback, suggesting that container rebalancing is a promising method.
Pongsakorn U.-Chupala, Yasuhiro Watashiba, Kohei Ichikawa, Susumu Date, Hajimu Iida
COMPSAC (1)5
2017 A simple multipath OpenFlow controller using topology-based algorithm for multipath TCP
abstract
Summary Multipath transmission control protocol(TCP), or MPTCP, is a widely‐researched mechanism that allows a single application‐level connection to be split to more than 1 TCP stream and, consequently, more than 1 network interface, as opposed to the traditional TCP/IP model. Being a transport layer protocol, MPTCP can easily interact between the application using it and the network supporting it. However, MPTCP does not have control of its own route. Default IP routing behavior generally takes all traffic through the shortest or best metric path. However, this behavior may actually cause paths to collide with each other, creating contention for bandwidth in a number of edges. This can result in a bottleneck that limits the throughput of the network. Therefore, a multipath routing mechanism is necessary to ensure smooth operation of MPTCP. We created smoc, a simple multipath OpenFlow controller, that uses only topology information of the network to avoid collision where possible. Evaluation of smoc in a virtual local‐area and a physical wide–area software‐defined networks showed favorable results as smoc provided better performance than simple or spanning tree–routing mechanisms.
Chawanat Nakasan, Kohei Ichikawa, Hajimu Iida, Putchong Uthayopas
Concurr. Comput. Pract. Exp.3
2017 Review participation in modern code review - An empirical study of the android, Qt, and OpenStack projects
Patanamon Thongtanunam, Shane McIntosh, Ahmed E. Hassan, Hajimu Iida
Empir. Softw. Eng.4
2016 A hosting service of multi-language historage repositories
abstract
In the research of Mining Software Repositories, source code repositories are one of the core sources since it contains the product and the process of software development. A source code repository stores the versions of files and makes it possible to browse the histories of files, such as modification dates, authors, messages, so on. Although such rich information of file histories is easily available, extracting the histories of methods/functions, which are elements of source code files, is not easy from general code repositories. To tackle this difficulty, we have developed Historage, a fine-grained version control system. Historage repository is a Git repository, which is built upon an original Git repository. Therefore, similar mining techniques for general Git repositories are applicable to Historage repositories. We also have developed Kataribe, a hosting service of Historage repositories, which contains hundreds of Historage repositories constructed from repositories in GitHub, which are written in C#, Java, Python and Ruby. The list of all Historage and original repositories are available at http://kataribe.naist.jp/public. With this dataset, we will promote in-depth and fine-grained software evolution research with diversity of programming languages.
Kyohei Uemura, Yusuke Saito, Shin Fujiwara, Daiki Tanaka, Kenji Fujiwara, Hajimu Iida, Ken-ichi Matsumoto
ICIS6
2016 An SDN-Based Multipath GridFTP for High-Speed Data Transfer
abstract
We demonstrate high-speed data transfer GridFTP using a multipath control mechanism with SDN (Software-Defined Networking). GridFTP is a typical tool that has been developed and widely used for bulk data transfer over a wide area network in the field. GridFTP supports a parallel high-speed data transfer scheme using multiple TCP streams. However, one of the shortest paths is used solely for data transfer in the default IP routing while there are multiple network paths (multipath) exist between widely-distributed sites. In this study, we propose a system that distributes the parallel TCP streams of GridFTP into multiple network paths by a traffic engineering technique brought by SDN. Our proposed system has achieved approximately 20% better performance than the conventional method in the best case in a global-scale real enviroment.
Che Huang, Chawanat Nakasan, Kohei Ichikawa, Hajimu Iida
ICDCS4
2016 Revisiting code ownership and its relationship with software quality in the scope of modern code review
abstract
Code ownership establishes a chain of responsibility for modules in large software systems. Although prior work uncovers a link between code ownership heuristics and software quality, these heuristics rely solely on the authorship of code changes. In addition to authoring code changes, developers also make important contributions to a module by reviewing code changes. Indeed, recent work shows that reviewers are highly active in modern code review processes, often suggesting alternative solutions or providing updates to the code changes. In this paper, we complement traditional code ownership heuristics using code review activity. Through a case study of six releases of the large Qt and OpenStack systems, we find that: (1) 67%--86% of developers did not author any code changes for a module, but still actively contributed by reviewing 21%--39% of the code changes, (2) code ownership heuristics that are aware of reviewing activity share a relationship with software quality, and (3) the proportion of reviewers without expertise shares a strong, increasing relationship with the likelihood of having post-release defects. Our results suggest that reviewing activity captures an important aspect of code ownership, and should be included in approximations of it in future studies.
Patanamon Thongtanunam, Shane McIntosh, Ahmed E. Hassan, Hajimu Iida
ICSE4
2016 A Hybrid Game Contents Streaming Method: Improving Graphic Quality Delivered on Cloud Gaming
Kar-Long Chan, Kohei Ichikawa, Yasuhiro Watashiba, Putchong Uthayopas, Hajimu Iida
ICEC5
2016 Detecting exploratory programming behaviors for introductory programming exercises
abstract
Developers often perform the repeating cycle of implementation and evaluation when they need to deal with the unfamiliar portion of the source code. This cycle is named as exploratory programming. We regard exploratory programming as an effective way not only to improve novice's programming skill but also to support educators in programming exercise in University. Because when novices often use the exploratory programming, it means novices struggle to solve their assignments. Therefore, educators should grasp which elements, APIs or blocks novices often used exploratory programming for. In this paper, firstly we propose the definition of novice's exploratory programming to collect logs of exploratory based on various granularity by novices. Secondly, we propose an algorithm based on our proposed definition to automatically detect exploratory programming behaviors. We also conducted a small case study. As a result of automatic detection, our proposed algorithm allows us to know what elements of program novices often feel difficult and struggle for.
Erina Makihara, Hiroshi Igaki, Norihiro Yoshida, Kenji Fujiwara, Hajimu Iida
ICPC5
2016 Mining the modern code review repositories: a dataset of people, process and product
abstract
In this paper, we present a collection of Modern Code Review data for five open source projects. The data showcases mined data from both an integrated peer review system and source code repositories. We present an easy-to-use and richer data structure to retrieve the (a) People, (b) Process, and (c) Product aspects of the peer review. This paper presents the extraction methodology, the dataset structure, and a collection of database dumps.
Xin Yang 0018, Raula Gaikovina Kula, Norihiro Yoshida, Hajimu Iida
MSR4
2015 A Multipath Controller for Accelerating GridFTP Transfer over SDN
abstract
A large amount of scientific data needs to be transferred from one site to another as fast as possible in the computational science fields. High-speed data transfer between sites is very important, especially in the Grid computing field, GridFTP has been widely used for bulk data transfer over a wide area network. GridFTP achieves greater performance by supporting parallel TCP streams. Using parallel TCP streams improves the throughput of slow-start algorithms and lossy networks even on a single path. This research proposes a traffic engineering technique that increases the data transfer performance by using multiple paths simultaneously for the parallel TCP streams. For this purpose, we use Software-Defined Network (SDN) technology and its implementation, OpenFlow. This paper presents the design and implementation of the proposed system. Our performance evaluation demonstrates that our proposed system can accelerate GridFTP Transfer in both virtual and real global-scale environments.
Che Huang, Chawanat Nakasan, Kohei Ichikawa, Hajimu Iida
e-Science4
2015 Quick Trigger on Stack Overflow: A Study of Gamification-Influenced Member Tendencies
abstract
In recent times, gamification has become a popular technique to aid online communities stimulate active member participation. Gamification promotes a reward-driven approach, usually measured by response-time. Possible concerns of gamification could a trade-off between speedy over quality responses. Conversely, bias toward easier question selection for maximum reward may exist. In this study, we analyze the distribution gamification-influenced tendencies on the Q&A Stack Overflow online community. In addition, we define some gamification-influenced metrics related to response time to a question post. We carried experiments of a four-month period analyzing 101,291 members posts. Over this period, we determined a Rapid Response time of 327 seconds (5.45 minutes). Key findings suggest that around 92% of SO members have fewer rapid responses that non-rapid responses. Accepted answers have no clear relationship with rapid responses. However, we did find that rapid responses significantly contain tags that did not follow their usual tagging tendencies.
Xin Yang 0018, Raula Gaikovina Kula, Eunjong Choi, Katsuro Inoue, Hajimu Iida
MSR6
2015 Investigating Code Review Practices in Defective Files: An Empirical Study of the Qt System
abstract
Software code review is a well-established software quality practice. Recently, Modern Code Review (MCR) has been widely adopted in both open source and proprietary projects. To evaluate the impact that characteristics of MCR practices have on software quality, this paper comparatively studies MCR practices in defective and clean source code files. We investigate defective files along two perspectives: 1) files that will eventually have defects (i.e., Future-defective files) and 2) files that have historically been defective (i.e., Risky files). Through an empirical study of 11,736 reviews of changes to 24,486 files from the Qt open source project, we find that both future-defective files and risky files tend to be reviewed less rigorously than their clean counterparts. We also find that the concerns addressed during the code reviews of both defective and clean files tend to enhance evolvability, i.e., Ease future maintenance (like documentation), rather than focus on functional issues (like incorrect program logic). Our findings suggest that although functionality concerns are rarely addressed during code review, the rigor of the reviewing process that is applied to a source code file throughout a development cycle shares a link with its defect proneness.
Patanamon Thongtanunam, Shane McIntosh, Ahmed E. Hassan, Hajimu Iida
MSR4
2015 Who should review my code? A file location-based code-reviewer recommendation approach for Modern Code Review
abstract
Software code review is an inspection of a code change by an independent third-party developer in order to identify and fix defects before an integration. Effectively performing code review can improve the overall software quality. In recent years, Modern Code Review (MCR), a lightweight and tool-based code inspection, has been widely adopted in both proprietary and open-source software systems. Finding appropriate code-reviewers in MCR is a necessary step of reviewing a code change. However, little research is known the difficulty of finding code-reviewers in a distributed software development and its impact on reviewing time. In this paper, we investigate the impact of reviews with code-reviewer assignment problem has on reviewing time. We find that reviews with code-reviewer assignment problem take 12 days longer to approve a code change. To help developers find appropriate code-reviewers, we propose RevFinder, a file location-based code-reviewer recommendation approach. We leverage a similarity of previously reviewed file path to recommend an appropriate code-reviewer. The intuition is that files that are located in similar file paths would be managed and reviewed by similar experienced code-reviewers. Through an empirical evaluation on a case study of 42,045 reviews of Android Open Source Project (AOSP), OpenStack, Qt and LibreOffice projects, we find that RevFinder accurately recommended 79% of reviews with a top 10 recommendation. RevFinder also correctly recommended the code-reviewers with a median rank of 4. The overall ranking of RevFinder is 3 times better than that of a baseline approach. We believe that RevFinder could be applied to MCR in order to help developers find appropriate code-reviewers and speed up the overall code review process.
Patanamon Thongtanunam, Chakkrit Tantithamthavorn, Raula Gaikovina Kula, Norihiro Yoshida, Hajimu Iida, Ken-ichi Matsumoto
SANER5
2014 Application-Oriented Bandwidth and Latency Aware Routing with Open Flow Network
abstract
Bandwidth and latency are two major factors that contribute the most to network application performance. Between each pair of switches in a network, there may be multiple paths connecting them. Each path has different properties because of multiple factors. Traditional shortest-path routing does not take this knowledge into consideration and may result in sub-optimal performance of applications and underutilization of network. We proposed a concept of "bandwidth and latency aware routing". The idea is that we could improve overall performance of the network by separating application into bandwidth-oriented and latency-oriented application and allocate different route for each type of application accordingly. We also proposed a design of this network system implemented using Open Flow. Routes are calculated from monitored information using Dijkstra algorithm and its variation. To support our design, we show a use case in which our design performs better than traditional routing as well as evaluation results.
Pongsakorn U.-Chupala, Kohei Ichikawa, Hajimu Iida, Nawawit Kessaraphong, Putchong Uthayopas, Susumu Date, Hirotake Abe, Hiroaki Yamanaka, Eiji Kawai
CloudCom3
2014 ReDA: A Web-Based Visualization Tool for Analyzing Modern Code Review Dataset
abstract
ReDA (http://reda.naist.jp/) is a web-based visualization tool for analyzing Modern Code Review (MCR) datasets for large Open Source Software (OSS) projects. MCR is a commonly practiced and lightweight inspection of source code using a support tool such as Gerrit system. Recently, mining code review history of such systems has received attention as a potentially effective method of ensuring software quality. However, due to increasing size and complexity of softwares being developed, these datasets are becoming unmanageable. ReDA aims to assist researchers of mining code review data by enabling better understand of dataset context and identifying abnormalities. Through real-time data interaction, users can quickly gain insight into the data and hone in on interesting areas to investigate. A video highlighting the main features can be found at: http://youtu.be/ fEoTRRas0U.
Patanamon Thongtanunam, Xin Yang 0018, Norihiro Yoshida, Raula Gaikovina Kula, Ana Erika Camargo Cruz, Kenji Fujiwara, Hajimu Iida
ICSME7
2014 Kataribe: a hosting service of historage repositories
abstract
In the research of Mining Software Repositories, code repository is one of the core source since it contains the product of software development. Code repository stores the versions of files, and makes it possible to browse the histories of files, such as modification dates, authors, messages, etc. Although such rich information of file histories is easily available, extracting the histories of methods, which are elements of source code files, is not easy from general code repositories. To tackle this difficulty, we have developed Historage, a fine-grained version control system. Historage repository is a Git repository which is built upon original Git repository. Therefore, similar mining techniques for general Git repositories are applicable to Historage repositories. Kataribe is a hosting service of Historage repositories, which enables researchers and developers to browse method histories on the web and clone Historage repositories to local. The Kataribe project aims to maintain and expand the datasets and features.
Kenji Fujiwara, Hideaki Hata, Erina Makihara, Yusuke Fujihara, Naoki Nakayama, Hajimu Iida, Ken-ichi Matsumoto
MSR6
2013 An Empirical Illustration to Validate a FLOSS Development Model Using S-Shaped Curves
abstract
Open source software (OSS) or Free/Libre OSS (FLOSS) has become an interesting source of research in software engineering. However, it has been criticized that FLOSS development is often considered as a homogeneous phenomenon grounded by assumptions rather than empirical evidence. Proper empirical methods that can shed light into FLOSS development are desirable. In this paper, we propose an empirical method to validate a software development model for FLOSS, the Adapted Staged Model for FLOSS. We mined some selected metrics from Apache Ivy and study their evolution using S-shaped curves. Our results indicate that S-shaped curves can model software evolution well for Ivy. Moreover, we demonstrated that our method can be used to identify successfully different stages of its development, validating part of the Adapted Staged Model for FLOSS.
Ana Erika Camargo Cruz, Hajimu Iida, Norbert Preining
ICSM2
2013 Who does what during a code review? datasets of OSS peer review repositories
abstract
We present four datasets that are focused on the general roles of OSS peer review members. With data mined from both an integrated peer review system and code source repositories, our rich datasets comprise of peer review data that was automatically recorded. Using the Android project as a case study, we describe our extraction methodology, the datasets and their application used for three separate studies. Our datasets are available online at http://sdlab.naist.jp/reviewmining/.
Kazuki Hamasaki, Raula Gaikovina Kula, Norihiro Yoshida, Ana Erika Camargo Cruz, Kenji Fujiwara, Hajimu Iida
MSR6
2013 Assessing Refactoring Instances and the Maintainability Benefits of Them from Version Archives
Kenji Fujiwara, Kyohei Fushida, Norihiro Yoshida, Hajimu Iida
PROFES4
2013 Micro process analysis of maintenance effort: an open source software case study using metrics based on program slicing
abstract
SUMMARY For any software project, most experts regard the maintenance phase as the most effort and cost intensive of all phases in the software development life cycle. This is due to the highmaintenance effort, time, and resources needed to effectively address issues during software maintenance (maintenance activities). Mismanagement of these efforts can lead to the degradation of software maintainability. Understanding the assessment of the related software processes can help sustain or improve maintainability during these maintenance activities. Recent studies have shown that current software process assessments are expensive, generic, and complex, especially for smaller organizations. In this paper, we investigate an alternative software process assessment approach performed by analyzing fine‐grained processes (micro processes) of maintenance activities. This approach assesses maintenance efforts based on micro processes in relation to their impact on source code. The approach derives maintenance effort from the complexity and duration of micro processes and uses proposed metrics based on program slicing to measure change impact. In this paper, we investigate an alternative software process assessment approach by analysing fine‐grained processes (micro processes) of maintenance activities. At statistically significant levels, results suggest that the level of the maintenance efforts correlates with its impact on source code. Copyright © 2012 John Wiley & Sons, Ltd.
Raula Gaikovina Kula, Kyohei Fushida, Norihiro Yoshida, Hajimu Iida
J. Softw. Evol. Process.4
2012 Understanding OSS Peer Review Roles in Peer Review Social Network (PeRSoN)
abstract
Due to the distributed collaborations and the volunteering nature of Open Source Software (OSS), OSS peer review processes differs from traditional approaches. Despite the latest research efforts to understand OSS peer review processes, very little is known. Unlike related work, this study investigates OSS peer review processes from a different perspective. We investigate the importance of OSS peer review contributor roles and their review activities by using social network analysis (SNA), proposed as PeRSoN (Peer Review Social Network). As a case study, we extracted and analyzed the review process of Android Open Source Project (AOSP). To the best of our knowledge, this is the first research constructing social networks from mining a peer review repository. Our preliminary results provided hints on relationships among the OSS peer review contributor roles, their activities, and the network structure. The results raised issues that will be used to refine our approach in the future.
Xin Yang 0018, Raula Gaikovina Kula, Ana Erika Camargo Cruz, Norihiro Yoshida, Kazuki Hamasaki, Kenji Fujiwara, Hajimu Iida
APSEC7
2011 A Process Complexity-Product Quality (PCPQ) Model Based on Process Fragment with Workflow Management Tables
Masaki Obana, Noriko Hanakawa, Hajimu Iida
PROFES3
2010 Analysis of Bug Fixing Processes Using Program Slicing Metrics
Raula Gaikovina Kula, Kyohei Fushida, Shinji Kawaguchi, Hajimu Iida
PROFES4
2010 Standardizing the Software Tag in Japan for Transparency of Development
Masateru Tsunoda, Tomoko Matsumura, Hajimu Iida, Kozo Kubo, Shinji Kusumoto, Katsuro Inoue, Ken-ichi Matsumoto
PROFES3
2009 Code Clone Graph Metrics for Detecting Diffused Code Clones
abstract
Code clones (duplicated source code in a software system) are one of the major factors in decreasing maintainability. Many code clone detection methods have been proposed to find code clones automatically from large-scale software. However, it is still hard to find harmful code clones to improve maintainability because there are many code clones that should remain. Thus, to help find harmful code clones, we propose a code clone visualization method and a metrics application on the visualized information. Our method enables the location of harmful code clones diffused in a software system. We apply our method to three open source software programs and visualize their code clone information.
Yoshihiko Fukushima, Raula Gaikovina Kula, Shinji Kawaguchi, Kyohei Fushida, Masataka Nagura, Hajimu Iida
APSEC6
2007 Email and Trouble Report Analysis for Revealing Context with the Project Replayer
Kimiharu Okura, Shinji Kawaguchi, Noriko Hanakawa, Hajimu Iida
APSEC4
2006 Project Replayer with Email Analysis - Revealing Contexts in Software Development
abstract
In many software development projects, people tend to repeat same mistakes due to lack of shared knowledge from past experiences. Generally, it is very difficult to manually find out valuable phenomena from huge data. Invisible context, which cannot be known directly from software documents or formal reports, is an important factor to these difficulties. We propose a new method to find contexts based on analysis to email archives in a project repository. In this method, we first apply natural language processing to extract keywords from email messages. Next, similarities among the messages are calculated based on the extracted keywords, and the messages are classified into clusters according to the similarities. The clustering result can be presented with other information such as code growth graph or schedule charts. This method is implemented as an extension to the Project Replayer, a tool to review past project data. Pilot analysis confirms that a researcher could grasp important contexts of failures in actual projects using the Project Replayer.
Kimiharu Okura, Keita Goto, Noriko Hanakawa, Shinji Kawaguchi, Hajimu Iida
APSEC5
2006 A Software Process Tailoring System Focusing to Quantitative Management Plans
Kazumasa Hikichi, Kyohei Fushida, Hajimu Iida, Ken-ichi Matsumoto
PROFES3
2005 Mega Software Engineering
Katsuro Inoue, Pankaj K. Garg, Hajimu Iida, Ken-ichi Matsumoto, Koji Torii
PROFES3
2002 Daibutsu-den: A Component-Based Framework for Organizational Process Asset Utilization
Hajimu Iida, Yasushi Tanaka, Ken-ichi Matsumoto
PROFES1
2000 A Practical Method for Watermarking Java Programs
abstract
Java programs distributed through the Internet are now suffering from program theft. This is because Java programs can be easily decomposed into reusable class files and even decompiled into source code by program users. We propose a practical method that discourages program theft by embedding Java programs with a digital watermark. Embedding a program developer's copyright notation as a watermark in Java class files will ensure the legal ownership of class files. Our embedding method is indiscernible by program users, yet enables us to identify an illegal program that contains stolen class files. The result of the experiment to evaluate our method showed most of the watermarks (20 out of 23) embedded in class files survived two kinds of attacks that attempt to erase watermarks: an obfuscactor attack, and a decompile-recompile attack.
Akito Monden, Hajimu Iida, Ken-ichi Matsumoto, Koji Torii, Katsuro Inoue
COMPSAC2
1999 Genereation of Object-Oriented Software Process Using Milestones
abstract
One of the major problems in object-oriented software projects is the lack of management's ability to comprehend and control the development progress of a project. This is because traditional phases of software development are not appropriate for object-oriented development. The project manager's "road map" is likely to be different with different phases, different milestones, and different checkpoints. This paper proposes a new framework which gives us a guideline for generating software process with relevant milestones for object-oriented development methods. The framework provides algorithms for identifying development phases and baseline products based on relationships among activities and products of the development method. In addition, the framework defines a software process model to manage a development progress in which milestones are established at the end of each phase in order to check the baseline product and establish goals for the following phase. Results of the application of the proposed framework show that the framework can generate a software process customized for well-known object-oriented development methods in a systematic way.
Noriko Hanakawa, Hajimu Iida, Ken-ichi Matsumoto, Koji Torii
Int. J. Softw. Eng. Knowl. Eng.2
1997 Conceptual Issues of an Object-Centered Process Model
abstract
We propose an object-centered software process description model. We also present the idea of a software development management environment based on the model. To use this model and environment, we illustrate the software development environment as it is, and provide a framework for software process description, management and improvement.
Makoto Matsushita, Makoto Oshita, Hajimu Iida, Katsuro Inoue
APSEC3
1996 A Framework of Generating Software Process Including Milestones for Object-Oriented Development Method
abstract
One of the major problems in object-oriented software projects is the lack of management ability to comprehend and control the development progress of a project. This is because traditional phases of software development are not appropriate for object-oriented development. The project manager's "road map" is likely to be different with different phases, different milestones, and different checkpoints. The paper proposes a new framework which gives one a rigorous guideline for generating software process with relevant milestones for various kinds of software development methods, especially for object-oriented development methods. The framework provides algorithms for identifying development phases and baseline products based on relationships among activities and products of the development method. In addition, the framework defines a software process model to manage a development progress in which milestones are established at the end of each phase in order to check the baseline product and establish goals for the following phase. Results of the application of the proposed framework show that the framework can generate a software process customized for well-known object-oriented development methods in a systematic way.
Noriko Hanakawa, Hajimu Iida, Ken-ichi Matsumoto, Koji Torii
APSEC2
1996 Simulation Model of Overlapping Development Process Based on Progress of Activities
abstract
In many cases of small scale software development, downstream activities are actually carried out before upstream activities have completely finished in order to reduce the development time and meet the deadline. However, overlapping activities also increase the total work effort because it demands more communications between activities. It is very important to be able to estimate the tradeoff between development time and total work effort since constraints may be placed on them. The paper proposes a simulation model to describe the overlapping development processes and the effects between overlapped activities. Based on this model, a software process simulator was developed. This simulator is applicable for supporting project planning (scheduling and staffing) with constraints on cycle time and total work effort.
Hajimu Iida, Jun Eijima, Satushi Yabe, Ken-ichi Matsumoto, Koji Torii
APSEC1
1996 An Interaction Support Mechanism in Software Development
abstract
The paper proposes a new modeling method of interactions in the software development process, which focuses on the interactions among the elements of the process, and a new software development environment based on the model. In this method, interactions in the software process are modeled as a set of agents and communication channels. An agent interacts with other agents with channels. Channels are classified according to their content and type of interaction. A prototype of the supporting environment for software development which is based on the model is also developed. The environment consists of a proxy program for the agent and integrated communication server, which provides mechanisms for interaction, process execution, and user navigation.
Makoto Matsushita, Katsuro Inoue, Hajimu Iida
APSEC3
1991 Generating software development environments from the description of product relations
abstract
A method is described for constructing software development support system from the description of software product relations. Logical structures of the products have ben expressed by a tree structure. Each product appearing through the software development corresponds to a leaf in a tree and a set of those products and other sets correspond to an internal node. Various kinds of software development are easily expressed by changing the products at the leaves. The description of product relation is translated into a script of process description language (PDL). A support system is obtained by executing the translated PDL script with the PDL interpreter which the authors have developed.>
Hajimu Iida, Yoshihiro Nishimura, Katsuro Inoue, Koji Torii
COMPSAC1
1991 Functional language for enacting software processes
abstract
In order to define software processes formally and to use the defined processes in actual development situations, the authors have designed the process description language (PDL) and an associated PDL System. PDL is a functional programming language based on an algebraic specification language, and development processes may be defined in PDL at various levels of abstraction. Abstract PDL scripts (programs) specify the general course of process execution. Essential features of the language for process description are discussed. The characteristics of the PDL System are also presented.>
Katsuro Inoue, Takeshi Ogihara, Hajimu Iida, Minoru Nitta
COMPSAC3