Kosei Horikawa

dblp:420/5372 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0006-6317-0754ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Do AI Agents Really Improve Code Readability?
abstract
Code readability is fundamental to software quality and maintainability. Poor readability extends development time, increases bug-inducing risks, and contributes to technical debt. With the rapid advancement of Large Language Models, AI agent-based approaches have emerged as a promising paradigm for automated refactoring, capable of decomposing complex tasks through autonomous planning and execution. While prior studies have examined refactoring by AI agents, these analyses cover all forms of refactoring, including performance optimization and structural improvement. As a result, the extent to which AI agent-based refactoring specifically improves code readability remains unclear.
Kyogo Horikawa, Kosei Horikawa, Yutaro Kashiwa, Hidetake Uwano, Hajimu Iida
MSR2
2026 Testing with AI Agents: An Empirical Study of Test Generation Frequency, Quality, and Coverage
abstract
Agent-based coding tools have transformed software development practices. Unlike prompt-based approaches that require developers to manually integrate generated code, these agent-based tools autonomously interact with repositories to create, modify, and execute code, including test generation. While many developers have adopted agent-based coding tools, little is known about how these tools generate tests in real-world development scenarios or how AI-generated tests compare to human-written ones.
Suzuka Yoshimoto, Shun Fujita, Kosei Horikawa, Daniel Feitosa, Yutaro Kashiwa, Hajimu Iida
MSR3
2025 An Empirical Investigation into Maintenance of Load Testing Scripts
abstract
Background: Modern software systems are expected to deliver high performance under a variety of different workloads. In order to automatically verify whether a system operates correctly under specific load conditions, load testing has become a widely adopted technique. As software systems evolve, their load requirements-such as performance thresholds and usage patterns-also change, necessitating updates to load tests. Aims: This study investigates the maintenance of load testing scripts to better understand how load requirements evolve and how these changes are reflected in the tests themselves. Method: We analyzed 35 open-source software (OSS) repositories that incorporate load testing. We examined the frequency and nature of load test updates. Results: Our analysis reveals that 45.7% of the studied projects do not update their load testing scripts after initial creation. However, a small subset of projects demonstrates extensive and ongoing maintenance of these scripts. Furthermore, we identified 20 distinct update types across 5 major categories of purposes for load testing script modifications. The most frequent update type is related to “Test Maintenance”, followed by “Test Scenario Modification.” Conclusions: Our findings suggest that load testing scripts are often left unmaintained over time in many projects. The updates, when performed, serve a wide range of purposes, with test maintenance being the most frequent.
Ibuki Nakamura, Kosei Horikawa, Brittany Reid, Yutaro Kashiwa, Hajimu Iida
ESEM2
2025 How Does Test Code Differ from Production Code in Terms of Refactoring? An Empirical Study
abstract
Refactoring is a widely applied practice for improving the internal structure of source code without altering its external behavior. Researchers have proposed approaches to detect refactoring operations and investigated their impact on the code quality. However, these studies often focus on production code, paying little attention to test code. It is still unclear whether developers perform refactoring on test code in the same way or for the same purpose. To fill this gap, we first investigate the types and prevalence of refactoring applied in production and test code, and then examine whether these refactorings impact the code quality in a different way. Our results show that certain refactorings are less common in the test code. Besides, while refactoring-related changes in production and/or test code improved readability, they had limited impact on most design smells. We also find that some specific refactoring types do impact certain design smells. These findings indicate the special attention needed for test code when analyzing refactorings.
Kosei Horikawa, Yutaro Kashiwa, Bin Lin 0008, Kenji Fujiwara, Hajimu Iida
ICSME1