VLDB 2026 Research / reviewers in the wild / expert
Kazumasa Shimari
dblp:242/2155
· DBLP profile ↗
17ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0001-8837-5090ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 14 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How AI Coding Agents Communicate: A Study of Pull Request Characteristics and Human Review ResponsesabstractThe rapid adoption of large language models has led to the emergence of AI coding agents that autonomously create pull requests on GitHub. However, how these agents differ in their pull request description characteristics, and how human reviewers respond to them, remains underexplored. In this study, we conduct an empirical analysis of pull requests created by five AI coding agents using the AIDev dataset. We analyze agent differences in pull request description characteristics, including structural features, and examine human reviewer response in terms of review activity, response timing, sentiment, and merge outcomes. We find that AI coding agents exhibit distinct PR description styles, which are associated with differences in reviewer engagement, response time, and merge outcomes. We observe notable variation across agents in both reviewer interaction metrics and merge rates. These findings highlight the role of pull request presentation and reviewer interaction dynamics in human–AI collaborative software development. Kan Watanabe, Rikuto Tsuchida, Takahiro Monno, Kazuma Yamasaki, Youmei Fan, Kazumasa Shimari, Ken-ichi Matsumoto |
MSR | 7 |
| 2026 | Who Writes the Docs in SE 3.0?: Agent vs. Human Documentation Pull RequestsabstractAs software engineering moves toward SE 3.0, AI agents are increasingly used to carry out development tasks and contribute changes to software projects. It is therefore important to understand the extent of these contributions and how human developers review and intervene, since these factors shape the risks of delegating work to AI agents. While recent studies have examined how AI agents support software development tasks (e.g., code generation, issue resolution, and PR automation), their role in documentation tasks remains underexplored–even though documentation is widely consumed and shapes how developers understand and use software. Kazuma Yamasaki, Joseph Ayobami Joshua, Tasha Settewong, Mahmoud Alfadel, Kazumasa Shimari, Ken-ichi Matsumoto |
MSR | 5 |
| 2025 | Round Outcome Prediction in VALORANT Using Tactical Features from Video AnalysisabstractRecently, research on predicting match outcomes in esports has been actively conducted, but much of it is based on match log data and statistical information. This research targets the FPS game VALORANT, which requires complex strategies, and aims to build a round outcome prediction model by analyzing minimap information in match footage. Specifically, based on the video recognition model TimeSformer, we attempt to improve prediction accuracy by incorporating detailed tactical features extracted from minimap information, such as character position information and other in-game events. This paper reports preliminary results showing that a model trained on a dataset augmented with such tactical event labels achieved approximately$\mathbf{8 1 \%}$prediction accuracy, especially from the middle phases of a round onward, significantly outperforming a model trained on a dataset with the minimap information itself. This suggests that leveraging tactical features from match footage is highly effective for predicting round outcomes in VALORANT. Nirai Hayakawa, Kazumasa Shimari, Kazuma Yamasaki, Hirotatsu Hoshikawa, Rikuto Tsuchida, Ken-ichi Matsumoto |
CoG | 2 |
| 2025 | eye2vec: Learning Distributed Representations of Eye Movement for Program Comprehension AnalysisabstractThis paper presents eye2vec, an infrastructure for analyzing software developers' eye movements while reading source code. In common eye-tracking studies in program comprehension, researchers must preselect analysis targets such as control flow or syntactic elements, and then develop analysis methods to extract appropriate metrics from the fixation for source code. Here, researchers can define various levels of AOIs like words, lines, or code blocks, and the difference leads to different results. Moreover, the interpretation of fixation for word/line can vary across the purposes of the analyses. Hence, the eye-tracking analysis is a difficult task that depends on the time-consuming manual work of the researchers. eye2vec represents continuous two fixations as transitions between syntactic elements using distributed representations. The distributed representation facilitates the adoption of diverse data analysis methods with rich semantic interpretations. Haruhiko Yoshioka, Kazumasa Shimari, Hidetake Uwano, Ken-ichi Matsumoto |
ETRA | 2 |
| 2025 | Mining for Lags in Updating Critical Security Threats: A Case Study of Log4j LibraryabstractThe Log4j-Core vulnerability, known as Log4Shell, exposed significant challenges to dependency management in software ecosystems. When a critical vulnerability is disclosed, it is imperative that dependent packages quickly adopt patched versions to mitigate risks. However, delays in applying these updates can leave client systems exposed to exploitation. Previous research has primarily focused on NPM, but there is a need for similar analysis in other ecosystems, such as Maven. Leveraging the 2025 mining challenge dataset of Java dependencies, we identify factors influencing update lags and categorize them based on version classification (major, minor, patch release cycles). Results show that lags exist, but projects with higher release cycle rates tend to address severe security issues more swiftly. In addition, over half of vulnerability fixes are implemented through patch updates, highlighting the critical role of incremental changes in maintaining software security. Our findings confirm that these lags also appear in the Maven ecosystem, even when migrating away from severe threats. Hidetake Tanaka, Kazuma Yamasaki, Momoka Hirose, Takashi Nakano, Youmei Fan, Kazumasa Shimari, Raula Gaikovina Kula, Ken-ichi Matsumoto |
MSR | 6 |
| 2025 | Do Developers Depend on Deprecated Library Versions? A Mining Study of Log4jabstractLog4j has become a widely adopted logging library for Java programs due to its long history and high reliability. Its widespread use is notable not only because of its maturity but also due to the complexity and depth of its features, which have made it an essential tool for many developers. However, Log4j 1.x, which reached its end of support (deprecated), poses significant security risks and has numerous deprecated features that can be exploited by attackers. Despite this, some clients may still rely on this library. We aim to understand whether clients are still using Log4j 1.x despite its official support ending. We utilized the Mining Software Repositories 2025 challenge dataset, which provides a large and representative sample of open-source software projects. We analyzed over 10,000 log entries from the Mining Software Repositories 2025 challenge dataset using the Goblin framework to identify trends in usage rates for both Log4j 1.x and Log4j-core 2.x. Specifically, our study addressed two key issues: (1) We examined the usage rates and trends for these two libraries, highlighting any notable differences or patterns in their adoption. (2) We demonstrate that projects initiated after a deprecated library has reached the end of its support lifecycle can still maintain significant popularity. These findings highlight how deprecated are still popular, with the next step being to understand the reasoning behind these adoptions. Haruhiko Yoshioka, Sila Lertbanjongngam, Masayuki Inaba, Youmei Fan, Takashi Nakano, Kazumasa Shimari, Raula Gaikovina Kula, Ken-ichi Matsumoto |
MSR | 6 |
| 2024 | Nigerian Software Engineer or American Data Scientist? GitHub Profile Recruitment Bias in Large Language ModelsabstractLarge Language Models (LLMs) have taken the world by storm, demonstrating their ability not only to automate tedious tasks, but also to show some degree of proficiency in completing software engineering tasks. A key concern with LLMs is their “black-box” nature, which obscures their internal workings and could lead to societal biases in their outputs. In the software engineering context, in this early results paper, we empirically explore how well LLMs can automate recruitment tasks for a geographically diverse software team. We use OpenAI's ChatGPT to conduct an initial set of experiments using GitHub User Profiles from four regions to recruit a six-person software development team, analyzing a total of 3,657 profiles over a five-year period (2019–2023). Results indicate that ChatGPT shows preference for some regions over others, even when swapping the location strings of two profiles (counterfactuals). Furthermore, ChatGPT was more likely to assign certain developer roles to users from a specific country, revealing an implicit bias. Overall, this study reveals insights into the inner workings of LLMs and has implications for mitigating such societal biases in these models. Takashi Nakano, Kazumasa Shimari, Raula Gaikovina Kula, Christoph Treude, Marc Cheong, Ken-ichi Matsumoto |
ICSME | 2 |
| 2024 | Comparing Execution Trace Using Merkle- Tree to Detect Backward IncompatibilitiesabstractThe use of libraries is crucial in software development. Library users should update their libraries to address bugs and vulnerabilities that are fixed in newer versions. However, updating libraries can lead to software malfunction due to backward incompatibilities. Therefore, it is necessary to carefully examine the changes in the library, identify incompatible behavior, and modify the software accordingly when applying updates. Identifying the cause of incompatibility is challenging as updates often include changes to APIs other than the one used by the user. We propose a method to detect candidate library methods that cause backward incompatibilities in client-side library updates using Merkle tree. Our approach involves conducting unit tests on the client software, which includes library API calls, before and after the library updates. The execution traces of these tests are collected at the Java bytecode instruction level. By constructing Merkle trees for each execution trace before and after the update, we efficiently compare the control structures and return values to identify the differences indicating backward incompatibilities. To validate the effectiveness of our method, we conducted a case study on three instances of incompatibility in open-source software. Atsuhito Yamaoka, Teyon Son, Kazumasa Shimari, Takashi Ishio, Ken-ichi Matsumoto |
SANER | 3 |
| 2024 | Evaluating the effectiveness of size-limited execution trace with near-omniscient debugging
Kazumasa Shimari, Takashi Ishio, Tetsuya Kanda 0001, Katsuro Inoue |
Sci. Comput. Program. | 1 |
| 2023 | Towards Assessment of Practicality of Introductory Programming Course Using Vocabulary of Textbooks, Assignments, and Actual ProjectsabstractIn an assignment-based introductory programming course, a teacher makes a daily class plan based on a textbook and assigns tasks to the students. Assignments are prepared by the teacher so that students can have a better understanding of programming language constructs explained in the course. On the other hand, it is unclear how those language constructs are useful for practical programming tasks. To analyze the practicality of a programming course, this study proposes to compare the vocabularies of code used in textbooks, assignments, and regular programming tasks. If the vocabularies of the textbooks and assignments are closer to that of source code in actual projects, the programming course is considered more practical. As a case study, we have applied the method to evaluate a programming course focusing on data science for graduate students. The result revealed inconsistency between the programming language constructs taught in the course and frequently used in data analysis programs on the Kaggle platform. Kazuki Fukushima, Takashi Ishio, Kazumasa Shimari, Ken-ichi Matsumoto |
CSEE&T | 3 |
| 2023 | Leveraging Execution Trace with ChatGPT: A Case Study on Automated Fault DiagnosisabstractChatGPT possesses the ability to identify potential causes of bugs in a program, which can be used for fault diagnosis. Although ChatGPT cannot always provide accurate responses, it can provide the confidence level that is trained to be correlated with the accuracy of the response, and the confidence level helps verify the accuracy of responses. Through preliminary trials, we found that the explanatory power of potential causes and the confidence level became low even for accurate responses when runtime information is needed to identify the causes of bugs. In this study, we propose a method to construct prompts based on the target program and its execution traces to improve the accuracy of fault diagnosis by ChatGPT. Through case studies using five bugs from Defects4J, we obtained the following two results: (1) For four bugs, the explanatory power of the responses to potential bug causes improved using the information contained in the execution traces. (2) For three bugs, the confidence level was higher for the accurate responses when execution traces were available than when they were not. These results suggest that by using program execution traces in prompts, the accuracy of fault diagnosis by ChatGPT can be improved. Takafumi Sakura, Ryo Soga, Hideyuki Kanuka, Kazumasa Shimari, Takashi Ishio |
ICSME | 4 |
| 2022 | Selecting Test Cases based on Similarity of Runtime Information: A Case Study of an Industrial SimulatorabstractRegression testing is required to check the changes in behavior whenever developers make any changes to a software system. The cost of regression testing is a major problem because developers have to frequently update dependent components to minimize security risks and potential bugs. In this paper, we report a current practice in a company that maintains an industrial simulator as a critical component of their business. The simulator automatically records all the users’ requests and the simulation results in storage. The feature provides a huge number of test cases for regression testing to developers; however, their time budget for testing is limited (i.e., at most one night). Hence, the developers need to select a small number of test cases to confirm both the simulation result and execution performance are unaffected by an update of a dependent component. In other words, the test cases should achieve high coverage while keeping diversity of execution time. To solve the problem, we have developed a clustering-based method to select test cases, using the similarity of execution traces produced by them. The developers have used the method for a half year; they recognize that the method is better than the previous rule-based method used in the company. Kazumasa Shimari, Masahiro Tanaka, Takashi Ishio, Makoto Matsushita, Katsuro Inoue, Satoru Takanezawa |
ICSME | 1 |
| 2022 | didiffff: a viewer for comparing changes in both code and execution tracesabstractOne of the important purposes of code review is to find potential defects caused by other developers' code changes. When reviewing bug fixes, it is important to check the program behavior is properly changed to remove the bug. On the other hand, it is also important to check the program behavior that is not related to the bug is not changed. To investigate the program behavior, omniscient debugging which records all the runtime events is proposed. With omniscient debugging techniques, existing tools visualize multiple execution paths and the states of local variables of a method, but they are not focusing on code changes. In this paper, we implemented a prototype tool that compares and visualizes the difference between two execution traces caused by code changes. Each variable has a maximum of two lists of values, before and after the code changes, so we proposed their categorization based on their difference of length and contents. We also developed a viewer to show both code changes and the difference of execution traces at a glance by extending our previous viewer for omniscient debugging. Tetsuya Kanda 0001, Kazumasa Shimari, Katsuro Inoue |
ICPC | 2 |
| 2022 | JISDLab: A web-based interactive literate debugging environmentabstractThe debugging process is a huge burden on developers, both in terms of time and mentality. Scriptable debugging approaches have been proposed to reduce the burden associated with such debugging work. Scriptable debuggers (SDs) enable to describe developers' debugging process and share the debug scripts to reduce debugging effort. However, SDs require an execution environment for those scripts, and they are unable to manage ancillary information such as execution results and prerequisites for using the script in one place. We extend the existing scriptable debugging and propose an interactive literate debugging environment that enables reproducible bug reporting. The proposed method provides an executable script description that manipulates the debugger, information obtained through the debugger by executing the script, its visualization format, and the ability to save the information in the form of a document that includes explanatory text. By using these documents, it is possible to observe the detailed behavior of a program at runtime and to share the situation in which the focused behavior occurs among developers. In this paper, we describe our proposed interactive literate debugging environment and introduce our prototype tool, JISDLab, which is a web application using Jupyter. The sample debug script used in our demonstration scenario can be accessed via https://github.com/tklabgroup/JISDLab/blob/master/debugspace/case-SANER2022-tooldemo.ipynb Sakutaro Sugiyama, Takashi Kobayashi 0001, Kazumasa Shimari, Takashi Ishio |
SANER | 3 |
| 2021 | NOD4J: Near-omniscient debugging tool for Java using size-limited execution traceabstractLogging is an important feature of a software system to record run-time information. Detailed logging allows developers to collect run-time information in situations where they cannot use an interactive debugger, such as continuous integration and web application server cases. However, extensive logging leads to larger execution traces because few instructions can be repeated many times. This paper presents our tool NOD4J, which monitors a Java program's execution within limited storage space constraints and annotates the source code with observed values in an HTML format. Developers can easily investigate the execution and share the report on a web server. We show two examples that our tool can debug defects using incomplete execution traces. Kazumasa Shimari, Takashi Ishio, Tetsuya Kanda 0001, Naoto Ishida, Katsuro Inoue |
Sci. Comput. Program. | 1 |
| 2019 | Near-Omniscient Debugging for Java Using Size-Limited Execution TraceabstractLogging is an important feature for a software system to record its run-time information. Detailed logging allows developers to collect information in situations where they cannot use an interactive debugger, such as continuous integration and web application server cases. However, extensive logging leads to larger execution traces because few instructions could be repeated many times. To record detailed program behavior within limited storage space constraints, we propose Near-Omniscient Debugging, a methodology that records an execution trace using fixed size buffers for each observed instruction. Our tool monitors a Java program's execution and annotates source code with observed values in an HTML format. Developers can easily investigate the execution and share the report on a web server. In case of DaCapo benchmark applications, our tool requires fewer than 1% of the complete execution traces to visualize all runtime values used by 66% of instructions that are executed less than 64 times. Developers also can obtain data dependencies with precision 91.8% and recall 79.0% using this tool. Kazumasa Shimari, Takashi Ishio, Tetsuya Kanda 0001, Katsuro Inoue |
ICSME | 1 |
| 2019 | PADLA: a dynamic log level adapter using online phase detectionabstractLogging is an important feature for a software system to record its run-time information. Although detailed logs are helpful to identify the cause of a failure in a program execution, constantly recording detailed logs of a long-running system is challenging because of its performance overhead and storage cost. To solve the problem, we propose PADLA (Phase-Aware Dynamic Log Level Adapter) that dynamically adjusts the log level of a running system so that the system can record irregular events such as performance anomalies in detail while recording regular events concisely. PADLA is an extension of Apache Log4j, one of the most popular logging framework for Java. It employs an online phase detection algorithm to recognize irregular events. It monitors run-time performance of a system and learns regular execution phases of a program. If it recognizes a performance anomalies, it automatically changes the log level of a system to record the detailed behavior. In the case study, PADLA successfully recorded a detailed log for performance analysis of a server system under high load while suppressing the amount of log data and performance overhead. Tsuyoshi Mizouchi, Kazumasa Shimari, Takashi Ishio, Katsuro Inoue |
ICPC | 2 |