Katherine R. Dearstyne

dblp:320/4113 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2025
0009-0003-9218-3544ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Evaluating Reinforcement Learning Safety and Trustworthiness in Cyber-Physical Systems
abstract
Cyber-Physical Systems (CPS) often leverage Reinforcement Learning (RL) techniques to adapt dynamically to changing environments and optimize performance. However, it is challenging to construct safety cases for RL components. We therefore propose the SAFE-RL (Safety and Accountability Framework for Evaluating Reinforcement Learning) for supporting the development, validation, and safe deployment of RL-based CPS. We adopt a design science approach to construct the framework and demonstrate its use in three RL applications in small Uncrewed Aerial systems (sUAS).
Katherine R. Dearstyne, Pedro Alarcon Granadeno, Theodore Chambers, Jane Cleland-Huang
CAIN1
2025 Can Llms Update Api Documentation?
abstract
Human-written API documentation often becomes outdated, requiring developers to update it manually. Researchers have proposed identifying outdated API name references in documentation, yet have not addressed updating API documentation. Now, emerging large language models (LLMs) are capable of generating code examples and text descriptions. Then, a key question arises: Can LLMs assist in updating API documentation? In this paper, we propose an approach for leveraging an LLM to update API documentation with code change information. To evaluate this approach, we select five open-source projects that manage documentation revisions on GitHub and analyze the differences in documentation between two releases to derive ground truths. We then assess the accuracy of LLM-generated updates by comparing them to the ground truths. Our results show that LLM-generated updates achieve higher METEOR than outdated API documentation (0.771 vs 0.679). It indicates that the LLM updates are more similar to the human updates than the outdated documentation. Our results also reveal that LLMs update code-related information in API documentation with a maximum F1 score of$\mathbf{0. 9 2 1}$.
Seonah Lee 0001, Jueun Heo, Katherine R. Dearstyne
ICSME3
2025 Intelligent Traceability to Support Software Maintainability and Accountability
abstract
As software systems grow in complexity, maintaining traceability—the ability to establish and manage relationships between software artifacts—becomes increasingly critical for supporting requirements validation, change impact analysis, and compliance assessment. Despite its recognized importance, the manual effort required to create and maintain trace links presents a significant barrier to adoption and limits traceability’s practical value. This research investigates how Large Language Models (LLMs) can address fundamental traceability challenges across four key areas: improving trace-link prediction in data-scarce environments, generating software artifacts and establishing traceability in projects lacking formal documentation, tailoring automated traceability to support accountability in regulated domains such as Cyber-Physical Systems, and enhancing software maintenance through traceability-supported differential testing workflows. By addressing these challenges, this research aims to make comprehensive traceability accessible across all software projects, enabling existing engineering practices to more effectively leverage trace relationships for enhanced maintainability, safety, and regulatory compliance.
Katherine R. Dearstyne
RE1
2025 QUESTRL: A Q&A Framework for Designing Trustworthy Reinforcement Learning Systems
abstract
Cyber-Physical Systems (CPS) increasingly leverage Reinforcement Learning (RL) to adapt dynamically to changing environments and optimize performance over time. While RL enhances efficiency and safety by enabling autonomous adjustments to unexpected conditions and hazard avoidance, it also introduces significant risks, as learned behaviors may lead to unpredictable or unsafe actions in real-world deployment. Therefore, integrating risk management into RL system design is essential. In this paper, we propose the QuestRL Framework, a question-driven approach that translates high-level safety guidelines into RL-specific considerations. This framework helps RL practitioners address key risks early in development, informing new or existing system requirements while ensuring traceability to risk management objectives. To evaluate its effectiveness, we conducted a study across two use cases, engaging six RL experts in developing system requirements with and without the framework. Our findings suggest that the framework promotes critical thinking and helps practitioners identify additional risk factors, ultimately supporting safer RL deployment.
Katherine R. Dearstyne, Pedro Alarcon Granadeno, Theodore Chambers, Jane Cleland-Huang
RE1
2024 Supporting Software Maintenance with Dynamically Generated Document Hierarchies
abstract
Software documentation supports a broad set of software maintenance tasks; however, creating and maintaining high-quality, multi-level software documentation can be incredibly time-consuming and therefore many code bases suffer from a lack of adequate documentation. We address this problem through presenting HGEN, a fully automated pipeline that leverages LLMs to transform source code through a series of six stages into a well-organized hierarchy of formatted documents. We evaluate HGEN both quantitatively and qualitatively. First, we use it to generate documentation for three diverse projects, and engage key developers in comparing the quality of the generated documentation against their own previously produced manually-crafted documentation. We then pilot HGEN in nine different industrial projects using diverse datasets provided by each project. We collect feedback from project stakeholders, and analyze it using an inductive approach to identify recurring themes. Results show that HGEN produces artifact hierarchies similar in quality to manually constructed documentation, with much higher coverage of the core concepts than the baseline approach. Stakeholder feedback highlights HGEN's commercial impact potential as a tool for accelerating code comprehension and maintenance tasks. Results and associated supplemental materials can be found at https://zenodo.org/records/11403244
Katherine R. Dearstyne, Alberto D. Rodriguez, Jane Cleland-Huang
ICSME1
2024 ROOT: Requirements Organization and Optimization Tool
abstract
Software engineering practices such as constructing requirements and establishing traceability help ensure systems are safe, reliable, and maintainable. However, they can be resource-intensive and are frequently underutilized. To alleviate the burden of these essential processes, we developed the Requirements Organization and Optimization Tool (ROOT). ROOT cen-tralizes project information and offers project visualizations and AI-based tools designed to streamline engineering processes. With ROOT's assistance, engineers benefit from improved oversight and early error detection, leading to the successful development of software systems. Link to screen cast: https://youtu.be/3rtMYRnsu24
Katherine R. Dearstyne, Alberto D. Rodriguez, Jane Cleland-Huang
ICSME1
2022 SAFA: A Tool for Supporting Safety Analysis in Evolving Software Systems
abstract
Many organizations seek to increase their agility in order to deliver more timely and competitive products. However, in safety-critical systems such as medical devices, autonomous vehicles, or factory floor robots, the release of new features has the potential to introduce hazards that potentially lead to run-time failures that impact software safety. As a result, many projects suffer from a phenomenon referred to as the big freeze. SAFA is designed to address this challenge. Through the use of cutting-edge deep-learning solutions, it generates trees of requirements, designs, code, tests, and other artifacts that visually depict how hazards are mitigated in the system, and it automatically warns the user when key artifacts are missing. It also uses a combination of colors, annotations, and recommendations to dynamically visualize change across software versions and augments safety cases with visual annotations to aid users in detecting and analyzing potentially adverse impacts of change upon system safety. A link to our tool demo can be found at https://www.youtube.com/watch?v=r-CwxerbSVA.
Alberto D. Rodriguez, Timothy Newman, Katherine R. Dearstyne, Jane Cleland-Huang
ASE3