Xiujing Guo

dblp:301/3192 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-5261-9126ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From noisy feedback to evidence-aware issue specifications: an agent-governed retrieval-augmented generation approach
abstract
Post-release user feedback is a major control signal for maintenance and evolution in modern software development, yet it is noisy, fragmented, and difficult to translate into developer-usable issue specifications. Large Language Models (LLMs) can assist this transformation, but they often hallucinate or over-commit when evidence is weak, conflicting, or incomplete, limiting their robustness in automated software engineering workflows. We propose AGR (Agent-Governed Retrieval-Augmented Generation), a framework that regulates evidence acquisition and generation decisions via agentic control. AGR first applies an agentic triage step to filter low-signal or off-topic feedback, then retrieves evidence from a three-category hierarchy comprising official documentation, historical bug reports, and targeted web sources. It further performs confidence-weighted fusion across authoritative categories and uses an agentic decision module to verify relevance and sufficiency, trigger additional retrieval or online search when needed, reuse prior reports via memory, and abstain when evidence-supported grounding cannot be established. We evaluate AGR on two open-source software ecosystems, Firefox and VS Code. Results show that AGR achieves strong decision accuracy in triage and evidence verification, and produces more actionable and engineering-useful issue specifications than both raw feedback and a strong LLM baseline, while reducing unsupported details.
Zhiyao Wang, Jialong Li 0001, Xiujing Guo, Tatsuhiro Tsuchiya
Autom. Softw. Eng.3
2025 Boundary Value Test Input Generation Using a Large Language Model: Fault Detection and Coverage Analysis
Xiujing Guo, Tatsuhiro Tsuchiya
ADMA (3)1
2025 RAG4Test: Retrieving GUI States for Multilingual Bug Report and Test Case Generation via LLMs
abstract
The scalability and efficiency of software testing are persistently hindered by a reliance on manual practices for creating test cases and bug reports. Moreover, the valuable insights from post-release user feedback are often lost due to the lack of an automated pipeline connecting them to regression testing. To overcome these challenges, we present a novel framework that synergizes the structural representation of software with the generative power of Large Language Models (LLMs). We first represent the application’s Graphical User Interface (GUI) as a directed graph, capturing its components and navigational logic. This queryable graph provides essential, structured context for a Retrieval-Augmented Generation (RAG) model, which then autonomously generates and populates high-quality test cases and bug reports. A key innovation of our work is a fully automated pipeline that processes unstructured user feedback and error reports, transforming them into standardized test cases. We selected some reviews from a popular mobile App for preliminary experiments and verified the feasibility and efficiency improvement of this method.
Zhiyao Wang, Xiujing Guo, Tatsuhiro Tsuchiya
APSEC2
2025 Graph-Centric Approaches for Coverage Optimization in Software Requirement Testing
abstract
In software testing, traceability links between software requirements and test cases are crucial for managing test coverage and detecting defects effectively. Accurate traceability enables comprehensive coverage analysis, identification of untested requirements, and targeted defect detection. The manual effort required to establish and update traceability links often leads to high labor costs and a greater risk of human error. Furthermore, as requirements evolve during the development lifecycle, the effort needed to maintain accurate links increases, driving up maintenance costs and further complicating test management.This study proposes a graph-based approach to establish and maintain traceability throughout the software testing process. By comparing various models for their ability to identify semantic relationships between software requirements and test cases, we selected the most effective method to create accurate traceability links. These links are further utilized through graph queries, enabling efficient analysis of test coverage, identification of untested requirements, and discovery of high-similarity requirement clusters, thereby enhancing the overall testing process. Automatically generating test cases based on query results, our approach seamlessly integrates into the software testing lifecycle, enhancing both coverage and efficiency. In a case study involving real-world industrial data, we effectively identified previously untested requirements and generated a substantial number of high-quality test cases. The results validate the applicability and effectiveness of our approach, demonstrating its potential to improve test traceability and reliability in practical software development environments.
Zhiyao Wang, Xiujing Guo, Tatsuhiro Tsuchiya
COMPSAC2
2025 Retrieval-Augmented Generation for Software Requirement-Based Test Case Generation
abstract
Testers often need to manually write black-box test cases based on software artifacts such as requirement documents. In agile development, this process is often time-consuming and is further complicated by frequent requirement changes, leading to continuous maintenance overhead. Automating this process is therefore essential. Given the strong natural language understanding and generation capabilities of large language models (LLMs), combined with Retrieval-Augmented Generation (RAG), we propose a RAG-based framework for automated test case generation. Before generation, we embed software artifacts to construct a vectorbased knowledge database. At runtime, software requirements are used as queries to retrieve relevant context, which is integrated into a prompt and passed to the LLM for test case generation. This approach addresses several shortcomings of LLMs, including limited context length, attention dilution over large inputs, and the tendency to hallucinate or over-look key domain-specific constraints. By providing query-specific external knowledge, RAG enhances both accuracy and efficiency. We deploy the framework with different models locally and conduct experiments on two open-source datasets. Compared with the manually written benchmark test cases, our method achieves full requirement coverage with fewer test cases, improved efficiency, reduced error potential, and realized better readability.
Zhiyao Wang, Xiujing Guo, Tatsuhiro Tsuchiya
QRS2
2024 Advancing Aspect-Based Sentiment Analysis Through Deep Learning Models
Chen Li 0027, Huidong Tang, Jinli Zhang, Xiujing Guo, Debo Cheng, Yasuhiko Morimoto
ADMA (5)4
2024 Optimal test case generation for boundary value analysis
abstract
Abstract Boundary value analysis (BVA) is a common technique in software testing that uses input values that lie at the boundaries where significant changes in behavior are expected. This approach is widely recognized and used as a natural and effective strategy for testing software. Test coverage is one of the criteria to measure how much the software execution paths are covered by the set of test cases. This paper focuses on evaluating test coverage with respect to BVA by defining a metric called boundary coverage distance (BCD). The BCD metric measures the extent to which a test set covers the boundaries. In addition, based on BCD, we consider the optimal test input generation to minimize BCD under the random testing scheme. We propose three algorithms, each representing a different test input generation strategy, and evaluate their fault detection capabilities through experimental validation. The results indicate that the BCD-based approach has the potential to generate boundary values and improve the effectiveness of software testing.
Xiujing Guo, Hiroyuki Okamura, Tadashi Dohi
Softw. Qual. J.1