EDBT 2026 Demo / reviewers in the wild / expert
Yifan Zhang 0016
dblp:57/4707-16
· DBLP profile ↗
11ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0003-1289-2192ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 10 · 7 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 6 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Standardized Evaluation of Metamorphic Relations: A Structured Rubric and Human-LLM Comparison
Yifan Zhang 0016, Dave Towey, Matthew Pike, Quang-Hung Luu, Huai Liu, Tsong Yueh Chen |
COMPSAC | 1 |
| 2025 | Evaluation of the Code Generated By Large Language Models: The State of the ArtabstractThe rapid development of Large Language Models (LLMs), such as ChatGPT and DeepSeek, has revolutionized software development, particularly in the domain of automated code generation. These models, built on architectures like the Transformer, have demonstrated remarkable capabilities in generating human-like text and source code, significantly enhancing developer productivity and reducing development time. However, the widespread adoption of LLMs for code generation raises concerns regarding the reliability, quality, and potential risks associated with the generated code. This article illustrates and analyzes the state of the art in evaluating LLM-generated code, summarizing research findings, and application areas. This paper highlights the challenges in distinguishing between machine-generated and human-written code, as well as the potential for LLMs to introduce security vulnerabilities and maintainability issues. We discuss the implications of these findings for both researchers and practitioners, emphasizing the need for continued research in the evaluation of LLM-generated code. Finally, we identify gaps in the literature and propose future research directions, such as the development of more robust benchmarks and improved evaluation metrics. By providing a thorough overview of the current landscape, this paper provides a valuable resource for researchers and practitioners interested in LLM’s code generation capabilities and limitations. We also highlight the importance of ongoing evaluation and refinement of these models to ensure their safe and effective integration into software-development practices. Zhihao Ying, Dave Towey, Yifan Zhang 0016 |
COMPSAC | 3 |
| 2025 | Comparative Analysis of Styles in LLM-Generated Code for LeetCode Problems: A Preliminary StudyabstractLarge language models (LLMs) have rapidly become a powerful tool in automated code generation, yet most research has focused on their correctness and efficiency rather than the stylistic patterns of their outputs. In this preliminary study, we analyze the code patterns generated by five popular LLMs—ChatGPT, Gemini, Claude, Grok, and DeepSeek—in their free versions, across three LeetCode problems, one top-ranking each from the easy, medium, and hard categories. Our evaluation employs key metrics including inline comment density, naming conventions, and edge case handling, highlighting both similarities and differences in verbosity, comprehensibility, and robustness among the codes generated by models. The findings of this study have important implications for software engineering and education, suggesting that LLM-generated code can serve as both a tool for rapid prototyping and an effective learning resource for beginners. Our future work will extend this analysis to a broader set of coding challenges and compare LLM outputs with human-written code to develop robust criteria for evaluating automated code generation. Yifan Zhang 0016, Tsong Yueh Chen, Rubing Huang, Matthew Pike, Dave Towey, Zhihao Ying, Zhiquan Zhou 0001 |
COMPSAC | 1 |
| 2025 | Exploring the Black-Box: Testing Image Synthesis Systems through Metamorphic ExplorationabstractThe increasing complexity of deep learning models, especially in black-box scenarios, presents significant challenges to traditional software testing methods. Due to the lack of transparency in neural networks’ decision-making processes and the non-deterministic nature of model outputs, traditional test oracle approaches become inadequate. To address this problem, Metamorphic Testing (MT) and its extended approach, Metamorphic Exploration (ME), provide new ideas for validating deep learning systems by defining Metamorphic Relations (MR) between inputs and outputs. However, existing image transformation-based MR faces new challenges in image synthesis scenarios, as these operations may destroy the contextual information and affect the model’s performance. This paper proposes a novel ME design for deep learning image synthesis networks and demonstrates its effectiveness using a visible-infrared image fusion network as the case study. The result identifies the performance degradation problem due to the tensor dimension manipulation error, which indicates that the ME not only detects defects but also helps developers deeply understand the internal mechanisms of complex systems through the Hypothesized Metamorphic Relation (HMR), thus providing unique value for software quality assurance (SQA) of AI-driven software. Zhihao Ying, Yifan Zhang 0016, Qian Zhang 0018, Dave Towey |
COMPSAC | 3 |
| 2025 | Enhancing autonomous driving simulations: A hybrid metamorphic testing framework with metamorphic relations generated by GPT
Yifan Zhang 0016, Tsong Yueh Chen, Matthew Pike, Dave Towey, Zhihao Ying, Zhiquan Zhou 0001 |
Inf. Softw. Technol. | 1 |
| 2024 | Enabling Effective Metamorphic- Relation Generation by Novice Testers: A Pilot StudyabstractThis paper presents a pilot study that examines the capacity of novice testers to generate Metamorphic Relations (MRs) for autonomous driving systems (ADSs), specifically fo-cusing on parking functions. By comparing MRs generated by human participants with those generated by artificial intelligence (AI), we seek to understand the variances in quality, particularly in terms of correctness, applicability, novelty, and utility. Our findings indicate that despite receiving only minimal training, human participants were capable of producing MRs with a wide range of effectiveness. Notably, humans exhibited a potential for creative thinking, contrasting with AI's ability to generate MRs that adhere closely to technical and applicability standards. The study underscores the need for improved educational strategies aimed at enhancing the quality and confidence of MRs produced by humans. Future research directions will explore the optimization of training approaches, particularly within a constrained timeframe to create a positive learning experience and maintain participant engagement, to fully harness the creative capabilities of human learners in the context of ADS testing. Yifan Zhang 0016, Dave Towey, Matthew Pike |
COMPSAC | 1 |
| 2024 | Enhancing ADS Testing: An Open Educational Resource for Metamorphic TestingabstractThis study introduces a website serving as an Open Educational Resource (OER), dedicated to Metamorphic Testing (MT) and Metamorphic Relation (MR) generation, with a specific focus on Autonomous Driving Systems (ADSs). It offers a comprehensive introduction to MT and ADSs, and presents a specially designed scenario template that simplifies the MR generation process for ADS functions. This template enhances accessibility, making it more user-friendly for a wider audience, and facilitates systematic application, ensuring that users can apply test case and MR generation in a structured and organized manner. The MR generation guidelines that work with the template lower the learning barrier for beginners in MT, thus facilitating easier adoption and application of MT to ADSs. Yifan Zhang 0016, Dave Towey, Matthew Pike, Zhiquan Zhou 0001, Tsong Yueh Chen |
COMPSAC | 1 |
| 2024 | Scenario-Driven Metamorphic Testing for Autonomous Driving SimulatorsabstractABSTRACT The proliferation of driver‐assistance features in vehicles has resulted in a growing interest among the public in fully autonomous driving systems (ADSs). However, the integration of software and hardware in these complex systems presents significant testing challenges, particularly with respect to ensuring passenger safety. To address these challenges, simulation has emerged as a crucial step in the testing of ADSs. This paper presents a solution to the challenges faced in testing ADSs, with a focus on the validation of ADS simulators. The proposed approach involves using simulations and metamorphic testing (MT) to generate multiple concrete metamorphic relations (MRs) for testing ADS simulators. In order to accomplish this goal, we introduce three metamorphic relation patterns (MRPs). Each MRP is accompanied by a metamorphic relation input pattern (MRIP) that aids in generating detailed MRs. These MRs are designed to identify potential issues within the ADS simulator. To simplify the testing process and facilitate MT for testers, a self‐evolving scenario‐testing framework is also presented. The framework allows testers to improve test cases and MRs iteratively until issues detected are confirmed. The benefits and limitations of the framework are demonstrated using an industry case study. Overall, this study offers a practical solution to the challenges in testing ADSs and provides useful insights into improving testing efficiency for researchers and practitioners in the field. Yifan Zhang 0016, Dave Towey, Matthew Pike, Jia Cheng Han, Zhiquan Zhou 0001, Chenghao Yin |
Softw. Test. Verification Reliab. | 1 |
| 2023 | Metamorphic Testing of an Automated Parking System: An Experience ReportabstractAutomated Driving Systems (ADSs) have gained popularity recently. However, the unstable and unsafe ADSs have caused many traffic accidents and received widespread attention. One way to alleviate such issues is to enhance the correctness and efficiency of testing ADSs. Due to the difficulty of checking ADSs’ behavior such as parking the car, confirming the correctness of the actual behavior may be non-trivial or impossible. This kind of problem is called the test oracle problem. Unlike traditional software testing, Metamorphic Testing (MT) does not focus on the correctness of the actual strategy but examines whether or not the inputs and outputs of multiple executions of a Software Under Test (SUT) satisfy certain relations of the SUT, called Metamorphic Relations (MRs). The paper also implements Mutation Analysis (MA) on Baidu Apollo ADS to evaluate our MT. MA involves small modifications to a program’s source code to see if test-cases can detect these changes. This work was part of a larger endeavour to create an Open Educational Resource (OER) to support learning about how to apply MT to ADSs. This paper reports on an experience of implementing MT to test the Automated Parking System (APS) of Apollo ADS and applying MA to evaluate the MT. Dave Towey, Zepei Luo, Ziqi Zheng, Peijian Zhou, Junbo Yang, Puttipatt Ingkasit, Changyang Lao, Matthew Pike, Yifan Zhang 0016 |
COMPSAC | 9 |
| 2023 | Automated Metamorphic-Relation Generation with ChatGPT: An Experience Report
Yifan Zhang 0016, Dave Towey, Matthew Pike |
COMPSAC | 1 |
| 2022 | Preparing Future SQA Professionals: An Experience Report of Metamorphic Exploration of an Autonomous Driving SystemabstractComputing systems are becoming increasingly complex and sophisticated. Technologies such as artificial intelligence, big data, and autonomous vehicles are pushing the boundaries of system size, complexity, and comprehensibility beyond anything seen before. These advances, however, have left the associated software quality assurance (SQA) tools and processes behind. This is compounded by many training and education programs also not attempting to address this inadequacy in the preparation of future software engineering professionals. We face a situation of extensively-deployed advanced computing systems, many of which lack sufficient SQA support. Metamorphic Testing (MT) and Metamorphic Exploration (ME) are SQA approaches that have a record of being able to alleviate some of the challenges associated with the advanced computer systems. This paper reports on an MT/ME experience with the Baidu Apollo autonomous driving system (ADS). The experience included identifying an apparent problem in Apollo, which was later confirmed to be a misunderstanding, but which illustrated the potential for ME to scaffold learning how to perform SQA on such complex systems. The report will be of benefit not only to other ADS developers and testers, but also to other SQA professionals, and especially to SQA trainers and educators. Yifan Zhang 0016, Matthew Pike, Dave Towey, Jia Cheng Han, Zhiquan Zhou 0001 |
EDUCON | 1 |