Steven Cho

dblp:380/3427 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0001-2548-4406ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Metamorphic Testing of Large Language Models for Natural Language Processing
abstract
Using Large Language Models (LLMs) to perform Natural Language Processing (NLP) tasks has been becoming increasingly pervasive in recent times. The versatile nature of LLMs makes them applicable to a wide range of such tasks. While the performance of recent LLMs is generally outstanding, several studies have shown that LLMs can often produce incorrect results. Automatically identifying these faulty behaviors is extremely useful for improving the effectiveness of LLMs. One obstacle to this is the limited availability of labeled datasets, necessitating an oracle to determine the correctness of LLM behaviors. Metamorphic Testing (MT) is a popular testing approach that alleviates this oracle problem. At the core of MT are Metamorphic Relations (MRs), defining the relationship between the outputs of related inputs. MT can expose faulty behaviors without the need for explicit oracles (e.g., labeled datasets). This paper presents the most comprehensive study of MT for LLMs to date. We conducted a literature review and collected 191 MRs for NLP tasks. We implemented a representative subset (36 MRs) to conduct a series of experiments with three popular LLMs, running$\sim 560 ~\mathrm{K}$metamorphic tests. The results shed light on the capabilities and opportunities of MT for LLMs, as well as its limitations.
Steven Cho, Stefano Ruberto, Valerio Terragni
ICSME1
2025 MDPMorph: An MDP-Based Metamorphic Testing Framework for Deep Reinforcement Learning Agents
abstract
Deep Reinforcement Learning (DRL) systems are widely used across various domains. However, testing these systems presents significant challenges. The DRL agent, which serves as the core decision-maker, generates continuous value estimates rather than discrete labels and operates under nonstationary policies within complex and stochastic environments. Consequently, there is no definitive “correct answer” for each state-action pair, complicating automated test generation due to the known oracle problem. To address this challenge, we propose a Metamorphic Testing (MT) framework (MDPMORPH) specifically designed for validating DRL agents. Our framework is based on Markov Decision Processes (MDP) and focuses on the core reasoning properties of agents to automatically uncover potential faults. To support MDPMORPH, we introduce a Metamorphic Relation (MR) design methodology tailored for DRL agents, based on the temporal characteristics of MDP. Using this method, we define nine generic MRs that encapsulate common and expected properties of an agent’s reasoning process. Furthermore, based on established assumptions and definitions within the context of MDPs, we theoretically demonstrate the soundness of these MRs. Finally, we specialize these generic MRs into environment-specific MRs by determining appropriate thresholds through training on three classic DRL environments. Our experimental results demonstrate that MDPMORPH and the proposed MRs are highly effective in automatically detecting mutants within these studied environments, with a 0.84 average mutation detection rate.
Yuning Xing, Daixu Ren, Steven Cho, Valerio Terragni
ISSRE5
2025 LLMorph: Automated Metamorphic Testing of Large Language Models
abstract
Automated testing is essential for evaluating and improving the reliability of Large Language Models (LLMs), yet the lack of automated oracles for verifying output correctness remains a key challenge. We present LLMorph, an automated testing tool specifically designed for LLMs performing NLP tasks, which leverages Metamorphic Testing (MT) to uncover faulty behaviors without relying on human-labeled data. MT uses Metamorphic Relations (MRs) to generate follow-up inputs from source test input, enabling detection of inconsistencies in model outputs without the need of expensive labelled data. LLMorph is aimed at researchers and developers who want to evaluate the robustness of LLM-based NLP systems. In this paper, we detail the design, implementation, and practical usage of LLMorph, demonstrating how it can be easily extended to any LLM, NLP task, and set of MRs. In our evaluation, we applied 36 MRs across four NLP benchmarks, testing three state-of-the-art LLMs: GPT-4, LLAMA3, and HERMES 2. This produced over 561,000 test executions. The results demonstrate LLMorph’s effectiveness in automatically exposing incorrect model behaviors at scale.The tool source code is available at https://github.com/steven-b-cho/llmorph. A screencast demo is available at https://youtu.be/sHmqdieCfw4.
Steven Cho, Stefano Ruberto, Valerio Terragni
ASE1
2025 Metamorphic Testing of Deep Reinforcement Learning Agents with MDPMorph
abstract
We present MDPMorph, a tool for metamorphic testing of Deep Reinforcement Learning (DRL) agents. MDPMorph is based on the Markov Decision Process (MDP) and targets the core reasoning properties of DRL agents to automatically uncover potential faults. It can generate metamorphic test suites and corresponding mutants directly from the DRL system under test. MDPMorph uses a subset of the metamorphic test suite and models to train the thresholds of the nine proposed Metamorphic Relations (MRs) using stochastic gradient descent. These MRs are based on the temporal characteristics of the MDP, and the training aims to determine the optimal threshold for each MR. After obtaining the optimal threshold, MDPMorph leverages the MRs to compare the execution results of different metamorphic test suites on the model under test and reports whether each test passes or fails. Finally, by collecting the execution results, MDPMorph calculates the mutant detection rate of MR to validate its effectiveness. Experimental results show that MDPMorph and the proposed MRs are highly effective in automatically detecting seeded faults (mutants).
Yuning Xing, Daixu Ren, Steven Cho, Valerio Terragni
ASE5