EDBT 2026 Demo / reviewers in the wild / expert
Donghwan Shin 0001
dblp:118/8296
· DBLP profile ↗
40ranked-venue papers
7as first author
28since 2021 · last 2026
0000-0002-0840-6449ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 38 · 7 first-author · 26 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Mutation Scheduling: Highly Parallel, Efficient Evaluation of Mutations for Rust Programs through Program Splitting
Zalán Lévai, Donghwan Shin 0001, Phil McMinn |
ICST | 2 |
| 2026 | mutest-rs: Flexible, Efficient Mutation Analysis Tool for Rust Programs, using Extensive Static AnalysisabstractDetermining the adequacy of software tests, and where testing gaps might lie, is crucial for improving and maintaining the strength of test suites. Mutation analysis facilitates this by evaluating tests against generated program faults; however, no mature mutation analysis tooling exists for the safety-focused Rust systems programming language, to date. This paper introduces mutest-rs, a mature, end-to-end mutation analysis tool for Rust programs that is based on extensive static program analysis, and integrates directly with the rustc Rust compiler. Our tool overcomes the numerous challenges of generating valid Rust code mutants, and does so efficiently through a Rustspecific meta-mutant approach. Our open-source tool, mutest-rs, is available online at https://mutest.rs. A video demonstrating mutest-rs is available at https://youtu.be/8yEYAU6P63I. Zalán Lévai, Donghwan Shin 0001, Phil McMinn |
ICST | 2 |
| 2026 | Multi-Fidelity Bayesian Optimization for Simulation Based Autonomous Driving Systems Testing
Olek Osikowicz, Phil McMinn, Donghwan Shin 0001 |
IV | 4 |
| 2026 | Automated testing of prevalent 3D user interactions in virtual reality applicationsabstractVirtual Reality (VR) technologies offer immersive user experiences across various domains, but present unique testing challenges compared to traditional software. Existing VR testing approaches enable scene navigation and interaction activation, but lack the ability to automatically synthesise realistic 3D user inputs (e.g, grab and trigger actions via hand-held controllers). Automated testing that generates and executes such input remains an unresolved challenge. Furthermore, existing metrics fail to robustly capture diverse interaction coverage. This paper addresses these gaps through four key contributions. First, we empirically identify four prevalent interaction types in nine open-source VR projects: fire , manipulate , socket , and custom . Second, we introduce the Interaction Flow Graph , a novel abstraction that systematically models 3D user interactions by identifying targets, actions, and conditions. Third, we construct XRBench3D , a benchmark comprising ten VR scenes that encompass 456 distinct user interactions for evaluating VR interaction testing. Finally, we present XRintTest , an automated testing approach that leverages this graph for dynamic scene exploration and interaction execution. Evaluation on XRBench3D shows that XRintTest achieves great effectiveness, reaching 93% coverage of fire , manipulate and socket interactions across all scenes, and performing 12x more effectively and 6x more efficiently than random exploration. Moreover, XRintTest can detect runtime exceptions and non-exception interaction issues, including subtle configuration defects. In addition, the Interaction Flow Graph can reveal potential interaction design smells that may compromise intended functionality and hinder testing performance for VR applications. Ruizhen Gu, José Miguel Rojas, Donghwan Shin 0001 |
Autom. Softw. Eng. | 3 |
| 2025 | Using Causal Inference to Test Systems with Hidden and Interacting Variables: An Evaluative Case StudyabstractSoftware systems with large parameter spaces, nondeterminism and high computational cost are challenging to test. Recently, software testing techniques based on causal inference have been successfully applied to systems that exhibit such characteristics, including scientific models and autonomous driving systems. One significant limitation is that these are restricted to test properties where all of the variables involved can be observed and where there are no interactions between variables. In practice, this is rarely guaranteed; the logging infrastructure may not be available to record all of the necessary runtime variable values, and it can often be the case that an output of the system can be affected by complex interactions between variables. To address this, we leverage two additional concepts from causal inference, namely effect modification and instrumental variable methods. We build these concepts into an existing causal testing tool and conduct an evaluative case study which uses the concepts to test three system-level requirements of CARLA, a high-fidelity driving simulator widely used in autonomous vehicle development and testing. The results show that we can obtain reliable test outcomes without requiring large amounts of highly controlled test data or instrumentation of the code, even when variables interact with each other and are not recorded in the test data. Michael Foster 0001, Robert M. Hierons, Donghwan Shin 0001, Neil Walkinshaw, Christopher Wild |
EASE | 3 |
| 2025 | Can Test Generation and Program Repair Inform Automated Assessment of Programming Projects?abstractComputer Science educators assessing student programming assignments are typically responsible for two challenging tasks: grading and providing feedback. Producing grades that are fair and feedback that is useful to students is a goal common to most educators. In this context, automated test generation and program repair offer promising solutions for detecting bugs and suggesting corrections in students' code which could be leveraged to inform grading and feedback generation. Previous research on the applicability of these techniques to simple programming tasks (e.g., single-method algorithms) has shown promising results, but their effectiveness for more complex programming tasks remains unexplored. To fill this gap, this paper investigates the feasibility of applying existing test generation and program repair tools for assessing complex programming assignment projects. In a case study using a real-world Java programming assignment project with 296 incorrect student submissions, we found that generated tests were insufficient in detecting bugs in over 50% of cases, while full repairs could only be automatically generated for only 2.1% of submissions. Our findings indicate significant limitations in current tools for detecting bugs and repairing student submissions, highlighting the need for more advanced techniques to support automated assessment of complex assignment projects. Ruizhen Gu, José Miguel Rojas, Donghwan Shin 0001 |
ICST | 3 |
| 2025 | XRintTest: An Automated Framework for User Interaction Testing in Extended Reality ApplicationsabstractExtended Reality (XR) technologies offer immersive user experiences across diverse application domains, presenting unique testing challenges due to their spatial interaction paradigms. While existing works test XR applications through scene navigation and interaction triggering, they fail to synthesise realistic spatial input via specialised XR devices, such as 6 degrees of freedom controller gestures, that are essential for modern XR user experiences. To address this gap, we present XRintTest, an automated testing framework for Unity-based XR applications. XRintTest starts by constructing an XR User Interaction Graph that models interaction targets and required events. Leveraging this graph, it then automatically explores the XR scene under test and generates user interactions. We evaluated XRintTest on XRBench3D, a novel benchmark comprising seven XR scenes containing 367 distinct 3D user interactions. XRintTest shows great effectiveness, achieving 97% coverage of trigger and grab interactions across all scenes, 9x more effective and 5x more efficient than random exploration, while detecting runtime exceptions and functional defects. We open-sourced our tool and dataset at https://github.com/ruizhengu/XRintTest and https://github.com/ruizhengu/XRBench3D, respectively. A video demo is available on YouTube at https://youtu.be/K0Q6waE47Us. Ruizhen Gu, José Miguel Rojas, Donghwan Shin 0001 |
ASE | 3 |
| 2025 | Unseen Data Detection using Routing Entropy in Mixture-of-Experts for Autonomous VehiclesabstractUnseen data that differ significantly from the training data can cause machine learning models to behave unpredictably, which is particularly problematic in safety-critical systems like autonomous vehicles. Detecting such data, commonly called out-of-distribution (OOD) data, is essential for ensuring the robustness of these models. Existing methods often rely on the model’s final output, which are limited since the model can be overconfident on unseen data. In this paper, we propose Routing Entropy, a novel OOD detection method that leverages the internal routing behavior of Mixture-of-Experts (MoE) models, a design increasingly adopted in modern neural networks. We hypothesize that MoE models exhibit high confidence routing for in-distribution (ID) inputs, but greater uncertainty for OOD inputs. We quantify this uncertainty by calculating the entropy of the routing scores for a given input. Experimental results on a MoE-based semantic segmentation model used for perception in autonomous driving demonstrate that Routing Entropy is effective on its own and, more importantly, provides a complementary signal to existing output-based methods. Combining Routing Entropy with an existing method significantly improves OOD detection performance. These results suggest that leveraging internal routing behavior of MoE models is a promising direction for robust OOD detection. Sang In Lee, Donghwan Shin 0001 |
ASE | 2 |
| 2025 | Software testing for extended reality applications: a systematic mapping studyabstractAbstract Extended Reality (XR) is an emerging technology spanning diverse application domains and offering immersive user experiences. However, its unique characteristics, such as six degrees of freedom interactions, present significant testing challenges distinct from traditional 2D GUI applications, demanding novel testing techniques to build high-quality XR applications. This paper presents the first systematic mapping study on software testing for XR applications. We selected 34 studies focusing on techniques and empirical approaches in XR software testing for detailed examination. The studies are classified and reviewed to address the current research landscape, test facets, and evaluation methodologies in the XR testing domain. Additionally, we provide a repository summarising the mapping study, including datasets and tools referenced in the selected studies, to support future research and practical applications. Our study highlights open challenges in XR testing and proposes actionable future research directions to address the gaps and advance the field of XR software testing. Ruizhen Gu, José Miguel Rojas, Donghwan Shin 0001 |
Autom. Softw. Eng. | 3 |
| 2024 | Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMsabstractLarge Language Models (LLMs) have shown remarkable capabilities in processing both natural and programming languages, which have enabled various applications in software engineering, such as requirement engineering, code generation, and software testing. However, existing code generation benchmarks do not necessarily assess the code understanding performance of LLMs, especially for the subtle inconsistencies that may arise between code and its semantics described in natural language. Donghwan Shin 0001 |
CAIN | 2 |
| 2024 | Autonomous Driving System Testing: Traffic Density Does Matter
Guannan Lou, Donghwan Shin 0001, Neil Walkinshaw, Robert M. Hierons |
ICTSS | 2 |
| 2024 | Systematic Evaluation of Deep Learning Models for Log-based Failure PredictionabstractAbstract With the increasing complexity and scope of software systems, their dependability is crucial. The analysis of log data recorded during system execution can enable engineers to automatically predict failures at run time. Several Machine Learning (ML) techniques, including traditional ML and Deep Learning (DL), have been proposed to automate such tasks. However, current empirical studies are limited in terms of covering all main DL types—Recurrent Neural Network (RNN), Convolutional Neural Network (CNN), and transformer—as well as examining them on a wide range of diverse datasets. In this paper, we aim to address these issues by systematically investigating the combination of log data embedding strategies and DL types for failure prediction. To that end, we propose a modular architecture to accommodate various configurations of embedding strategies and DL-based encoders. To further investigate how dataset characteristics such as dataset size and failure percentage affect model accuracy, we synthesised 360 datasets, with varying characteristics, for three distinct system behavioural models, based on a systematic and automated generation approach. Using the F1 score metric, our results show that the best overall performing configuration is a CNN-based encoder with Logkey2vec. Additionally, we provide specific dataset conditions, namely a dataset size $$>350$$ > 350 or a failure percentage $$>7.5\%$$ > 7.5 % , under which this configuration demonstrates high accuracy for failure prediction. Fatemeh Hadadi, Joshua Heneage Dawes, Donghwan Shin 0001, Domenico Bianculli, Lionel C. Briand |
Empir. Softw. Eng. | 3 |
| 2024 | Impact of log parsing on deep learning-based anomaly detectionabstractSoftware systems log massive amounts of data, recording important runtime information. Such logs are used, for example, for log-based anomaly detection, which aims to automatically detect abnormal behaviors of the system under analysis by processing the information recorded in its logs. Many log-based anomaly detection techniques based on deep learning models include a pre-processing step called log parsing. However, understanding the impact of log parsing on the accuracy of anomaly detection techniques has received surprisingly little attention so far. Investigating what are the key properties log parsing techniques should ideally have to help anomaly detection is therefore warranted. In this paper, we report on a comprehensive empirical study on the impact of log parsing on anomaly detection accuracy, using 13 log parsing techniques, seven anomly detection techniques (five based on deep learning and two based on traditional machine learning) on three publicly available log datasets. Our empirical results show that, despite what is widely assumed, there is no strong correlation between log parsing accuracy and anomaly detection accuracy, regardless of the metric used for measuring log parsing accuracy. Moreover, we experimentally confirm existing theoretical results showing that it is a property that we refer to as distinguishability in log parsing results-as opposed to their accuracy-that plays an essential role in achieving accurate anomaly detection. Zanis Ali Khan, Donghwan Shin 0001, Domenico Bianculli, Lionel C. Briand |
Empir. Softw. Eng. | 2 |
| 2024 | An Extensible Modeling Method Supporting Ontology-Based Scenario Specification and Domain-Specific ExtensionabstractScenario-based techniques, also known as scenario methods, have been actively employed to resolve intricate problems for engineering complex software systems. Scenarios are powerful tools that allow engineers to analyze the dynamics and contexts of complex systems. Despite the widespread use, there is a lack of a well-established reference framework that systematically organizes key concepts and attributes of scenarios. This has left engineers without a systematic guidance at the method level, hindering their ability to utilize the scenario methods effectively. To address the challenges associated with scenario methods, this study aims to provide a reference framework and modeling method. By conducting a literature review and suggesting a Conceptual Scenario Framework (CSF), we establish a conceptual basis that systematically presents the core concepts and characteristics of scenarios. Additionally, we introduce the Extensible Scenario Modeling Method (ESMM) that empowers engineers to perform scenario modeling and domain-specific extensions using the framework. With the inclusion of the Extensible Scenario Modeling Language (ESML), which comprises domain-general model types and classes for scenario description and ontological analysis, ESMM facilitates flexible design of domain-specific scenario elements through language-level extensions. This study assesses the proposed method in comparison to existing scenario development methods in the automated driving system domain. Through an analysis of their ability to represent scenario data, it was established that the language constructs of ESML possess semantic expressiveness suitable for serving as a reference framework. Furthermore, the findings from the case study validate the extensibility of ESMM for specialization in creating a scenario modeling language tailored to specific domains, while also effectively supporting the ontological analysis of particular application domains. Young Min Baek, Esther Cho, Donghwan Shin 0001, Doo-Hwan Bae |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2024 | Virtual Environment Model Generation for CPS Goal Verification using Imitation LearningabstractCyber-Physical Systems (CPS) continuously interact with their physical environments through embedded software controllers that observe the environments and determine actions. Field Operational Tests (FOT) are essential to verify to what extent the CPS under analysis can achieve certain CPS goals, such as satisfying the safety and performance requirements, while interacting with the real operational environment. However, performing many FOTs to obtain statistically significant verification results is challenging due to its high cost and risk in practice. Simulation-based verification can be an alternative to address the challenge, but it still requires an accurate virtual environment model that can replace the real environment interacting with the CPS in a closed loop. In this article, we propose ENVI (ENVironment Imitation), a novel approach to automatically generate an accurate virtual environment model, enabling efficient and accurate simulation-based CPS goal verification in practice.To do this, we first formally define the problem of the virtual environment model generation and solve it by leveraging Imitation Learning (IL), which has been actively studied in machine learning to learn complex behaviors from expert demonstrations. The key idea behind the model generation is to leverage IL for training a model that imitates the interactions between the CPS controller and its real environment as recorded in (possibly very small) FOT logs. We then statistically verify the goal achievement of the CPS by simulating it with the generated model. We empirically evaluate ENVI by applying it to the verification of two popular autonomous driving assistant systems. The results show that ENVI can reduce the cost of CPS goal verification while maintaining its accuracy by generating accurate environment models from only a few FOT logs. The use of IL in virtual environment model generation opens new research directions, further discussed at the end of the article. Donghwan Shin 0001, Doo-Hwan Bae |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2024 | Rigorous Assessment of Model Inference Accuracy using Language CardinalityabstractModels such as finite state automata are widely used to abstract the behavior of software systems by capturing the sequences of events observable during their execution. Nevertheless, models rarely exist in practice and, when they do, get easily outdated; moreover, manually building and maintaining models is costly and error-prone. As a result, a variety of model inference methods that automatically construct models from execution traces have been proposed to address these issues. However, performing a systematic and reliable accuracy assessment of inferred models remains an open problem. Even when a reference model is given, most existing model accuracy assessment methods may return misleading and biased results. This is mainly due to their reliance on statistical estimators over a finite number of randomly generated traces, introducing avoidable uncertainty about the estimation and being sensitive to the parameters of the random trace generative process. This article addresses this problem by developing a systematic approach based on analytic combinatorics that minimizes bias and uncertainty in model accuracy assessment by replacing statistical estimation with deterministic accuracy measures. We experimentally demonstrate the consistency and applicability of our approach by assessing the accuracy of models inferred by state-of-the-art inference tools against reference models from established specification mining benchmarks. Donato Clun, Donghwan Shin 0001, Antonio Filieri, Domenico Bianculli |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2023 | Towards Log SlicingabstractAbstract This short paper takes initial steps towards developing a novel approach, called log slicing, that aims to answer a practical question in the field of log analysis: Can we automatically identify log messages related to a specific message (e.g., an error message)? The basic idea behind log slicing is that we can consider how different log messages are “computationally related” to each other by looking at the corresponding logging statements in the source code. These logging statements are identified by 1) computing a backwards program slice, using as criterion the logging statement that generated a problematic log message; and 2) extending that slice to include relevant logging statements. The paper presents a problem definition of log slicing, describes an initial approach for log slicing, and discusses a key open issue that can lead towards new research directions. Joshua Heneage Dawes, Donghwan Shin 0001, Domenico Bianculli |
FASE | 2 |
| 2023 | Many-Objective Reinforcement Learning for Online Testing of DNN-Enabled SystemsabstractDeep Neural Networks (DNNs) have been widely used to perform real-world tasks in cyber-physical systems such as Autonomous Driving Systems (ADS). Ensuring the correct behavior of such DNN-Enabled Systems (DES) is a crucial topic. Online testing is one of the promising modes for testing such systems with their application environments (simulated or real) in a closed loop, taking into account the continuous interaction between the systems and their environments. However, the environmental variables (e.g., lighting conditions) that might change during the systems' operation in the real world, causing the DES to violate requirements (safety, functional), are often kept constant during the execution of an online test scenario due to the two major challenges: (1) the space of all possible scenarios to explore would become even larger if they changed and (2) there are typically many requirements to test simultaneously. In this paper, we present MORLOT (Many-Objective Rein-forcement Learning for Online Testing), a novel online testing approach to address these challenges by combining Reinforcement Learning (RL) and many-objective search. MORLOT leverages RL to incrementally generate sequences of environmental changes while relying on many-objective search to determine the changes so that they are more likely to achieve any of the uncovered objectives. We empirically evaluate MORLOT using CARLA, a high-fidelity simulator widely used for autonomous driving research, integrated with Transfuser, a DNN-enabled ADS for end-to-end driving. The evaluation results show that MORLOT is significantly more effective and efficient than alternatives with a large effect size. In other words, MORLOT is a good option to test DES with dynamically changing environments while accounting for multiple safety requirements. Fitash Ul Haq, Donghwan Shin 0001, Lionel C. Briand |
ICSE | 2 |
| 2023 | Identifying the Hazard Boundary of ML-Enabled Autonomous Systems Using Cooperative Coevolutionary SearchabstractIn Machine Learning (ML)-enabled autonomous systems (MLASs), it is essential to identify thehazard boundaryof ML Components (MLCs) in the MLAS under analysis. Given that such boundary captures the conditions in terms of MLC behavior and system context that can lead to hazards, it can then be used to, for example, build a safety monitor that can take any predefined fallback mechanisms at runtime when reaching the hazard boundary. However, determining suchhazard boundaryfor an ML component is challenging. This is due to the problem space combining system contexts (i.e., scenarios) and MLC behaviors (i.e., inputs and outputs) being far too large for exhaustive exploration and even to handle using conventional metaheuristics, such as genetic algorithms. Additionally, the high computational cost of simulations required to determine any MLAS safety violations makes the problem even more challenging. Furthermore, it is unrealistic to consider a region in the problem space deterministically safe or unsafe due to the uncontrollable parameters in simulations and the non-linear behaviors of ML models (e.g., deep neural networks) in the MLAS under analysis. To address the challenges, we propose MLCSHE (ML Component Safety Hazard Envelope), a novel method based on a Cooperative Co-Evolutionary Algorithm (CCEA), which aims to tackle a high-dimensional problem by decomposing it into two lower-dimensional search subproblems. Moreover, we take aprobabilisticview of safe and unsafe regions and define a novel fitness function to measure the distance from the probabilistic hazard boundary and thus drive the search effectively. We evaluate the effectiveness and efficiency of MLCSHE on a complex Autonomous Vehicle (AV) case study. Our evaluation results show that MLCSHE is significantly more effective and efficient compared to a standard genetic algorithm and random search. Sepehr Sharifi, Donghwan Shin 0001, Lionel C. Briand, Nathan Aschbacher |
IEEE Trans. Software Eng. | 2 |
| 2022 | Efficient Online Testing for DNN-Enabled Systems using Surrogate-Assisted and Many-Objective OptimizationabstractWith the recent advances of Deep Neural Networks (DNNs) in real-world applications, such as Automated Driving Systems (ADS) for self-driving cars, ensuring the reliability and safety of such DNN-enabled Systems emerges as a fundamental topic in software testing. One of the essential testing phases of such DNN-enabled systems is online testing, where the system under test is embedded into a specific and often simulated application environment (e.g., a driving environment) and tested in a closed-loop mode in interaction with the environment. However, despite the importance of online testing for detecting safety violations, automatically generating new and diverse test data that lead to safety violations presents the following challenges: (1) there can be many safety requirements to be considered at the same time, (2) running a high-fidelity simulator is often very computationally-intensive, and (3) the space of all possible test data that may trigger safety violations is too large to be exhaustively explored. Fitash Ul Haq, Donghwan Shin 0001, Lionel C. Briand |
ICSE | 2 |
| 2022 | Guidelines for Assessing the Accuracy of Log Message Template Identification TechniquesabstractLog message template identification aims to convert raw logs containing free-formed log messages into structured logs to be processed by automated log-based analysis, such as anomaly detection and model inference. While many techniques have been proposed in the literature, only two recent studies provide a comprehensive evaluation and comparison of the techniques using an established benchmark composed of real-world logs. Nevertheless, we argue that both studies have the following issues: (1) they used different accuracy metrics without comparison between them, (2) some ground-truth (oracle) templates are incorrect, and (3) the accuracy evaluation results do not provide any information regarding incorrectly identified templates. Zanis Ali Khan, Donghwan Shin 0001, Domenico Bianculli, Lionel C. Briand |
ICSE | 2 |
| 2022 | Correction to: Can Offline Testing of Deep Neural Networks Replace Their Online Testing?
Fitash Ul Haq, Donghwan Shin 0001, Shiva Nejati 0001, Lionel C. Briand |
Empir. Softw. Eng. | 2 |
| 2022 | PRINS: scalable model inference for component-based system logsabstractAbstract Behavioral software models play a key role in many software engineering tasks; unfortunately, these models either are not available during software development or, if available, quickly become outdated as implementations evolve. Model inference techniques have been proposed as a viable solution to extract finite state models from execution logs. However, existing techniques do not scale well when processing very large logs that can be commonly found in practice. In this paper, we address the scalability problem of inferring the model of a component-based system from large system logs, without requiring any extra information. Our model inference technique, called PRINS, follows a divide-and-conquer approach. The idea is to first infer a model of each system component from the corresponding logs; then, the individual component models are merged together taking into account the flow of events across components, as reflected in the logs. We evaluated PRINS in terms of scalability and accuracy, using nine datasets composed of logs extracted from publicly available benchmarks and a personal computer running desktop business applications. The results show that PRINS can process large logs much faster than a publicly available and well-known state-of-the-art tool, without significantly compromising the accuracy of inferred models. Donghwan Shin 0001, Domenico Bianculli, Lionel C. Briand |
Empir. Softw. Eng. | 1 |
| 2021 | Digital Twins Are Not Monozygotic - Cross-Replicating ADAS Testing in Two Industry-Grade Automotive SimulatorsabstractThe increasing levels of software- and data-intensive driving automation call for an evolution of automotive soft-ware testing. As a recommended practice of the Verification and Validation (V&V) process of ISO/PAS 21448, a candidate standard for safety of the intended functionality for road vehicles, simulation-based testing has the potential to reduce both risks and costs. There is a growing body of research on devising test automation techniques using simulators for Advanced Driver-Assistance Systems (ADAS). However, how similar are the results if the same test scenarios are executed in different simulators? We conduct a replication study of applying a Search-Based Software Testing (SBST) solution to a real-world ADAS (PeVi, a pedestrian vision detection system) using two different commercial simulators, namely, TASS/Siemens PreScan and ESI Pro-SiVIC. Based on a minimalistic scene, we compare critical test scenarios generated using our SBST solution in these two simulators. We show that SBST can be used to effectively generate critical test scenarios in both simulators, and the test results obtained from the two simulators can reveal several weaknesses of the ADAS under test. However, executing the same test scenarios in the two simulators leads to notable differences in the details of the test outputs, in particular, related to (1) safety violations revealed by tests, and (2) dynamics of cars and pedestrians. Based on our findings, we recommend future V&V plans to include multiple simulators to support robust simulation-based testing and to base test objectives on measures that are less dependant on the internals of the simulators. Markus Borg, Raja Ben Abdessalem, Shiva Nejati 0001, François-Xavier Jegeden, Donghwan Shin 0001 |
ICST | 5 |
| 2021 | Automatic test suite generation for key-points detection DNNs using many-objective search (experience paper)abstractAutomatically detecting the positions of key-points (e.g., facial key-points or finger key-points) in an image is an essential problem in many applications, such as driver's gaze detection and drowsiness detection in automated driving systems. With the recent advances of Deep Neural Networks (DNNs), Key-Points detection DNNs (KP-DNNs) have been increasingly employed for that purpose. Nevertheless, KP-DNN testing and validation have remained a challenging problem because KP-DNNs predict many independent key-points at the same time---where each individual key-point may be critical in the targeted application---and images can vary a great deal according to many factors. Fitash Ul Haq, Donghwan Shin 0001, Lionel C. Briand, Thomas Stifter, Jun Wang 0020 |
ISSTA | 2 |
| 2021 | Log-based slicing for system-level test casesabstractRegression testing is arguably one of the most important activities in software testing. However, its cost-effectiveness and usefulness can be largely impaired by complex system test cases that are poorly designed (e.g., test cases containing multiple test scenarios combined into a single test case) and that require a large amount of time and resources to run. One way to mitigate this issue is decomposing such system test cases into smaller, separate test cases---each of them with only one test scenario and with its corresponding assertions---so that the execution time of the decomposed test cases is lower than the original test cases, while the test effectiveness of the original test cases is preserved. This decomposition can be achieved with program slicing techniques, since test cases are software programs too. However, existing static and dynamic slicing techniques exhibit limitations when (1) the test cases use external resources, (2) code instrumentation is not a viable option, and (3) test execution is expensive. Salma Messaoudi, Donghwan Shin 0001, Annibale Panichella, Domenico Bianculli, Lionel C. Briand |
ISSTA | 2 |
| 2021 | A Theoretical Framework for Understanding the Relationship Between Log Parsing and Anomaly Detection
Donghwan Shin 0001, Zanis Ali Khan, Domenico Bianculli, Lionel C. Briand |
RV | 1 |
| 2021 | Can Offline Testing of Deep Neural Networks Replace Their Online Testing?abstractAbstract We distinguish two general modes of testing for Deep Neural Networks (DNNs): Offline testing where DNNs are tested as individual units based on test datasets obtained without involving the DNNs under test, and online testing where DNNs are embedded into a specific application environment and tested in a closed-loop mode in interaction with the application environment. Typically, DNNs are subjected to both types of testing during their development life cycle where offline testing is applied immediately after DNN training and online testing follows after offline testing and once a DNN is deployed within a specific application environment. In this paper, we study the relationship between offline and online testing. Our goal is to determine how offline testing and online testing differ or complement one another and if offline testing results can be used to help reduce the cost of online testing? Though these questions are generally relevant to all autonomous systems, we study them in the context of automated driving systems where, as study subjects, we use DNNs automating end-to-end controls of steering functions of self-driving vehicles. Our results show that offline testing is less effective than online testing as many safety violations identified by online testing could not be identified by offline testing, while large prediction errors generated by offline testing always led to severe safety violations detectable by online testing. Further, we cannot exploit offline testing results to reduce the cost of online testing in practice since we are not able to identify specific situations where offline testing could be as accurate as online testing in identifying safety requirement violations. Fitash Ul Haq, Donghwan Shin 0001, Shiva Nejati 0001, Lionel C. Briand |
Empir. Softw. Eng. | 2 |
| 2020 | Comparing Offline and Online Testing of Deep Neural Networks: An Autonomous Car Case StudyabstractThere is a growing body of research on developing testing techniques for Deep Neural Networks (DNNs). We distinguish two general modes of testing for DNNs: Offline testing where DNNs are tested as individual units based on test datasets obtained independently from the DNNs under test, and online testing where DNNs are embedded into a specific application and tested in a close-loop mode in interaction with the application environment. In addition, we identify two sources for generating test datasets for DNNs: Datasets obtained from real-life and datasets generated by simulators. While offline testing can be used with datasets obtained from either sources, online testing is largely confined to using simulators since online testing within real-life applications can be time consuming, expensive and dangerous. In this paper, we study the following two important questions aiming to compare test datasets and testing modes for DNNs: First, can we use simulator-generated data as a reliable substitute to real-world data for the purpose of DNN testing? Second, how do online and offline testing results differ and complement each other? Though these questions are generally relevant to all autonomous systems, we study them in the context of automated driving systems where, as study subjects, we use DNNs automating end-to-end control of cars' steering actuators. Our results show that simulator-generated datasets are able to yield DNN prediction errors that are similar to those obtained by testing DNNs with real-life datasets. Further, offline testing is more optimistic than online testing as many safety violations identified by online testing could not be identified by offline testing, while large prediction errors generated by offline testing always led to severe safety violations detectable by online testing. Fitash Ul Haq, Donghwan Shin 0001, Shiva Nejati 0001, Lionel C. Briand |
ICST | 2 |
| 2019 | Empirical evaluation of mutation-based test case prioritization techniquesabstractSummary In this paper, we propose a new test case prioritization technique that combines both mutation‐based and diversity‐aware approaches. The diversity‐aware mutation‐based technique relies on the notion of mutant distinguishment, which aims to distinguish one mutant's behaviour from another, rather than from the original program. The relative cost and effectiveness of the mutation‐based prioritization techniques (i.e., using both the traditional mutant kill and the proposed mutant distinguishment) are empirically investigated with 352 real faults and 553,477 developer‐written test cases. The empirical evaluation considers both the traditional and the diversity‐aware mutation criteria in various settings: single‐objective greedy, hybrid, and multi‐objective optimization. The results show that there is no single dominant technique across all the studied faults. To this end, the reason why each one of the mutation‐based prioritization criteria performs poorly is discussed, using a graphical model called Mutant Distinguishment Graph that demonstrates the distribution of the fault‐detecting test cases with respect to mutant kills and distinguishment. © 2018 John Wiley & Sons, Ltd. Donghwan Shin 0001, Shin Yoo, Mike Papadakis, Doo-Hwan Bae |
Softw. Test. Verification Reliab. | 1 |
| 2018 | Are mutation scores correlated with real fault detection?: a large scale empirical study on the relationship between mutants and real faultsabstractEmpirical validation of software testing studies is increasingly relying on mutants. This practice is motivated by the strong correlation between mutant scores and real fault detection that is reported in the literature. In contrast, our study shows that correlations are the results of the confounding effects of the test suite size. In particular, we investigate the relation between two independent variables, mutation score and test suite size, with one dependent variable the detection of (real) faults. We use two data sets, CoreBench and Defects4J, with large C and Java programs and real faults and provide evidence that all correlations between mutation scores and real fault detection are weak when controlling for test suite size. We also find that both independent variables significantly influence the dependent one, with significantly better fits, but overall with relative low prediction power. By measuring the fault detection capability of the top ranked, according to mutation score, test suites (opposed to randomly selected test suites of the same size), we find that achieving higher mutation scores improves significantly the fault detection. Taken together, our data suggest that mutants provide good guidance for improving the fault detection of test suites, but their correlation with fault detection are weak. Mike Papadakis, Donghwan Shin 0001, Shin Yoo, Doo-Hwan Bae |
ICSE | 2 |
| 2018 | A Theoretical and Empirical Study of Diversity-Aware Mutation Adequacy CriterionabstractDiversity has been widely studied in software testing as a guidance towards effective sampling of test inputs in the vast space of possible program behaviors. However, diversity has received relatively little attention in mutation testing. The traditional mutation adequacy criterion is a one-dimensional measure of the total number of killed mutants. We propose a novel, diversity-aware mutation adequacy criterion called distinguishing mutation adequacy criterion, which is fully satisfied when each of the considered mutants can be identified by the set of tests that kill it, thereby encouraging inclusion of more diverse range of tests. This paper presents the formal definition of the distinguishing mutation adequacy and its score. Subsequently, an empirical study investigates the relationship among distinguishing mutation score, fault detection capability, and test suite size. The results show that the distinguishing mutation adequacy criterion detects 1.33 times more unseen faults than the traditional mutation adequacy criterion, at the cost of a 1.56 times increase in test suite size, for adequate test suites that fully satisfies the criteria. The results show a better picture for inadequate test suites; on average, 8.63 times more unseen faults are detected at the cost of a 3.14 times increase in test suite size. Donghwan Shin 0001, Shin Yoo, Doo-Hwan Bae |
IEEE Trans. Software Eng. | 1 |
| 2016 | A Theoretical Framework for Understanding Mutation-Based Testing MethodsabstractIn the field of mutation analysis, mutation is the systematic generation of mutated programs (i.e., mutants) from an original program. The concept of mutation has been widely applied to various testing problems, including test set selection, fault localization, and program repair. However, surprisingly little focus has been given to the theoretical foundation of mutation-based testing methods, making it difficult to understand, organize, and describe various mutation-based testing methods. This paper aims to consider a theoretical framework for understanding mutation-based testing methods. While there is a solid testing framework for general testing, this is incongruent with mutation-based testing methods, because it focuses on the correctness of a program for a test, while the essence of mutation-based testing concerns the differences between programs (including mutants) for a test. In this paper, we begin the construction of our framework by defining a novel testing factor, called a test differentiator, to transform the paradigm of testing from the notion of correctness to the notion of difference. We formally define behavioral differences of programs for a set of tests as a mathematical vector, called a d-vector. We explore the multi-dimensional space represented by d-vectors, and provide a graphical model for describing the space. Based on our framework and formalization, we interpret existing mutation-based fault localization methods and mutant set minimization as applications, and identify novel implications for future work. Donghwan Shin 0001, Doo-Hwan Bae |
ICST | 1 |
| 2016 | Comprehensive analysis of FBD test coverage criteria using mutants
Donghwan Shin 0001, Eunkyoung Jee, Doo-Hwan Bae |
Softw. Syst. Model. | 1 |
| 2015 | Quality Based Software Project Staffing and Scheduling with Cost BoundabstractSoftware project planning is becoming more complicated and important as the size of software project grows. Many approaches have been proposed to help project managers by providing optimal staffing and scheduling in terms of minimizing the cost (i.e., necessary expanse) or time (i.e., time span or duration) required for the software project. Unfortunately, the software quality, another critical factor in software project planning, is largely overlooked in previous work. In this paper, we propose the quality based software project staffing and scheduling approach using a genetic algorithm (GA). We define a quality score by considering practical issues in software project planning in addition to task severity and defect amplification model. Further, the cost is utilized as a cost-bound in the GA to consider not only quality but also cost. Case study shows that the proposed approach improves the quality while the cost is optimized as the same as the cost-based approach. In other words, we provide better software project plans considering both cost and quality for software project managers. Also, we show the relationship between the quality and the cost in terms of software project planning. Dongwon Seo, Donghwan Shin 0001, Doo-Hwan Bae |
APSEC | 2 |
| 2015 | Efficient Testing of Self-Adaptive Behaviors in Collective Adaptive SystemsabstractCollective adaptive systems (CAS) consist of multiple agents that adapt to changing system and environmental conditions in order to satisfy system goals and quality requirements. As more applications involve using CAS in a critical context, ensuring the correct and safe adaptive behaviors of quality-driven CAS has become more important. In this paper, we propose Collective Adaptive System Testing (CAST), a scalable and efficient approach to testing self-adaptive behaviors of CAS. We propose a selective method to instantiate and execute test cases relevant to the current adaptation context. This enables testers to focus testing on key self-adaptive behaviors while dealing with the scale and dynamicity of the system. An experimental evaluation using a traffic monitoring system is performed to validate its scalability, efficiency, and fault-detection effectiveness. The experimental results provide insights into how CAST can serve as a feasible and effective assurance technique for CAS. Yoo Jin Lim, Eunkyoung Jee, Donghwan Shin 0001, Doo-Hwan Bae |
COMPSAC | 3 |
| 2015 | Human Resource Allocation in Software Project with Practical ConsiderationsabstractSoftware planning is very important for the success of a software project. Even if the same developers work on the same project, the time span of the project and the quality of software may change based on the project plan. When software managers plan a software project, they strive to allocate human resources in a more efficient way to produce a better software with less cost. The planning process is, however, time-consuming and complicated, especially when the size of the software project is large. Many approaches have been proposed to help software project managers by providing optimal human resource allocations in terms of minimizing the cost. Previous approaches, however, only concentrated on minimizing the cost, and no existing works have considered the practical issues affecting project schedules in practice. We elicited the practical considerations relating to the human resource allocation problem through discussions with a group of software project experts. The practical considerations can affect the project schedule in practice, but their importance has not been taken into consideration in previous approaches. Reflecting the practical considerations, we propose an approach for solving the human resource allocation problem using a genetic algorithm (GA). We compare our approach to an approach that only considers minimization of the time span. Our evaluation shows that the proposed algorithm considers the practical considerations well, in terms of continuous allocation on relevant tasks, minimization of developer multitasking time, and balance of allocation. We also conducted a survey targeting software developers and managers, and the responses showed that practical considerations are as important as minimizing the cost, and our approach would be helpful to software managers. We also investigate the effect of weight factors and coefficient between sub-scores, and find that it is difficult to consider some practical considerations at the same time. Dongwon Seo, Gwangui Hong, Donghwan Shin 0001, Jimin Hwa, Doo-Hwan Bae |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2014 | Practical Human Resource Allocation in Software Projects Using Genetic Algorithm
Dongwon Seo, Gwangui Hong, Donghwan Shin 0001, Jimin Hwa, Doo-Hwan Bae |
SEKE | 4 |
| 2014 | Automated test case generation for FBD programs implementing reactor protection system softwareabstractSUMMARY Automated and effective testing for function block diagram (FBD) programs has become an important issue, as FBD is increasingly used in implementing safety‐critical systems. This work describes an automated test case generation technique for FBD programs and its associated tool—FBDTester. Given an FBD program and desired test coverage criteria, FBDTester generates test requirements and invokes the Satisfiability Modulo Theories solver iteratively to derive a set of test cases. An industrial case study using reactor protection system software shows that the automatically generated test suites detected at least 82% of the known faults, whereas manually generated test cases only detected approximately 35%. Mutation analysis revealed that the automatically generated test suites substantially outperformed manually generated ones. Although test sequence generation requires some manual effort in the current FBDTester, it is apparent that the proposed approach significantly improves the efficiency and the reliability of FBD testing. Copyright © 2014 John Wiley & Sons, Ltd. Eunkyoung Jee, Donghwan Shin 0001, Sung Deok Cha, Jang-Soo Lee, Doo-Hwan Bae |
Softw. Test. Verification Reliab. | 2 |
| 2012 | Empirical Evaluation on FBD Model-Based Test Coverage Criteria Using Mutation Analysis
Donghwan Shin 0001, Eunkyoung Jee, Doo-Hwan Bae |
MoDELS | 1 |