Shin Hong

dblp:28/7539 · DBLP profile ↗
← Back
27ranked-venue papers
9as first author
13since 2021 · last 2025
0000-0003-4217-6031ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 26 · 8 first-author · 13 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 ZigZagFuzz: Interleaved Fuzzing of Program Options and Files
abstract
Command-line options (e.g., -l , -F , -R for ls ) given to a command-line program can significantly alternate the behaviors of the program. Thus, fuzzing not only file input but also program options can improve test coverage and bug detection. In this article, we propose ZigZagFuzz which achieves higher test coverage and detects more bugs than the state-of-the-art fuzzers by separately mutating program options and file inputs in an iterative/interleaving manner. ZigZagFuzz applies the following three core ideas. First, to utilize different characteristics of the program option domain and the file input domain, ZigZagFuzz separates phases of mutating program options from ones of mutating file inputs and performs two distinct mutation strategies on the two different domains. Second, to reach deep segments of a target program that are accessed through an interleaving sequence of program option checks and file inputs checks, ZigZagFuzz continuously interleaves phases of mutating program options with phases of mutating file inputs. Finally, to improve fuzzing performance further, ZigZagFuzz periodically shrinks input corpus by removing similar test inputs based on their function coverage. The experiment results on the 20 real-world programs show that ZigZagFuzz improves test coverage and detects 1.9 to 10.6 times more bugs than the state-of-the-art fuzzers that mutate program options such as AFL++-argv, AFL++-all, Eclipser, CarpetFuzz, ConfigFuzz, and POWER. We have reported the new bugs detected by ZigZagFuzz, and the original developers confirmed our bug reports.
Ahcheong Lee, Youngseok Choi, Shin Hong, Kyutae Cho, Moonzoo Kim
ACM Trans. Softw. Eng. Methodol.3
2024 BugOss: A benchmark of real-world regression bugs for empirical investigation of regression fuzzing techniques
Jeewoong Kim, Shin Hong
J. Syst. Softw.2
2023 Poster: BugOss: A Regression Bug Benchmark for Empirical Study of Regression Fuzzing Techniques
abstract
This paper presents BugOss, a benchmark of real-world regression bugs found in OSS-Fuzz for experimenting with regression fuzzing techniques. To reproduce the real project context where the bugs were introduced, each study artifact of BugOss indicates the exact bug-inducing commit, and provides the information about the target bug, together with the existing bugs in the same commit. The experiment results with five fuzzing techniques show that the 18 C/C++ artifacts currently registered for BugOss encompass various cases of regression bugs in real-world. We believe that BugOss offers a useful basis for empirically investigating regression fuzzing techniques.
Jeewoong Kim, Shin Hong
ICST2
2023 Introduction to the special issue on automation of software test and test code quality
abstract
We are pleased to present the papers selected for inclusion in the special issue devoted to the 1st International Conference on Automation of Software Test, which was held virtually in colocation with the 42nd IEEE/ACM International Conference on Software Engineering (ICSE 2020). Software testing is an integral and important part of the software engineering (SE) discipline. Over the past decades, a significant amount of SE research has focused on automation of software test (AST), including the automation of test case generation, test case selection and prioritization, test execution, test verdict analysis, and debugging. AST practice has also moved forward significantly, and in recent years, many test tools, frameworks, and methodologies have been developed and have significantly enhanced the quality assurance and the productivity in SE practices. Despite the significant achievements, AST remains challenging. To achieve total automation of software testing, a huge amount of code for the test cases and the supporting test infrastructure is needed, which yields itself in turn to large maintenance costs. However, if on the one side it is now generally accepted that disciplined procedures and quality standards should be applied in production code development, on the other side, comparable levels of rigor and quality are not demanded for the code written for testing that production code. Indeed, several recent empirical studies point to how test code is affected by many problems, bugs, unjustified assumptions, hidden dependencies from other tests or environment, flakiness, and performance issues. Because of such problems, test effectiveness is impacted and several false alarms are raised in regression testing that increase the test costs. Recently, both researchers and practitioners have proposed solutions toward this problem by identifying test code smells or test code quality issues and providing techniques to automatically detect and repair test code bugs and flakiness. In consideration of this active research thread, the 1st ACM/IEEE International Conference on Automation of Software Test (AST 2020) featured “Who Tests The Tests?” as the conference theme. AST 2020 was run for the first time in the format of a colocated conference to ICSE, after a successful series of 14 workshops under the same name. This special issue offers a venue for researchers and practitioners to share the advances in software test automation, especially on test code quality. In particular, we invited the authors of the best papers of AST 2020 to submit extended versions of their conference publications. Moreover, we also encouraged the authors and participants of the AST series and the broader SE community to submit novel original papers on the themes of software test automation and test code quality, including approaches for the following: understanding the dimensions and characteristics of problems with unreliable, low-quality test code; identifying and preventing test code smells, test code bugs, and flaky tests; impact of test code quality in scaling up test automation of very large, complex system; metrics for test code quality and robustness; automated repair of test code bugs and flakiness; and employing Artificial Intelligence and Machine Learning methods to help test automation. Each of the 10 submitted papers underwent a rigorous review process by independent referees to ensure that the paper had a sound, novel, and original contribution to the field of software test automation. Finally, five papers were accepted for inclusion in the special issue; three of these are extended versions of the best papers of AST 2020, and two are novel external submissions. In particular, the paper titled “Quantum Software Testing—State of the Art,” by Antonio García de la Barrera, Ignacio García-Rodríguez de Guzmán, Macario Polo, and Mario Piattini, provides a systematic mapping study assessing the state of the art in testing of quantum computing applications, indeed an emerging new paradigm that promises exponential speed up in solving highly demanding computational problems. In “An Empirical Study on How Sapienz Achieves Coverage and Crash Detection”, Iván Arcuschin, Juan Pablo Galeotti, and Diego Garbervetsky report the results from an empirical study aiming at better understanding how the main features of Sapienz, a powerful tool for automated testing of Android applications using evolutionary algorithms, impact its effectiveness. Many program analysis and testing tools need to incorporate string solver mechanisms. After observing that adequate tools and benchmarks for comparing existing solvers were lacking, in “ZaligVinder: A Generic Test Framework for String Solvers”, Mitja Kulczynski, Florin Manea, Dirk Nowotka, and Danny Bøgsted Poulsen propose an extensible framework gathering several string solver benchmarks that can be used for analysis and debugging purposes. “ExVivoMicroTest: Ex-Vivo Testing of Microservices” by Luca Gazzola, Maayan Goldstein, Leonardo Mariani, Marco Mobilio, Itai Segall, Alessandro Tundo, and Luca Ussi focuses on ex vivo testing of microservices during deployment. The authors claim that regression testing of such services may not be adequate prior to deployment due to a lack of knowledge regarding new use scenarios. The authors then propose a technique, named ExVivoMicroTest, that analyzes the behavior of deployed microservices and generates test cases for testing future updates. In “Fight Silent Horror Unit Test Methods by Consulting a TestWizard”, Maura Cerioli, Giovanni Lagorio, Maurizio Leotta, and Filippo Ricca focus on a practical problem, that is, the existence of incorrect tests. While a totally automated solution to identify invalid tests seems impractical, the idea of TestWizard is to assess individual tests' quality from the point of view of their coherence to specifications. We would like to thank the Editors-in-Chief of the Journal of Software: Evolution and Process for giving us the opportunity to publish this special issue. We appreciate all reviewers for their efforts in providing thoughtful and constructive comments to improve the quality of the publications. We also thank the authors of all submissions and publications of this special issue.
Antonia Bertolino, Shin Hong, Aditya P. Mathur
J. Softw. Evol. Process.2
2022 Inferring Fine-grained Traceability Links between Javadoc Comment and JUnit Test Code
abstract
This work presents DOTELINK, a technique that infers fine-grained traceability links between Javadoc comments and JUnit test code. To resolve the limitation of method-level traceability links, DOTELINK establishes links in sentence-level for Javadoc comments and code region-level for JUnit test methods. DOTELINK first segregates each Javadoc comment into multiple sentences, and each JUnit test method into coherent code snippets. And then, DOTELINK associates a Javadoc sentence with a code snippet if their lexical similarity is high. DOTELINK identifies 62.4% of the true fine-grained traceability links in the experiments with 5 real-world projects. We believe that DOTELINK effectively helps developers utilize the duality of the two sorts of requirement representations to improve test quality.
Jeewoong Kim, Shin Hong
ICSME2
2022 Learning-based Mutant Reduction Using Fine-grained Mutation Operators
abstract
This is an extended abstract of the article: Yunho Kim and SHin Hong, Learning-based Mutant Reduction Using Fine-grained Mutation Operators, Journal of Software Testing, Verification and Reliability, e1786, https://doi.org/10.1002/stvr.1786
Shin Hong
ICST2
2022 Repairing Fragile GUI Test Cases Using Word and Layout Embedding
abstract
Smartphone vendors apply both device and brand-specific customisations to the underlying operating systems, resulting in a wide range of device configurations. It is crucial that all of the device variations provide compatibility with the default version of the underlying operating system, such as Android. To ensure that widely and commonly used apps run on each of these device variations without any problem, vendors depend on automated GUI level testing of widely and commonly used apps: the failure of a GUI test script that emulates a routine usage of these apps would raise an alarm that a recent change made to a specific device variation may have caused a regression fault. These GUI level compatibility smoke tests are unique in the sense that they are GUI level automated test scripts that are written outside the software development life cycle of the target apps: they are written and maintained by the engineers of the smartphone vendors, and not the app developers. As such, these test scripts are extra vulnerable to the fragility of GUI test scripts, which are already known to be fragile when maintained by app developers. This paper introduces a repair technique for View Identification Failures (VIFs) in those smoke tests so that the smartphone vendors can quickly update their GUI test scripts when they break due to changed view ids. Our technique matches view ids between old and new versions of the target app based on various similarity metrics such as the semantic embedding similarity between ids and GUI labels, and layout similarity based on node embeddings of the GUI layout tree. We evaluate the proposed technique using 512 VIFs collected from real-world Android mobile apps. The proposed technique can repair 72 % of the 512 studied VIFs with only one attempt, compared to 28 % repaired using lexical distance-based matching.
Juyeon Yoon, Seungjoon Chung, Kihyuck Shin, Jinhan Kim, Shin Hong, Shin Yoo
ICST5
2022 Learning-based mutant reduction using fine-grained mutation operators
abstract
Summary For mutation testing, the huge cost of running test suites on a large number of mutants has been a serious obstacle. To resolve this problem, we propose a learning‐based mutant reduction techniqueMuTrain.MuTrainuses cost‐considerate linear regression (i.e., CLARS) to learn amutation model, which predicts the mutation score of a test suite based on the mutation testing results of a previous version of a target program. Then,MuTrainapplies the mutation model for subsequent versions to predict mutation scores with significantly fewer mutants. For effective mutant reduction and accurate mutation score prediction,MuTrainuses fine‐grained mutation operators refined from the existing coarse‐grained mutation operators. The experiment results show thatMuTrainreduces the number of mutants effectively (i.e., selecting only 1.6% of mutants). Moreover,MuTrainpredicts mutation score far more accurately than the existing mutant reduction techniques and random mutant selection. We also found thatMuTrainachieves much greater mutant reduction when it uses the fine‐grained mutation operators than the traditional coarse‐grained mutation operators (i.e., 1.6% vs. 14.6%).
Shin Hong
Softw. Test. Verification Reliab.2
2022 Predictive Mutation Analysis via the Natural Language Channel in Source Code
abstract
Mutation analysis can provide valuable insights into both the system under test and its test suite. However, it is not scalable due to the cost of building and testing a large number of mutants. Predictive Mutation Testing (PMT) has been proposed to reduce the cost of mutation testing, but it can only provide statistical inference about whether a mutant will be killed or not by the entire test suite. We propose Seshat, a Predictive Mutation Analysis (PMA) technique that can accurately predict the entire kill matrix , not just the Mutation Score (MS) of the given test suite. Seshat exploits the natural language channel in code, and learns the relationship between the syntactic and semantic concepts of each test case and the mutants it can kill, from a given kill matrix. The learnt model can later be used to predict the kill matrices for subsequent versions of the program, even after both the source and test code have changed significantly. Empirical evaluation using the programs in Defects4J shows that Seshat can predict kill matrices with an average F-score of 0.83 for versions that are up to years apart. This is an improvement in F-score by 0.14 and 0.45 points over the state-of-the-art PMT technique and a simple coverage-based heuristic, respectively. Seshat also performs as well as PMT for the prediction of the MS only. When applied to a mutant-based fault localisation technique, the predicted kill matrix by Seshat is successfully used to locate faults within the top 10 position, showing its usefulness beyond prediction of MS. Once Seshat trains its model using a concrete mutation analysis, the subsequent predictions made by Seshat are on average 39 times faster than actual test-based analysis. We also show that Seshat can be successfully applied to automatically generated test cases with an experiment using EvoSuite.
Jinhan Kim, Juyoung Jeon, Shin Hong, Shin Yoo
ACM Trans. Softw. Eng. Methodol.3
2021 Improving Mutation-Based Fault Localization with Plausible-code Generating Mutation Operators
abstract
This paper proposes a new mutation operator using neural network to generate plausible code elements to improve performance of mutation-based fault localization on omission faults. Unlike the existing mutation operators, the proposed mutation operator synthesizes new code elements at a given mutation site with a neural language model. We extended MUSE to use the proposed mutation operator, and conducted a case study with 3 omission faults found in JFreeChart of Defects4J. As a result, the accuracy of MUSE with the new mutation operator increased significantly in all three faults.
Juyoung Jeon, Shin Hong
ASE2
2021 Improving Configurability of Unit-level Continuous Fuzzing: An Industrial Case Study with SAP HANA
abstract
This paper presents industrial experiences on enhancing the configurability of a fuzzing framework for effective continuous fuzzing of the SAP HANA components. We propose five new mutation scheduling strategies for effective uses of grammar-aware mutators in the unit-level fuzzing framework, and three new seed corpus selection strategies to configure a fuzzing campaign to check on changed code in priority. The empirical results show that the proposed extension gives users chances to improve fuzzing effectiveness and efficiency by configuring the framework specifically for each target component.
Hanyoung Yoo, Jingun Hong, Lucas Bader, Dongwon Hwang, Shin Hong
ASE5
2021 Empirical Study of Effectiveness of EvoSuite on the SBST 2020 Tool Competition Benchmark
Robert Sebastian Herlim, Shin Hong, Moonzoo Kim
SSBSE2
2021 DEMINER: test generation for high test coverage through mutant exploration
abstract
Summary Most software testing techniques test a target program as it is and fail to utilize valuable information of diverse test executions on many variants/mutants of the original program in test generation. This paper proposes a new test generation technique DEMINER, which utilizes mutant executions to guide test generation on the original program for high test coverage. DEMINER first generates various mutants of an original target program and then extracts runtime information of mutant executions, which covered unreached branches by the mutation effects. Using the obtained runtime information, DEMINER inserts guideposts, artificial branches to replay the observed mutation effects, to the original target programs. Finally, DEMINER runs automated test generation on the original program with guideposts and achieves higher test coverage. We implemented DEMINER for C programs through software mutation and guided test generation such as concolic testing and fuzzing. We have shown the effectiveness of DEMINER on six real‐world target programs: Busybox‐ls, Busybox‐printf, Coreutils‐sort, GNU‐find, GNU‐grep and GNU‐sed. The experiment results show that DEMINER improved branch coverage by 63.4% and 19.6% compared with those of the conventional concolic testing techniques and the conventional fuzzing techniques on average, respectively.
Shin Hong
Softw. Test. Verification Reliab.2
2020 Using SMT Solver and Logic Puzzles for Teaching Computational Logics in Discrete Mathematics Class
abstract
Computational logics is one of the core languages for undergraduate students to build the fundamental knowledge and understanding of the computer science principles. However, unlike other beginner-level programming language courses, most students learn computational logics only with deduction or paper-and-pencil problem solving without any programming experiences. This poster shows our case of providing a SMT-solver (e.g., Z3) and programming problems of writing logic puzzle solvers (e.g., Sudoku, Numbrix) as effective learning materials that beginner-level students can understand and practice computational logics as rigorous programming languages. These problems involve students modeling a problem as a satisfiability problem, finding and specifying constraints of a problem as predicate logic formulas, and calling a SMT solver from the programs, and other relevant problem-solving activities. We present a set of SMT-solver-based programming assignments (with different difficulty-levels) designed for beginner-level computer science course, instructions for students to equip basic skills for using Z3, useful set-up to challenge and support students at the same time. We also discuss our teaching experiences, open issues and future work.
Shin Hong
SIGCSE1
2019 Classifying False Positive Static Checker Alarms in Continuous Integration Using Convolutional Neural Networks
abstract
Static code analysis in Continuous Integration (CI) environment can significantly improve the quality of a software system because it enables early detection of defects without any test executions or user interactions. However, being a conservative over-approximation of system behaviours, static analysis also produces a large number of false positive alarms, identification of which takes up valuable developer time. We present an automated classifier based on Convolutional Neural Networks (CNNs). We hypothesise that many false positive alarms can be classified by identifying specific lexical patterns in the parts of the code that raised the alarm: human engineers adopt a similar tactic. We train a CNN based classifier to learn and detect these lexical patterns, using a total of about 10K historical static analysis alarms generated by six static analysis checkers for over 27 million LOC, and their labels assigned by actual developers. The results of our empirical evaluation suggest that our classifier can be highly effective for identifying false positive alarms, with the average precision across all six checkers of 79.72%.
Seongmin Lee 0001, Shin Hong, Jungbae Yi, Taeksu Kim, Chul-Joo Kim, Shin Yoo
ICST2
2019 Target-driven compositional concolic testing with function summary refinement for effective bug detection
abstract
Concolic testing is popular in unit testing because it can detect bugs quickly in a relatively small search space. But, in system-level testing, it suffers from the symbolic path explosion and often misses bugs. To resolve this problem, we have developed a focused compositional concolic testing technique, FOCAL, for effective bug detection. Focusing on a target unit failure v (a crash or an assert violation) detected by concolic unit testing, FOCAL generates a system-level test input that validates v. This test input is obtained by building and solving symbolic path formulas that represent system-level executions raising v. FOCAL builds such formulas by combining function summaries one by one backward from a function that raised v to main. If a function summary φa of function a conflicts with the summaries of the other functions, FOCAL refines φa to φa′ by applying a refining constraint learned from the conflict. FOCAL showed high system-level bug detection ability by detecting 71 out of the 100 real-world target bugs in the SIR benchmark, while other relevant cutting edge techniques (i.e., AFL-fast, KATCH, Mix-CCBSE) detected at most 40 bugs. Also, FOCAL detected 13 new crash bugs in popular file parsing programs.
Shin Hong, Moonzoo Kim
ESEC/SIGSOFT FSE2
2018 Invasive Software Testing: Mutating Target Programs to Diversify Test Exploration for High Test Coverage
Shin Hong, Bongseok Ko, Duy Loc Phan, Moonzoo Kim
ICST2
2017 MUSEUM: Debugging real-world multilingual programs using mutation analysis
abstract
Context: The programming language ecosystem has diversified over the last few decades. Non-trivial programs are likely to be written in more than a single language to take advantage of various control/data abstractions and legacy libraries. Objective: Debugging multilingual bugs is challenging because language interfaces are difficult to use correctly and the scope of fault localization goes beyond language boundaries. To locate the causes of real-world multilingual bugs, this article proposes a mutation-based fault localization technique (MUSEUM). Method: MUSEUM modifies a buggy program systematically with our new mutation operators as well as conventional mutation operators, observes the dynamic behavioral changes in a test suite, and reports suspicious statements. To reduce the analysis cost, MUSEUM selects a subset of mutated programs and test cases. Results: Our empirical evaluation shows that MUSEUM is (i) effective: it identifies the buggy statements as the most suspicious statements for both resolved and unresolved non-trivial bugs in real-world multilingual programming projects; and (ii) efficient: it locates the buggy statements in modest amount of time using multiple machines in parallel. Also, by applying selective mutation analysis (i.e., selecting subsets of mutants and test cases to use), MUSEUM achieves significant speedup with marginal accuracy loss compared to the full mutation analysis. Conclusion: It is concluded that MUSEUM locates real-world multilingual bugs accurately. This result shows that mutation analysis can provide an effective, efficient, and language semantics agnostic analysis on multilingual code. Our light-weight analysis approach would play important roles as programmers write and debug large and complex programs in diverse programming languages.
Shin Hong, Taehoon Kwak, Byeongcheol Lee, Yiru Jeon, Bongseok Ko, Moonzoo Kim
Inf. Softw. Technol.1
2015 Systematic Testing of Reactive Software with Non-Deterministic Events: A Case Study on LG Electric Oven
abstract
Most home appliance devices such as electric ovens are reactive systems which repeat receiving a user input/event through an event handler, updating their internal state based on the input, and generating outputs. A challenge to test a reactive program is to check if the program correctly reacts to various non-deterministic sequence of events because an unexpected sequence of events may make the system fail due to the race conditions between the main loop and asynchronous event handlers. Thus, it is important to systematically generate/test various sequences of events by controlling the order of events and relative timing of event occurrences with respect to the main loop execution. In this paper, we report our industrial experience to solve the aforementioned problem by developing a systematic event generation framework based on concolic testing technique. We have applied the framework to a LG electric oven and detected several critical bugs including one that makes the oven ignore user inputs due to the illegal state transition.
Yongbae Park, Shin Hong, Moonzoo Kim, Dongju Lee
ICSE (2)2
2015 Mutation-Based Fault Localization for Real-World Multilingual Programs (T)
abstract
Programmers maintain and evolve their software in a variety of programming languages to take advantage of various control/data abstractions and legacy libraries. The programming language ecosystem has diversified over the last few decades, and non-trivial programs are likely to be written in more than a single language. Unfortunately, language interfaces such as Java Native Interface and Python/C are difficult to use correctly and the scope of fault localization goes beyond language boundaries, which makes debugging multilingual bugs challenging. To overcome the aforementioned limitations, we propose a mutation-based fault localization technique for real-world multilingual programs. To improve the accuracy of locating multilingual bugs, we have developed and applied new mutation operators as well as conventional mutation operators. The results of the empirical evaluation for six non-trivial real-world multilingual bugs are promising in that the proposed technique identifies the buggy statements as the most suspicious statements for all six bugs.
Shin Hong, Byeongcheol Lee, Taehoon Kwak, Yiru Jeon, Bongsuk Ko, Moonzoo Kim
ASE1
2015 A survey of race bug detection techniques for multithreaded programmes
abstract
Summary As multithreaded programmes become popular to fully utilize multicore CPUs, many race bug detection techniques have been developed to find concurrency errors in multithreaded programmes effectively. Because these techniques have different views on target programme execution and detect race bugs of various types, it is difficult to characterize, compare and improve race bug detection techniques. This paper presents a formalexecution model, which can uniformly represent various views of race bug detection techniques on target programme execution. Then, this paper classifies 43 race bug detection techniques according to their target race bugs. We classify race bugs on whether or not a bug violatesoperation blockspecification and/ordata associationspecification. This survey provides researchers with a clear top‐down view of various race bug detection techniques. In addition, the concrete examples of various race bugs in this survey can help field engineers avoid race bugs in their multithreaded programmes. Copyright © 2014 John Wiley & Sons, Ltd.
Shin Hong, Moonzoo Kim
Softw. Test. Verification Reliab.1
2015 Are concurrency coverage metrics effective for testing: a comprehensive empirical investigation
abstract
Summary Testing multithreaded programs is inherently challenging, as programs can exhibit numerous thread interactions. To help engineers test these programs cost‐effectively, researchers have proposed concurrency coverage metrics. These metrics are intended to be used as predictors for testing effectiveness and provide targets for test generation. The effectiveness of these metrics, however, remains largely unexamined. In this work, we explore the impact of concurrency coverage metrics on testing effectiveness and examine the relationship between coverage, fault detection, and test suite size. We study eight existing concurrency coverage metrics and six new metrics formed by combining complementary metrics. Our results indicate that the metrics are moderate to strong predictors of testing effectiveness and effective at providing test generation targets. Nevertheless, metric effectiveness varies across programs, and even combinations of complementary metrics do not consistently provide effective testing. These results highlight the need for additional work on concurrency coverage metrics. Copyright © 2014 John Wiley & Sons, Ltd.
Shin Hong, Matthew Staats, Moonzoo Kim, Gregg Rothermel
Softw. Test. Verification Reliab.1
2014 Detecting Concurrency Errors in Client-Side Java Script Web Applications
abstract
As web technologies have evolved, the complexity of dynamic web applications has increased significantly and web applications suffer concurrency errors due to unexpected orders of interactions among web browsers, users, the network, and so forth. In this paper, we present WAVE (Web Applications Virtual Environment), a testing framework to detect concurrency errors in client-side web applications written in JavaScript. WAVE generates various sequences of operations as test cases for a web application and executes a sequence of operations by dynamically controlling interactions of a target web application with the execution environment. We demonstrate that WAVE is effective and efficient for detecting concurrency errors through experiments on eight examples and five non-trivial real-world web applications.
Shin Hong, Yongbae Park, Moonzoo Kim
ICST1
2013 The Impact of Concurrent Coverage Metrics on Testing Effectiveness
abstract
When testing multithreaded programs, the number of possible thread interactions makes exploring all interactions infeasible in practice. In response, researchers have developed concurrent coverage metrics for multithreaded programs. These metrics allow them to estimate how well they have exercised concurrent program behavior, just as branch and statement coverage metrics do for sequential program testing. However, unlike sequential coverage metrics, the effectiveness of concurrent coverage metrics in testing remains largely unexamined. In this paper, we explore the relationship between concurrent coverage and fault detection effectiveness by studying the application of eight concurrent coverage metrics in testing nine concurrent programs. Our results show that existing concurrent coverage metrics are often moderate to strong predictors of concurrent testing effectiveness, and are generally reasonable targets for test suite generation. Nevertheless, their relative effectiveness as predictors and test generation targets varies across programs, and thus additional work is needed in this area.
Shin Hong, Matthew Staats, Moonzoo Kim, Gregg Rothermel
ICST1
2013 Effective pattern-driven concurrency bug detection for operating systems
Shin Hong, Moonzoo Kim
J. Syst. Softw.1
2012 Testing concurrent programs to achieve high synchronization coverage
abstract
The effectiveness of software testing is often assessed by measuring coverage of some aspect of the software, such as its code. There is much research aimed at increasing code coverage of sequential software. However, there has been little research on increasing coverage for concurrent software. This paper presents a new technique that aims to achieve high coverage of concurrent programs by generating thread schedules to cover uncovered coverage requirements. Our technique first estimates synchronization-pair coverage requirements, and then generates thread schedules that are likely to cover uncovered coverage requirements. This paper also presents a description of a prototype tool that we implemented in Java, and the results of a set of studies we performed using the tool on a several open-source programs. The results show that, for our subject programs, our technique achieves higher coverage faster than random testing techniques; the estimation-based heuristic contributes substantially to the effectiveness of our technique.
Shin Hong, Moonzoo Kim, Mary Jean Harrold
ISSTA1
2012 Understanding user understanding: determining correctness of generated program invariants
abstract
Recently, work has begun on automating the generation of test oracles, which are necessary to fully automate the testing process. One approach to such automation involves dynamic invariant generation which extracts invariants from program executions. To use such invariants as test oracles, however, it is necessary to distinguish correct from incorrect invariants, a process that currently requires human intervention. In this work we examine this process. In particular, we examine the ability of 30 users, across two empirical studies, to classify invariants generated from three Java programs. Our results indicate that users struggle to classify generated invariants: on average, they misclassify 9.1% to 31.7% of correct invariants and 26.1%-58.6% of incorrect invariants. These results contradict prior studies that suggest that classification by users is easy, and indicate that further work needs to be done to bridge the gap between the effectiveness of dynamic invariant generation in theory, and the ability of users to apply it in practice. Along these lines, we suggest several areas for future work.
Matthew Staats, Shin Hong, Moonzoo Kim, Gregg Rothermel
ISSTA2