Marinos Kintis

dblp:78/9105 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
2since 2021 · last 2022
0000-0002-9068-6514ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 16 · 8 first-author · 2 since 2021
YearPublicationVenuePosition
2022 Mutation analysis and its industrial applications
abstract
We are pleased to introduce the first part of the special issue on Mutation Testing of the Journal of Software Testing, Verification & Reliability. The issue features three papers selected by a review board of 40 mutation testing experts. Our board was selected from an extensive pool of experts comprising of reviewers from the Mutation workshop series, authors of published manuscripts in established journals related to mutation testing as well as the specific subdomain where mutation was applied. Mutation Testing was introduced in 1970s as a rigorous and automated alternative to fault injection. With over four decades of research and development, it is now accepted as a premier measure of the fault reveling potential of test suites. Mutation Testing is a fault-based testing technique that introduces candidate faults by making syntactic changes. The fault revealing ability of a test suite is then measured as the ratio of mutants it can detect against the total number of detectable mutants. A test suite is mutation testing adequate if the test suite detects all detectable mutants. While mutation testing is usually limited to simple faults, the twin axioms of mutation testing—finite neighborhood and coupling effect—provide confidence that a test suite that is mutation adequate can also reveal complex faults composed of simple faults. While the field of mutation testing has seen active research, numerous important problems remain open. These include (1) how to adapt mutation testing to non-traditional fields such as machine learning, (2) how to apply mutation testing for automatic test generation such as fuzzers, and (3) how to effectively solve the question of equivalent mutants. We thank the editors, authors, and reviewers for their efforts to successfully complete this special issue.
Rahul Gopinath, Jie Zhang 0050, Marinos Kintis, Mike Papadakis
Softw. Test. Verification Reliab.3
2022 Mutation analysis and its industrial applications
abstract
We are pleased to introduce the second part of the special issue on Mutation Testing of the Journal of Software Testing, Verification & Reliability. We thank the editors, authors, and reviewers for their efforts to successfully complete this special issue.
Rahul Gopinath, Jie Zhang 0050, Marinos Kintis, Mike Papadakis
Softw. Test. Verification Reliab.3
2019 An Industrial Study on the Differences between Pre-Release and Post-Release Bugs
abstract
Software bugs constitute a frequent and common issue of software development. To deal with this problem, modern software development methodologies introduce dedicated quality assurance procedures. At the same time researchers aim at developing techniques capable of supporting the early discovery and fix of bugs. One important factor that guides such research attempts is the characteristics of software bugs and bug fixes. In this paper, we present an industrial study on the characteristics and differences between pre-release bugs, i.e. bugs detected during software development, and post-release bugs, i.e. bugs that escaped to production. Understanding such differences is of paramount importance as it will improve our understanding on the testing and debugging support that practitioners require from the research community, on the validity of the assumptions of several research techniques, and, most importantly, on the reasons why bugs escape to production. To this end, we analyze 37 industrial projects from BGL BNP Paribas and document the differences between pre-release bugs and post-release bugs. Our findings suggest that post-release bugs are more complex to fix, requiring developers to modify several source code files, written in different programming languages, and configuration files, as well. We also find that approximately 82% of the post-release bugs involve code additions and can be characterized as 'omission' bugs. Finally, we conclude the paper with a discussion on the implications of our study and provide guidance to future research directions.
Renaud Rwemalika, Marinos Kintis, Mike Papadakis, Yves Le Traon, Pierre Lorrach
ICSME2
2019 On the Evolution of Keyword-Driven Test Suites
abstract
Many companies rely on software testing to verify that their software products meet their requirements. However, test quality and, in particular, the quality of end-to-end testing is relatively hard to achieve. The problem becomes challenging when software evolves, as end-to-end test suites need to adapt and conform to the evolved software. Unfortunately, end-to-end tests are particularly fragile as any change in the application interface, e.g., application flow, location or name of graphical user interface elements, necessitates a change in the tests. This paper presents an industrial case study on the evolution of Keyword-Driven test suites, also known as Keyword-Driven Testing (KDT). Our aim is to demonstrate the problem of test maintenance, identify the benefits of Keyword-Driven Testing and overall improve the understanding of test code evolution (at the acceptance testing level). This information will support the development of automatic techniques, such as test refactoring and repair, and will motivate future research. To this end, we identify, collect and analyze test code changes across the evolution of industrial KDT test suites for a period of eight months. We show that the problem of test maintenance is largely due to test fragility (most commonly-performed changes are due to locator and synchronization issues) and test clones (over 30% of keywords are duplicated). We also show that the better test design of KDT test suites has the potential for drastically reducing (approximately 70%) the number of test code changes required to support software evolution. To further validate our results, we interview testers from BGL BNP Paribas and report their perceptions on the advantages and challenges of keyword-driven testing.
Renaud Rwemalika, Marinos Kintis, Mike Papadakis, Yves Le Traon, Pierre Lorrach
ICST2
2019 Ukwikora: continuous inspection for keyword-driven testing
abstract
Automation of acceptance test suites becomes necessary in the context of agile software development practices, which require rapid feedback on the quality of code changes. To this end, companies try to automate their acceptance tests as much as possible. Unfortunately, the growth of the automated test suites, by several automation testers, gives rise to potential test smells, i.e., poorly designed test code, being introduced in the test code base, which in turn may increase the cost of maintaining the code and creating new one. In this paper, we investigate this problem in the context of our industrial partner, BGL BNP Paribas, and introduce Ukwikora, an automated tool that statically analyzes acceptance test suites, enabling the continuous inspection of the test code base. Ukwikora targets code written in the Robot Framework syntax, a popular framework for writing Keyword-Driven tests. Ukwikora has been successfully deployed at BGL BNP Paribas, detecting issues otherwise unknown to the automation testers, such as the presence of duplicated test code, dead test code and dependency issues among the tests. The success of our case study reinforces the need for additional research and tooling for acceptance test suites.
Renaud Rwemalika, Marinos Kintis, Mike Papadakis, Yves Le Traon, Pierre Lorrach
ISSTA2
2018 Are mutants really natural?: a study on how "naturalness" helps mutant selection
abstract
Background: Code is repetitive and predictable in a way that is similar to the natural language. This means that code is "natural" and this "naturalness" can be captured by natural language modelling techniques. Such models promise to capture the program semantics and identify source code parts that `smell', i.e., they are strange, badly written and are generally error-prone (likely to be defective). Aims: We investigate the use of natural language modelling techniques in mutation testing (a testing technique that uses artificial faults). We thus, seek to identify how well artificial faults simulate real ones and ultimately understand how natural the artificial faults can be. Our intuition is that natural mutants, i.e., mutants that are predictable (follow the implicit coding norms of developers), are semantically useful and generally valuable (to testers). We also expect that mutants located on unnatural code locations (which are generally linked with error-proneness) to be of higher value than those located on natural code locations. Method: Based on this idea, we propose mutant selection strategies that rank mutants according to a) their naturalness (naturalness of the mutated code), b) the naturalness of their locations (naturalness of the original program statements) and c) their impact on the naturalness of the code that they apply to (naturalness differences between original and mutated statements). We empirically evaluate these issues on a benchmark set of 5 open-source projects, involving more than 100k mutants and 230 real faults. Based on the fault set we estimate the utility (i.e. capability to reveal faults) of mutants selected on the basis of their naturalness, and compare it against the utility of randomly selected mutants. Results: Our analysis shows that there is no link between naturalness and the fault revelation utility of mutants. We also demonstrate that the naturalness-based mutant selection performs similar (slightly worse) to the random mutant selection. Conclusions: Our findings are negative but we consider them interesting as they confute a strong intuition, i.e., fault revelation is independent of the mutants' naturalness.
Matthieu Jimenez, Thierry Titcheu Chekam, Maxime Cordy, Mike Papadakis, Marinos Kintis, Yves Le Traon, Mark Harman
ESEM5
2018 How effective are mutation testing tools? An empirical analysis of Java mutation testing tools with manual analysis and real faults
Marinos Kintis, Mike Papadakis, Andreas Papadopoulos, Evangelos Valvis, Nicos Malevris, Yves Le Traon
Empir. Softw. Eng.1
2018 Detecting Trivial Mutant Equivalences via Compiler Optimisations
abstract
Mutation testing realises the idea of fault-based testing, i.e., using artificial defects to guide the testing process. It is used to evaluate the adequacy of test suites and to guide test case generation. It is a potentially powerful form of testing, but it is well-known that its effectiveness is inhibited by the presence of equivalent mutants. We recently studied Trivial Compiler Equivalence (TCE) as a simple, fast and readily applicable technique for identifying equivalent mutants for C programs. In the present work, we augment our findings with further results for the Java programming language. TCE can remove a large portion of all mutants because they are determined to be either equivalent or duplicates of other mutants. In particular, TCE equivalent mutants account for 7.4 and 5.7 percent of all C and Java mutants, while duplicated mutants account for a further 21 percent of all C mutants and 5.4 percent Java mutants, on average. With respect to a benchmark ground truth suite (of known equivalent mutants), approximately 30 percent (for C) and 54 percent (for Java) are TCE equivalent. It is unsurprising that results differ between languages, since mutation characteristics are language-dependent. In the case of Java, our new results suggest that TCE may be particularly effective, finding almost half of all equivalent mutants.
Marinos Kintis, Mike Papadakis, Yue Jia 0001, Nicos Malevris, Yves Le Traon, Mark Harman
IEEE Trans. Software Eng.1
2017 Assessing and Improving the Mutation Testing Practice of PIT
abstract
Mutation testing is extensively used in software testing studies. However, popular mutation testing tools use a restrictive set of mutants which does not conform to the community standards and mutation testing literature. This can be problematic since the effectiveness of mutation strongly depends on the used mutants. To investigate this issue we form an extended set of mutants and implement it on a popular mutation testing tool named PIT. We then show that in real-world projects the original mutants of PIT are easier to kill and lead to tests that score statistically lower than those of the extended set of mutants for a range of 35% to 70% of the studied classes. These results raise serious concerns regarding the validity of mutation-based experiments that use PIT. To further show the strengths of the extended mutants we also performed an analysis using a benchmark with mutation-adequate test cases and identified equivalent mutants. Our results confirmed that the extended mutants are more effective than a) the original version of PIT and b) two other popular mutation testing tools (major and muJava). In particular, our results demonstrate that the extended mutants are more effective by 23%, 12% and 7% than the mutants of the original PIT, major and muJava. They also show that the extended mutants are at least as strong as the mutants of all the other three tools together. To support future research, we make the new version of PIT, which is equipped with the extended mutants, publicly available.
Thomas Laurent 0003, Mike Papadakis, Marinos Kintis, Christopher Henard, Yves Le Traon, Anthony Ventresque
ICST3
2016 Analysing and Comparing the Effectiveness of Mutation Testing Tools: A Manual Study
abstract
Mutation testing is considered as one of the most powerful testing methods. It operates by asking testers to design tests that reveal a set of mutants, which are purpose-made injected defects. Evidently, the strength of the method strongly depends on the used mutants. However, this dependence raises concerns regarding the mutation testing practice that is implemented by existing tools. Thus, it is probable that implementation inadequacies can lead to incompetent results. In this paper, we cross-evaluate three popular mutation testing tools for Java, namely MUJAVA, MAJOR and PIT, with respect to their effectiveness. We perform an empirical study of 3,324 manually analysed mutants from real-world projects and we find that there are large differences between the tools' effectiveness, ranging from 76% to 88%, with MUJAVA achieving the best results. We also demonstrate that no tool is able to subsume the others and provide practical recommendations on how to strengthen each one of the studied tools. Finally, our analysis shows that 11%, 12% and 7% of the mutants generated by MUJAVA, MAJOR and PIT are equivalent, respectively.
Marinos Kintis, Mike Papadakis, Andreas Papadopoulos, Evangelos Valvis, Nicos Malevris
SCAM1
2015 MEDIC: A static analysis framework for equivalent mutant identification
Marinos Kintis, Nicos Malevris
Inf. Softw. Technol.1
2015 Employing second-order mutation for isolating first-order equivalent mutants
abstract
Summary The equivalent mutant problem is a major hindrance to mutation testing. Being undecidable in general, it is only susceptible to partial solutions. In this paper, mutant classification is utilised for isolating likely to be first‐order equivalent mutants. A new classification technique,Isolating Equivalent Mutants (I‐EQM), is introduced and empirically investigated. The proposed approach employs a dynamic execution scheme that integrates theimpacton the program execution of first‐order mutants with the impact on the output of second‐order mutants. An experimental study, conducted using two independently created sets of manually classified mutants selected from real‐world programs revalidates previously published results and provides evidence for the effectiveness of the proposed technique. Overall, the study shows that I‐EQM substantially improves previous methods by retrieving a considerably higher number of killable mutants, thus, amplifying the quality of the testing process. Copyright © 2014 John Wiley & Sons, Ltd.
Marinos Kintis, Mike Papadakis, Nicos Malevris
Softw. Test. Verification Reliab.1
2013 Identifying More Equivalent Mutants via Code Similarity
abstract
Equivalent mutants are one of the major costs of mutation testing. The undecidable nature of this problem makes a fully automated solution unattainable and necessitates the manual analysis of live mutants. This paper introduces the concept of mirrored mutants, ones that affect similar code fragments. It is argued that mirrored mutants exhibit analogous behavior with respect to their equivalence. Thus, if one of them is equivalent, then the other mirrored mutants should be too. An empirical study, conducted on real world programs, investigates this argument, focusing on both intra-method and inter-method mirrored mutants. The obtained results suggest that mirrored mutants indeed exhibit this kind of behavior and thus can be utilized to ameliorate the adverse effects of the equivalent mutant problem.
Marinos Kintis, Nicos Malevris
APSEC (1)1
2012 Isolating First Order Equivalent Mutants via Second Order Mutation
abstract
In this paper, a technique named I-EQM, able to dynamically isolate first order equivalent mutants, is proposed. I-EQM works by employing a novel dynamic execution scheme that integrates both first and second order mutation. The proposed approach combines the "impact" on the program execution of the first order mutants with the "impact" on the output of second order ones, to isolate likely to be first order equivalent mutants. Experimental results on a benchmark set of manually classified mutants, selected from real word programs, reveals that I-EQM achieves to classify equivalent mutants with a 71% and 82% classification precision and recall respectively. These results improve the previously proposed approaches by selecting (retrieving) a considerably higher number of killable mutants with only a limited loss on the classification precision.
Marinos Kintis, Mike Papadakis, Nicos Malevris
ICST1
2010 Evaluating Mutation Testing Alternatives: A Collateral Experiment
abstract
Mutation testing while being a successful fault revealing technique for unit testing, it is a rather expensive one for practical use. To bridge these two aspects there is a need to establish approximation techniques able to reduce its expenses while maintaining its effectiveness. In this paper several second order mutation testing strategies are introduced, assessed and compared along with weak mutation against strong. The experimental results suggest that they both constitute viable alternatives for mutation as they establish considerable effort reductions without greatly affecting the test effectiveness. The experimental assessment of weak mutation suggests that it reduces significantly the number of the produced equivalent mutants on the one hand and that the test criterion it provides is not as weak as is thought to be on the other. Finally, an approximation of the number of first order mutants needed to be killed in order to also kill the original mutant set is presented. The findings indicate that only a small portion of a set of mutants needs to be targeted in order to be killed while the rest can be killed collaterally.
Marinos Kintis, Mike Papadakis, Nicos Malevris
APSEC1
2010 Mutation Testing Strategies - A Collateral Approach
Mike Papadakis, Nicos Malevris, Marinos Kintis
ICSOFT (2)3