Tanja E. J. Vos

dblp:v/TanjaEJVos · DBLP profile ↗
← Back
60ranked-venue papers
8as first author
24since 2021 · last 2026
0000-0002-6003-9113ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 40 · 6 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 7 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Computer networks · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Evaluating the Practical Applicability of Defect Taxonomies in Industrial Bug Repositories
Lianne V. Hufkens, Robin R. Bouwmeester, Fernando Pastor Ricós, Beatriz Marín, Tanja E. J. Vos
ENASE (1)5
2026 Teaching Testing Seriously in Academia
Tanja E. J. Vos, Bart Knaack, Beatriz Marín, Niels Doorn, Nikè van Vugt-Hage
ENASE (1)1
2026 Programming Smart Playtesting
abstract
Until recently the game industry heavily relied on manual playtesting to test the games it produces. Even if the benefits of introducing automated testing are acknowledged, it is rarely done in practice. Some of the main hurdles include the lack of automated testing tools that can target computer games as well as the complexity of automated game plays which are much more difficult to program than typical simple test sequences. This article presents an agent-based testing framework called aplib that comes with a Domain Specific Language (DSL) that allows complex playtests to be programmed more abstractly. A so-called goal structure is used to abstractly formulate a playtest scenario in terms of main goals and their decomposition into subgoals. Scenarios that are not too complicated can be formulated using static goal structures. More complex scenarios may need a test agent that can dynamically adapt its play according to the situation that evolves during the play. To handle such cases, aplib allows dynamic goals to be expressed as well. Invariants and pre-/post-conditions are used to assert the properties that a play is expected to satisfy. They include differential properties that allow constraints on the current state to be related to that of past states. Three case studies are included in the article. The first one aims to evaluate the performance of playtests programmed with aplib . The second shows that the approach can also be combined with other automated testing approaches, in this case reinforcement learning. The third shows the applicability of such playtests in a 3D setup and for non-functional testing.
I. S. W. B. Prasetya, Mehdi Dastani, Rui Prada, Tanja E. J. Vos, Frank Dignum, Fitsum Meshesha Kifetew, Guido Mintjes, Samira Shirzadehhajimahmood, Saba Gholizadeh Ansari
ACM Trans. Softw. Eng. Methodol.4
2025 Design of a Serious Game on Exploratory Software Testing to Improve Student Engagement
Niels Doorn, Tanja E. J. Vos, Beatriz Marín
ENASE2
2025 The Scent of Test Effectiveness: Can Scriptless Testing Reveal Code Smells?
abstract
This paper presents an industrial experience applying random scriptless GUI testing to the Yoho web application developed by Marviq. The study was motivated by several key challenges faced by the company, including the need to optimise testing resources, explore how random testing can complement manual testing, and investigate new coverage metrics, such as “code smell coverage”, to assess software quality and maintainability. We conducted an experiment to explore the impact of the number and length of random GUI test sequences on traditional adequacy metrics, the complementarity of random with manual testing, and the relationship between code smell coverage and traditional code coverage. Using Testar for scriptless testing and SonarQube code smell identification, results show that longer random test sequences yielded better test adequacy metrics and increased code smell coverage. In addition, random testing offers promising efficiency in test coverage and detects unique smells that m anual testing might overlook. Additionally, including code smell coverage provides valuable insights into long-term code maintainability, revealing gaps that traditional metrics may not capture. These findings highlight the benefits of combining functional testing with metrics assessing code quality, particularly in resource-constrained environments.
Olivia Rodríguez-Valdés, Domenico Amalfitano, Otto Sybrandi, Beatriz Marín, Tanja E. J. Vos
ENASE5
2025 LLM-Empowered Scriptless Functional Testing
abstract
Scriptless testing generates test sequences dynamically by automatically exploring the Graphical User Interface (GUI). Instead of relying on predefined scripts-which have proven expensive to maintain-scriptless tools detect available widgets, derive possible actions, and select actions on the fly using exploratory techniques such as random selection, model-based inference, or reinforcement learning. While scriptless testing is a valuable complement to scripted approaches, current techniques lack the intelligence needed to strategically select and execute GUI actions that fulfill specific functional testing goals-such as those derived from requirements, use cases, or user stories. Unsurprisingly, this leads companies to question the viability of scriptless testing tools and continue relying on scripts for test automation. This paper reports on the integration of Large Language Models (LLMs) into a scriptless GUI testing tool for action selection, aiming to determine whether it can generate effective action sequences to test specific functional requirements. Our results demonstrate that a multi-objective test goal structure, combined with historical and feedback context, enables LLM-empowered scriptless testing to automate functional testing. Although further research is needed to tackle challenges in complex test scenarios, our findings offer promising results that LLM-empowered scriptless testing can reduce reliance on the maintenance-heavy aspects of traditional scripted testing.
Colin Van Hooren, Fernando Pastor Ricós, Stefano Bromuri, Tanja E. J. Vos, Beatriz Marín
QRS4
2025 Behavior Driven Development for 3D games
abstract
Computer 3D games are complex software environments that require novel testing processes to ensure high-quality standards. The Intelligent Verification/Validation for Extended Reality Based Systems ( iv4XR ) framework addresses this need by enabling the implementation of autonomous agents to automate game testing scenarios. This framework facilitates the automation of regression test cases for complex 3D games like Space Engineers. Nevertheless, the technical expertise required to define test scripts using iv4XR can constrain seamless collaboration between developers and testers. This paper reports how integrating a Behavior-Driven Development (BDD) approach with the iv4XR framework allows the industrial company behind Space Engineers to automate regression testing. The success of this industrial collaboration has inspired the iv4XR team to integrate the BDD approach to improve the automation of play-testing for the experimental 3D game LabRecruits. Furthermore, the iv4XR framework has been extended with tactical programming to enable the automation of long-play test scenarios in Space Engineers. These results underscore the versatility of the iv4XR framework in supporting diverse testing approaches while showcasing how BDD empowers users to create, manage, and execute automated game tests using comprehensive and human-readable statements.
Fernando Pastor Ricós, Beatriz Marín, I. S. W. B. Prasetya, Tanja E. J. Vos, Joseph Davidson, Karel Hovorka
Data Knowl. Eng.4
2024 Grammar-Based Action Selection Rules for Scriptless Testing
abstract
Scriptless testing at the GUI level involves generating test sequences on the fly. These test sequences mimic user interactions on the GUI. The creation of these sequences works through action selection rules, which is most commonly based on stochastic methods. Script-less tests are reliable because they work with the actual state of the System Under Test (SUT). However, the tests are less specific, harder to interpret, and it is difficult to test concrete use cases or workflows. We want to tackle this drawback of scriptless tests by introducing action selection rules that are easier to guide than pure stochastic methods. In this paper, a new approach based on a grammar for the action selection rules is proposed, enabling scriptless testing tools to mimic user behaviour when interacting with web applications. While grammars have been used in software testing to generate input data for test cases, the proposed approach uses grammars to specify action selection rules to generate test sequences that mimic testing strategies employed by human testers. An empirical study has been performed to evaluate the effectiveness and the efficiency of the grammar-based action selection rules to filing web forms in comparison with random action selection rules. In the study, two SUTs were used: WebformSUT and Parabank. The average success rate for the grammar-based approach was 95.9% against random's 57.0% for WebformSUT and 99.8% against 55.7% for Parabank. For the widget interaction grammar-based had an average deviation from the ideal ratio of 0.06165 (WebformSUT) and 0.0180 (Parabank), compared random's 0.4318 (WebformSUT) and 0.7774 (Parabank). The results demonstrate the effectiveness of the grammar-based approach and the improvement in the use of resources.
Lianne V. Hufkens, Fernando Pastor Ricós, Beatriz Marín, Tanja E. J. Vos
AST4
2024 Towards Understanding Students' Sensemaking of Test Case Design: A One-Page Summary
abstract
This study examines sensemaking in student test case design, showing a reliance on conceptual knowledge learned during programming courses over exploratory testing strategies. Of the three identified approaches taken by students, the “developer approach” is used most often, suggesting a gap in software engineering education. We hypothesise that software testing should be taught in computer science programs using a design paradigm based on empiricism instead of rationalism. Based on these results, and our hypothesis, we will further analyse the sensemaking processes of both students and experts, and create an instructional design to improve software testing education in computer science programs.
Niels Doorn, Tanja E. J. Vos, Beatriz Marín
CSEE&T2
2024 Scriptless Testing for an Industrial 3D Sandbox Game
abstract
Computer games have reached unprecedented importance, exceeding two billion users in the early 2020s. Human game testers bring invaluable expertise to evaluate complex games like 3D sandbox games. However, the sheer scale and diversity of game content constrain their ability to explore all scenarios manually. Recognizing the significance and inherent complexity of game testing, our research aims to investigate new automated testing approaches. To achieve this goal, we have integrated scriptless testing into the industrial game Space Engineers, enabling an automated approach to explore and test sandbox game scenarios. Our approach involves the development of a Space Engineers-plugin, leveraging the Intelligent Verification and Validation for Extended Reality-Based Systems (IV4XR) framework and extending the capabilities of the open-source scriptless testing tool TESTAR. Through this research, we unveil the potential of a scriptless agent to explore 3D sandbox game scenarios autonomously. Results demonstrate the effectiveness of an autonomous scriptless agent in achieving spatial coverage when exploring and (dis)covering elements within the 3D sandbox game.
Fernando Pastor Ricós, Beatriz Marín, Tanja E. J. Vos, Joseph Davidson, Karel Hovorka
ENASE3
2024 State of the Practice in Software Testing Teaching in Four European Countries
abstract
Software testing is an indispensable component of software development, yet it often receives insufficient attention. The lack of a robust testing culture within computer science and informatics curricula contributes to a shortage of testing expertise in the software industry. Addressing this problem at its root -education- is paramount. In this paper, we conduct a comprehensive mapping review of software testing courses, elucidating their core attributes and shedding light on prevalent subjects and instructional methodologies. We mapped 117 courses offered by Computer Science (and related) degrees in 49 academic institutions from four Western European countries, namely Belgium, Italy, Portugal and Spain. The testing subjects were mapped against the conceptual framework provided by the ISO/IEC/IEEE 29119 standard on software testing. Among the results, the study showed that dedicated software testing courses are offered by only 39% of the analysed universities, whereas the basics of software testing are taught in at least one course at every university. The analysis of the software testing topics highlights the gaps that need to be filled in order to better align the current academic offerings with the real industry needs.
Porfirio Tramontana, Beatriz Marín, Ana C. R. Paiva, Alexandra Mendes, Tanja E. J. Vos, Domenico Amalfitano, Felix Cammaerts, Monique Snoeck, Anna Rita Fasolino
ICST5
2024 Novelty-Driven Evolutionary Scriptless Testing
Lianne V. Hufkens, Tanja E. J. Vos, Beatriz Marín
RCIS (2)2
2024 An Industrial Experience Leveraging the iv4XR Framework for BDD Testing of a 3D Sandbox Game
Fernando Pastor Ricós, Beatriz Marín, I. S. W. B. Prasetya, Tanja E. J. Vos, Joseph Davidson, Karel Hovorka
RCIS (1)4
2024 Scriptless and Seamless: Leveraging Probabilistic Models for Enhanced GUI Testing in Native Android Applications
Olivia Rodríguez-Valdés, Kevin van der Vlist, Robbert van Dalen, Beatriz Marín, Tanja E. J. Vos
RCIS (2)5
2023 Set the right example when teaching programming: Test Informed Learning with Examples (TILE)
abstract
Many educators face problems with integrating testing into programming education. For instance: existing courses are already fully packed; testing requires skills that students might not yet have; and testing is, although considered important, not always given priority by students. Educators, in general, do not have time to overhaul a programming course to fully integrate testing, resulting in a situation in which the improvement of testing education seems to have slowed down. In this paper, we propose Test Informed Learning with Examples (TILE), a new concept to create test-awareness in introductory programming courses. TILE aims to introduce testing as early as possible and in a subtle way. As a result, integration into existing curricula can be done seamlessly and requires less effort than completely overhauling existing programming courses. The contributions of this paper are: the presentation of TILE; experiences of having applied this method in the classroom; and an open repository with assignments using our approach. Applying TILE seems to be a promising approach to introduce testing in early programming. Moreover, some TILEs can be added to existing courses with almost no effort from day one. More research is needed to gain confidence in the benefits of using TILE over time and to collect evidence that we reached the final aim of TILE, i.e. students that test because that inherently belongs to programming, and not because it is explicitly asked from them.
Niels Doorn, Tanja E. J. Vos, Beatriz Marín, Erik Barendsen
ICST2
2023 Domain TILEs: Test Informed Learning with Examples from the Testing Domain
Niels Doorn, Tanja E. J. Vos, Beatriz Marín, Christoph Bockisch, Steffen Dick, Erik Barendsen
RCIS2
2023 Using GUI Change Detection for Delta Testing
Fernando Pastor Ricós, Rick Neeft, Beatriz Marín, Tanja E. J. Vos, Pekka Aho
RCIS4
2023 Reinforcement Learning for Scriptless Testing: An Empirical Investigation of Reward Functions
Olivia Rodríguez-Valdés, Tanja E. J. Vos, Beatriz Marín, Pekka Aho
RCIS2
2023 Towards understanding students' sensemaking of test case design
abstract
Software testing is the most used technique for quality assurance in industry. However, in computer science education software testing is still treated as a second-class citizen and students are unable to test their software well enough. One reason for this is that teaching the subject of software testing is difficult as it is a complex intellectual activity for which students need to allocate multiple cognitive resources at the same time. A myriad of primary and secondary studies have tried to solve this problem in education, however still with very limited results. Before we can design interventions to improve our pedagogical approaches, we need to gain more in-depth understanding and recognition of sensemaking as it is happening when students design test cases. An initial exploratory study identified four different sensemaking approaches used by students while creating test models. In this paper we present a follow-up study with 50 students from a large university in Spain. The used methodology was based on the previous study with the improvements that originated from its evaluation. We asked the participants to create a test model based on a description of a test problem using a specialized web-based tool for modeling test cases. We measured how well these models fit the test problem, the sensemaking process that students went through when creating the models, and the students’ perception of the modeling task. The participants received no compensation for their efforts, and we scheduled the experiment during a regular class. Apart from the created models and their metadata, we also collected recordings of the students’ computer screens made during the experiment and used a questionnaire to study their perspectives on the assignment. All the collected textual, graphical, and video data was analyzed using an iterative inductive analysis process to allow new information about the different sensemaking approaches to emerge. We gained better insights into the sensemaking processes of students while modeling test cases for a problem. The results enabled us to refine our previous findings, and we identified new sensemaking approaches. Based on these results, we can further investigate ways to influence the sensemaking process in education, the possible misconceptions that have a negative influence on it, and the desired mental model we want our students to have to design test cases.
Niels Doorn, Tanja E. J. Vos, Beatriz Marín
Data Knowl. Eng.2
2023 Scripted and scriptless GUI testing for web applications: An industrial case
abstract
Automation is required in the software development to reduce the high costs of producing software and to address the short release cycles of modern development processes. Lot of effort has been performed to automate testing, which is one of the most resource-consuming development phases. Automation of testing through the Graphical User Interface (GUI) has been researched to improve the system testing. We aim to evaluate the complementarity of automated GUI testing tools in a real industrial context, which refers to the capability of the tools to work usefully together. To address the objective, we conduct an exploratory case study in an IT development company from The Netherlands. We select two representative tools for automated GUI testing, one for scripted and another for scriptless testing. We measure the complementarity by measuring the effectiveness, the efficiency, and subjective satisfaction of the tools. It can be observed that the scripted tool performs better in detecting process failures, and the scriptless tool performs better in detecting visible failures and also reaching higher coverage. Both tools perform in a similar way in terms of efficiency. Additionally, both tools were perceived to be useful in the survey performed for the subjective satisfaction. We conclude that scriptless and scripted testing approaches are complementary, and they can improve the effectiveness compared to manual testing processes performed in an industrial context by detecting different failures and reducing the effort and time to find these failures and to reproduce them.
Axel Bons, Beatriz Marín, Pekka Aho, Tanja E. J. Vos
Inf. Softw. Technol.4
2023 Distributed state model inference for scriptless GUI testing
abstract
State model inference of software applications through the Graphical User Interface (GUI) is a technique that identifies GUI states and transitions, and maps them into a model. Scriptless GUI testing tools can benefit substantially from the availability of these state models, for example, to improve the exploration, or have sophisticated test oracles. However, inferring models for large systems requires a long execution time. Our goal is to improve the speed of the state model inference process. To achieve this goal, this paper presents a distributed state model inference approach with an open source scriptless GUI testing tool. Moreover, in order to be able to infer a suitable model, we design a set of strategies to deal with abstraction challenges and to distinguish GUI states and transitions in the model. To validate it, we conduct an experiment with two open source web applications that have been tested with the distributed architecture using one to six Docker containers sharing the same state model. With the obtained results, we can conclude that it is feasible to infer a model with a distributed approach and that using the distributed approach reduces the time required for inferring a state model.
Fernando Pastor Ricós, Arend Slomp, Beatriz Marín, Pekka Aho, Tanja E. J. Vos
J. Syst. Softw.5
2022 Scriptless GUI Testing on Mobile Applications
abstract
Traditionally, end-to-end testing of mobile apps is either performed manually or automated with test scripts. However, manual GUI testing is expensive and slow, and test scripts are fragile for GUI changes, resulting in high maintenance costs. Scriptless testing attempts to address the costs associated with GUI testing. Existing scriptless approaches for mobile testing do not seem to fit the requirements of the industry, specifically those of the ING. This study presents an extension to open source TESTAR tool to support scriptless GUI testing of Android and iOS applications. We present an initial validation of the tool on an industrial setting at the ING. From the validation, we determine that the extended TESTAR outperforms two other state-of-the-art scriptless testing tools for Android in terms of code coverage, and achieves similar performance as the scripted test automation already in use at the ING. Moreover, we see that the scriptless approach covers parts of the application under test that the existing test scripts did not cover, showing the complementarity of the approaches, providing more value for the testers.
Thorn Jansen, Fernando Pastor Ricós, Yaping Luo, Kevin van der Vlist, Robbert van Dalen, Pekka Aho, Tanja E. J. Vos
QRS7
2022 State Model Inference Through the GUI Using Run-Time Test Generation
Ad Mulders, Olivia Rodríguez-Valdés, Fernando Pastor Ricós, Pekka Aho, Beatriz Marín, Tanja E. J. Vos
RCIS6
2021 testar - scriptless testing through graphical user interface
abstract
Summary Covering all the possible paths of the graphical user interface (GUI) with test scripts would take too much effort and result in serious maintenance issues. We propose complementing scripted testing with scriptless test automation using the open‐source testar tool. This paper gives a comprehensive overview of testar and its latest extensions together with the ongoing and future research. With this paper, we hope we can help and encourage other researchers to use testar for their GUI testing‐related research and pave the way for an international research agenda in GUI testing built upon stable and open‐source infrastructure.
Tanja E. J. Vos, Pekka Aho, Fernando Pastor Ricós, Olivia Rodríguez-Valdés, Ad Mulders
Softw. Test. Verification Reliab.1
2020 Tutorial on a Gamification Toolset for Improving Engagement of Students in Software Engineering Courses
abstract
Few if any would dispute that educating software engineering is a challenging endeavour. Although programming and creating new artefacts can motivate the creativity of students. Other software engineering topics (like e.g. requirement specifications and testing) are not considered very exciting by students. However, these topics are important to develop quality software and insufficient knowledge of students - Europe's future software engineers - in the long run contributes to failing software. The EU Erasmus+ project IMPRESS was set to explore the use of gamification in educating software engineering at the university level. The objective has been to develop a toolset that can help to improve students' engagement, and hence their appreciation, for the taught subjects like software testing and specifications. The proposed tutorial will guide participants through the set of tools developed by the project and introduce how they can use them to improve students' engagement.
Tanja E. J. Vos, Gordon Fraser 0001, Iván Martínez-Ortiz, Rui Prada, António Manuel Ferreira Rito da Silva, I. S. W. B. Prasetya
CSEE&T1
2020 Agent-based Testing of Extended Reality Systems
abstract
Testing for quality assurance (QA) is a crucial step in the development of Extended Reality (XR) systems that typically follow iterative design and development cycles. Bringing automation to these testing procedures will increase the productivity of XR developers. However, given the complexity of the XR environments and the User Experience (UX) demands, achieving this is highly challenging. We propose to address this issue through the creation of autonomous cognitive test agents that will have the ability to cope with the complexity of the interaction space by intelligently explore the most prominent interactions given a test goal and support the assessment of affective properties of the UX by playing the role of users.
Rui Prada, I. S. W. B. Prasetya, Fitsum Meshesha Kifetew, Frank Dignum, Tanja E. J. Vos, Jason Lander, Jean-Yves Donnart, Alexandre Kazmierowski, Joseph Davidson, Pedro M. Fernandes
ICST5
2020 Deploying TESTAR to Enable Remote Testing in an Industrial CI Pipeline: A Case-Based Evaluation
Fernando Pastor Ricós, Pekka Aho, Tanja E. J. Vos, Ismael Torres Boigues, Ernesto Calás Blasco, Héctor Martínez Martínez
ISoLA (1)3
2020 Scriptless Testing at the GUI Level in an Industrial Setting
Hatim Chahim, Mehmet Duran, Tanja E. J. Vos, Pekka Aho, Nelly Condori-Fernández
RCIS3
2019 IMPRESS: Improving Engagement in Software Engineering Courses Through Gamification
Tanja E. J. Vos, I. S. W. B. Prasetya, Gordon Fraser 0001, Iván Martínez-Ortiz, Iván J. Pérez-Colado, Rui Prada, José Bernardo Rocha, António Manuel Ferreira Rito da Silva
PROFES1
2019 Offline Oracles for Accessibility Evaluation with the TESTAR Tool
abstract
To manage the complexity of today's information systems, we need to investigate novel approaches for automated testing. In this paper we present two extensions to TESTAR, a state-of-the-art tool for testing systems through the GUI. We extend this tool with 1) a systematic and powerful approach for storing and querying test results through a graph database for offline oracles, and 2) support for accessibility evaluation of general applications for stakeholders with disabilities utilizing offline oracles. Furthermore, we conduct a preliminary validation of these extensions through a case study on the popular VLC Media Player.
Floren de Gier, Davy Kager, Stijn de Gouw, Tanja E. J. Vos
RCIS4
2018 Towards Automated Testing of the Internet of Things: Results Obtained with the TESTAR Tool
Mirella Martínez, Anna Esparcia-Alcázar, Tanja E. J. Vos, Pekka Aho, Joan Fons
ISoLA (3)3
2017 Effectiveness Assessment of an Early Testing Technique using Model-Level Mutants
abstract
While modern software development technologies enhance the capabilities of model-based/driven development, they introduce challenges for testers such as how to perform early testing at model level to ensure the quality of the model. In this context, we have developed an early testing technique supported by the CoSTest tool to validate requirements at model level. In this paper we describe an empirical evaluation of CoSTest with respect to its effectiveness in terms of its fault detection and test suite adequacy. This evaluation is carried out by model-level mutation testing using first order mutants (created by injection of a single fault) and high order mutants (containing more than one fault) with seven conceptual schemas (of different sizes) that represent the functionality of different software systems in different domains. Our findings show that the tests generated by CoSTest are effective at killing a large number of mutants. However, there are also some fault types (e.g. delete the references to a class attribute or an operation call in a constraint) that our test suites were not able to detect. CoSTest was more effective in terms of detecting fault types using high order mutants that first order mutants. Thus, CoSTest's effectiveness is affected by the mutant type tested.
María Fernanda Granda, Nelly Condori-Fernández, Tanja E. J. Vos, Oscar Pastor 0001
EASE3
2017 Evolving Rules for Action Selection in Automated Testing via Genetic Programming - A First Approach
Anna Esparcia-Alcázar, Francisco Almenar, Urko Rueda, Tanja E. J. Vos
EvoApplications (2)4
2017 Overview of the ICST International Software Testing Contest
abstract
In the software testing contest, practitioners and researcher's are invited to test their test approaches against similar approaches to evaluate pros and cons and which is perceivably the best. The 2017 iteration of the contest focused on Graphical User Interface-driven testing, which was evaluated on the testing tool TESTONA. The winner of the competition was announced at the closing ceremony of the international conference on software testing (ICST), 2017.
Emil Alégroth, Shinsuke Matsuki, Tanja E. J. Vos, Kinji Akemine
ICST3
2017 CoSTest: A Tool for Validation of Requirements at Model Level
abstract
We present CoSTest, a tool that supports the validation of Conceptual Schemas by using testing. The tool implements techniques for transforming instantiations from a Requirements Model into test case implementations by supporting a Model-driven architecture.
María Fernanda Granda, Nelly Condori-Fernández, Tanja E. J. Vos, Oscar Pastor 0001
RE3
2016 Mutation Operators for UML Class Diagrams
María Fernanda Granda, Nelly Condori-Fernández, Tanja E. J. Vos, Oscar Pastor 0001
CAiSE3
2016 Automated Localisation Testing in Industry with Test ^* ∗
Mireilla Martínez, Anna Esparcia-Alcázar, Urko Rueda, Tanja E. J. Vos, Carlos Ortega
ICTSS4
2015 What do we know about the defect types detected in conceptual models?
abstract
In Model-Driven Development (MDD), defects are managed at the level of conceptual models because the other artefacts are generated from them, such as more refined models, test cases and code. Although some studies have reported on defect types at model level, there still does not exist a clear and complete overview of the defect types that occur at the abstraction level. This paper presents a systematic mapping study to identify the model defect types reported in the literature and determine how they have been detected. Among the 282 articles published in software engineering area, 28 articles were selected for analysis. A total of 226 defects were identified, classified and their results analysed. For this, an appropriate defect classification scheme was built based on appropriate dimensions for models in an MDD context.
María Fernanda Granda, Nelly Condori-Fernández, Tanja E. J. Vos, Oscar Pastor 0001
RCIS3
2014 Evaluating the TESTAR tool in an industrial case study
abstract
[Context] Automated test case design and execution at the GUI level of applications is not a fact in industrial practice. Tests are still mainly designed and executed manually. In previous work we have described TESTAR, a tool which allows to set-up fully automatic testing at the GUI level of applications to find severe faults such as crashes or non-responsiveness. [Method] This paper aims at the evaluation of TESTAR with an industrial case study. The case study was conducted at SOFTEAM, a French software company, while testing their Modelio SaaS system, a cloud-based system to manage virtual machines that run their popular graphical UML editor Modelio. [Goal] The goal of the study was to evaluate how the tool would perform within the context of SOFTEAM and on their software application. On the other hand, we were interested to see how easy or difficult it is to learn and implant our academic prototype within an industrial setting. [Results] The effectiveness and efficiency of the automated tests generated with TESTAR can definitely compete with that of the manual test suite. [Conclusions] The training materials as well as the user and installation manual of TESTAR need to be improved using the feedback received during the study. Finally, the need to program Java-code to create sophisticated oracles for testing created some initial problems and some resistance. However, it became clear that this could be solved by explaining the need for these oracles and compare them to the alternative of more expensive and complex human oracles. The need to raise consciousness that automated testing means programming solved most of the initial problems.
Sebastian Bauersfeld, Tanja E. J. Vos, Nelly Condori-Fernández, Alessandra Bagnato, Etienne Brosse
ESEM2
2014 Evaluating rogue user testing in industry: An experience report
abstract
Testing applications with a graphical user interface (GUI) is an important, though challenging and time consuming task. The state of the art in the industry are still capture and replay tools, which may simplify the recording and execution of input sequences, but do not support the tester in finding fault-sensitive test cases and leads to a huge overhead on maintenance of the test cases when the GUI changes. In earlier works we presented the Rogue User Testing Tool, an automated approach to testing applications at the GUI level whose objective is to solve part of the maintenance problem by automatically generating test cases based on a structure that is automatically derived from the GUI. In this paper we report on our experiences obtained when implanting the Rogue User testing Tool with the Spanish software vendor Clavei who decided to apply the tool to stress test a component of one of their ERP applications. Our main goal was to identify potential problems that arise during the setup of the Rogue User. While carrying out our tests, we discovered critical and previously unknown faults in the application under test.
Sebastian Bauersfeld, Antonio de Rojas, Tanja E. J. Vos
RCIS3
2013 Combinatorial Testing Tool Learnability in an Industrial Environment
abstract
[Context] Numerous combinatorial testing techniques are available for generating test cases. However, many of them are never used in practice. [Objective] Considering that learn ability plays a vital role in initial adoption or rejection of a technology, in this paper we aim to investigate the learnability of a combinatorial testing tool in an industrial environment. [Method] A case study research method was designed and conducted, by including i) the definition of learnability measures for test cases models built using a combinatorial testing tool. ii) A training program was also implemented. iii) Qualitative and quantitative evaluation based on a three-level strategy was carried out (Reaction, Learning, and Performance). [Results] At the first level, the tool was perceived as easy to learn by the trainees (from a five-point ordinal scale). However, at the second level, during hands-on learning, it changed slightly: According to the working diaries, there were major difficulties. At third level, analyzing the learning curve of each trainee, we observe that semantic errors made per each subject were reduced slightly over the time.
Peter M. Kruse, Nelly Condori-Fernández, Tanja E. J. Vos, Alessandra Bagnato, Etienne Brosse
ESEM3
2013 Evaluating the FITTEST Automated Testing Tools: An Industrial Case Study
abstract
This paper aims at evaluating a set of automated tools of the FITTEST EU project within an industrial case study. The case study was conducted at the IBM Research lab in Haifa, by a team responsible for building the testing environment for future development versions of an IBM system management product. The main function of that product is resource management in a networked environment. This case study has investigated whether current IBM Research testing practices could be improved or complemented by using some of the automated testing tools that were developed within the FITTEST EU project. Although the existing Test Suite from IBM Research (TSibm) that was selected for comparison is substantially smaller than the Test Suite generated by FITTEST (TSfittest), the effectiveness of TSfittest, measured by the injected faults coverage is significantly higher (50% vs 70%). With respect to efficiency, by normalizing the execution times, we found the TSfittest runs faster (9.18 vs. 6.99). This is due to the fact that the TSfittest includes shorter tests. Within IBM Research and for the testing of the target product in the simulated environment: the FITTEST tools can increase the effectiveness of the current practice and the test cases automatically generated by the FITTEST tools can help in more efficient identification of the source of the identified faults. Moreover, the FITTEST tools have shown the ability to automate testing within a real industry case.
Duy Cu Nguyen, Bilha Mendelson, Daniel Citron, Onn Shehory, Tanja E. J. Vos, Nelly Condori-Fernández
ESEM5
2013 An empirical approach for evaluating the usability of model-driven tools
Nelly Condori-Fernández, José Ignacio Panach, Arthur I. Baars, Tanja E. J. Vos, Oscar Pastor 0001
Sci. Comput. Program.4
2013 Evolutionary functional black-box testing in an industrial setting
Tanja E. J. Vos, Felix F. Lindlar, Benjamin Wilmes, Andreas Windisch, Arthur I. Baars, Peter M. Kruse, Hamilton Gross, Joachim Wegener
Softw. Qual. J.1
2013 Using a functional size measurement procedure to evaluate the quality of models in MDD environments
abstract
Models are key artifacts in Model-Driven Development (MDD) methods. To produce high-quality software by using MDD methods, quality assurance of models is of paramount importance. To evaluate the quality of models, defect detection is considered a suitable approach and is usually applied using reading techniques. However, these reading techniques have limitations and constraints, and new techniques are required to improve the efficiency at finding as many defects as possible. This article presents a case study that has been carried out to evaluate the use of a Functional Size Measurement (FSM) procedure in the detection of defects in models of an MDD environment. To do this, we compare the defects and the defect types found by an inspection group with the defects and the defect types found by the FSM procedure. The results indicate that the FSM is useful since it finds all the defects related to a specific defect type, it finds different defect types than an inspection group, and it finds defects related to the correctness and the consistency of the models.
Beatriz Marín, Giovanni Giachetti, Oscar Pastor 0001, Tanja E. J. Vos, Alain Abran
ACM Trans. Softw. Eng. Methodol.4
2012 GUITest: a Java library for fully automated GUI robustness testing
abstract
Graphical User Interfaces (GUIs) are substantial parts of today's applications, no matter whether these run on tablets, smartphones or desktop platforms. Since the GUI is often the only component that humans interact with, it demands for thorough testing to ensure an efficient and satisfactory user experience. Being the glue between almost all of an application's components, GUIs also lend themselves for system level testing. However, GUI testing is inherently difficult and often involves great manual labor, even with modern tools which promise automation. This paper introduces a Java library called GUITest, which allows to generate fully automated GUI robustness tests for complex applications, without the need to manually generate models or input sequences. We will explain how it operates and present first results on its applicability and effectivity during a test involving Microsoft Word.
Sebastian Bauersfeld, Tanja E. J. Vos
ASE2
2012 Industrial Case Studies for Evaluating Search Based Structural Testing
abstract
Evolutionary structural testing has been researched and promising results have been presented. However, it has hardly been applied to real-world complex systems and as such, little is known about the scalability, applicability and acceptability of it in an industrial setting. The European project EvoTest (IST-33472) team has been working from 2006 till 2009 to improve this situation and this paper informs about the results. We start with an overview of tools and techniques which we have developed for automated evolutionary structural testing. Subsequently, we describe the empirical setup used to study the applicability of evolutionary structural testing in industry through two case studies. The test objects used for the studies are selected functions (handwritten and generated) from production systems at Daimler and Berner & Mattner Systemtechnik (BMS) like, for example, Rear Window Defroster, Global Powertrain Engine Controller, Window Lift Control System, etc. The results of the case studies are described and research questions are assessed based on the obtained results. In summary, the results indicate that evolutionary structural testing in an industrial setting is worthwhile and profitable. Hardly any detailed knowledge of evolutionary computation is required to search for interesting test data. The case studies also research the benefits of using techniques like automated parameter tuning and search space smoothing.
Tanja E. J. Vos, Arthur I. Baars, Felix F. Lindlar, Andreas Windisch, Benjamin Wilmes, Hamilton Gross, Peter M. Kruse, Joachim Wegener
Int. J. Softw. Eng. Knowl. Eng.1
2011 Search-Based Testing, the Underlying Engine of Future Internet Testing
Arthur I. Baars, Kiran Lakhotia, Tanja E. J. Vos, Joachim Wegener
FedCSIS3
2011 Testing and Remote Maintenance of Real Future Internet Scenarios, Towards FITTEST and FastFix Advanced Software Engineering
Alessandra Bagnato, Anna Esparcia-Alcázar, Tanja E. J. Vos, Beatriz Marín, José Oliver Murillo, Salvador I. Folgado, Auxiliadora Carlos Alberola
FedCSIS3
2011 Towards an Experimental Framework for Measuring Usability of Model-Driven Tools
José Ignacio Panach, Nelly Condori-Fernández, Arthur I. Baars, Tanja E. J. Vos, Ignacio Romeu, Oscar Pastor 0001
INTERACT (4)4
2011 Symbolic search-based testing
abstract
We present an algorithm for constructing fitness functions that improve the efficiency of search-based testing when trying to generate branch adequate test data. The algorithm combines symbolic information with dynamic analysis and has two key advantages: It does not require any change in the underlying test data generation technique and it avoids many problems traditionally associated with symbolic execution, in particular the presence of loops. We have evaluated the algorithm on industrial closed source and open source systems using both local and global search-based testing techniques, demonstrating that both are statistically significantly more efficient using our approach. The test for significance was done using a one-sided, paired Wilcoxon signed rank test. On average, the local search requires 23.41% and the global search 7.78% fewer fitness evaluations when using a symbolic execution based fitness function generated by the algorithm.
Arthur I. Baars, Mark Harman, Youssef Hassoun, Kiran Lakhotia, Phil McMinn, Paolo Tonella, Tanja E. J. Vos
ASE7
2011 Towards testing future Web applications
abstract
The current Web applications are in continuous evolution to provide new and more complex functionalities, which can improve the user experience by means of adaptivity and dynamic changes. Since testing is the most frequently used technique to evaluate the quality of software applications in industry, novel testing approaches will be necessary to evaluate the quality of future (and more complex) web applications. In this paper, we investigate the testing challenges of future web applications and propose a testing methodology that addresses these challenges by the integration of search-based testing, model-based testing, oracle learning, concurrency testing, combinatorial testing, regression testing, and coverage analysis. This paper also presents a testing metamodel that states testing concepts and their relationships, which are used as the theoretical basis of the proposed testing methodology.
Beatriz Marín, Tanja E. J. Vos, Giovanni Giachetti, Arthur I. Baars, Paolo Tonella
RCIS2
2011 Early Usability Measurement in Model-Driven Development: Definition and Empirical Evaluation
abstract
Usability is currently a key feature for developing quality systems. A system that satisfies all the functional requirements can be strongly rejected by end-users if it presents usability problems. End-users demand intuitive interfaces and an easy interaction in order to simplify their work. The first step in developing usable systems is to determine whether a system is or is not usable. To do this, there are several proposals for measuring the system usability. Most of these proposals are focused on the final system and require a large amount of resources to perform the evaluation (end-users, video cameras, questionnaires, etc.). Usability problems that are detected once the system has been developed involve a lot of reworking by the analyst since these changes can affect the analysis, design, and implementation phases. This paper proposes a method to minimize the resources needed for the evaluation and reworking of usability problems. We propose an early usability evaluation that is based on conceptual models. The analyst can measure the usability of attributes that depend on conceptual primitives. This evaluation can be automated taking as input the conceptual models that represent the system abstractly.
José Ignacio Panach, Nelly Condori-Fernández, Tanja E. J. Vos, Nathalie Aquino, Francisco Valverde
Int. J. Softw. Eng. Knowl. Eng.3
2010 Evaluating the usefulness of a functional size measurement procedure to detect defects in MDD models
abstract
Models are key artifacts in Model-Driven Development (MDD) methods. To evaluate the quality of models, defect detection is considered to be a suitable approach, which is usually applied using reading techniques. However, new techniques are required in order to find as many defects as possible. This paper presents a case study to evaluate the usefulness of a Functional Size Measurement (FSM) procedure to detect defects in models of a MDD environment. The results indicate that the FSM is useful in finding all the defects that are related to a defect type as well as finding different defect types than an inspection team does.
Beatriz Marín, Giovanni Giachetti, Oscar Pastor 0001, Tanja E. J. Vos, Alain Abran
ESEM4
2010 Industrial Scaled Automated Structural Testing with the Evolutionary Testing Tool
abstract
Evolutionary testing has been researched and promising results have been presented. However, evolutionary testing has remained predominately a research-based activity not practiced within industry. Although attempts have been made, such as Daimler's Evolutionary Structural Test (EST) prototype, until now, no such tool has been suitable for industrial adoption. The European project EvoTest (IST-33472) team has been working from 2006 till 2009 to improve this situation. This paper describes the final version of the Evolutionary Testing Framework (ETF) resulting from the EvoTest project. In specific we will present the EvoTest Structural Testing tool for fully automatic structural testing that has been demonstrated to be suitable within an industrial setting. The paper concentrates on how to use it and interpret the results. The paper starts with introducing the concepts of Evolutionary Testing in general and Structural Testing in specific. Subsequently, the ETF and the EvoTest Structural Testing tool built on-top of it will be described. We will concentrate on the usage, the architecture, and remaining limitations of the tool. The paper concludes describing the results of using the EvoTest Structural Testing tool in practice on real-world systems in an industrial setting.
Tanja E. J. Vos, Arthur I. Baars, Felix F. Lindlar, Peter M. Kruse, Andreas Windisch, Joachim Wegener
ICST1
2008 Trace-based Reflexive Testing of OO Programs with T2
abstract
This paper presents an automatic trace-based unit testing approach to test object-oriented programs. Most automated testing tools test a class C by testing each of its methods in isolation. Such an approach works poorly if specifications are only partial, which is usually the case in practice. In contrast, our approach generates sequences of calls to the methods of C that are checked on-the-fly. This is more interactive, and has the side effect that methods are checking each other. Although simple, it seems to work quite well, even when specifications are only partially provided. We implement the approach in a tool called T2. It targets Java. It can test internal errors, Hoare triple specifications, class invariant, and even temporal properties. Furthermore, T2 accepts ’in-code’ specifications, these are specifications written in the specified class itself, and are written in plain Java; hence reducing the cost usually needed to maintain specifications to minimum.
I. S. W. B. Prasetya, Tanja E. J. Vos, Arthur I. Baars
ICST2
2006 Web Cube
I. S. W. B. Prasetya, Tanja E. J. Vos, S. Doaitse Swierstra
FORTE2
2005 Building Verification Condition Generators by Compositional Extensions
abstract
This paper describes a technique that combines algebraic datatypes and monads to build derivative verification condition generators (VCGs) by extending a base VCG. Extensions are compositional and can be stacked while the base VCG is left unchanged. The technique can be used to build a set of weaker VCGs to do light weight verification. Moreover, it enables us to add an ability to generate validation traces. The paper explains the technique through an example that extends a simple language L/sub 0/ with new constructs to handle exceptions. To deal with exceptions, not only the logic of L/sub 0/ has to be extended with new rules, its structure also needs to be changed. We show that using our technique the extension can be implemented in a simple and compositional way, without any change to the underlying logic.
I. S. W. B. Prasetya, A. Azurat, Tanja E. J. Vos, Arthur van Leeuwen
SEFM3
2004 A UNITY-Based Framework Towards Component Based Systems
I. S. W. B. Prasetya, Tanja E. J. Vos, A. Azurat, S. Doaitse Swierstra
OPODIS2
1997 Make your Enemies Transparent
Tanja E. J. Vos, S. Doaitse Swierstra
WG1