VLDB 2026 Research / reviewers in the wild / expert
José Miguel Rojas
dblp:94/5122 · also José Miguel Rojas Siles
· DBLP profile ↗
27ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0002-0079-5355ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 24 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Theory of computation · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated testing of prevalent 3D user interactions in virtual reality applicationsabstractVirtual Reality (VR) technologies offer immersive user experiences across various domains, but present unique testing challenges compared to traditional software. Existing VR testing approaches enable scene navigation and interaction activation, but lack the ability to automatically synthesise realistic 3D user inputs (e.g, grab and trigger actions via hand-held controllers). Automated testing that generates and executes such input remains an unresolved challenge. Furthermore, existing metrics fail to robustly capture diverse interaction coverage. This paper addresses these gaps through four key contributions. First, we empirically identify four prevalent interaction types in nine open-source VR projects: fire , manipulate , socket , and custom . Second, we introduce the Interaction Flow Graph , a novel abstraction that systematically models 3D user interactions by identifying targets, actions, and conditions. Third, we construct XRBench3D , a benchmark comprising ten VR scenes that encompass 456 distinct user interactions for evaluating VR interaction testing. Finally, we present XRintTest , an automated testing approach that leverages this graph for dynamic scene exploration and interaction execution. Evaluation on XRBench3D shows that XRintTest achieves great effectiveness, reaching 93% coverage of fire , manipulate and socket interactions across all scenes, and performing 12x more effectively and 6x more efficiently than random exploration. Moreover, XRintTest can detect runtime exceptions and non-exception interaction issues, including subtle configuration defects. In addition, the Interaction Flow Graph can reveal potential interaction design smells that may compromise intended functionality and hinder testing performance for VR applications. Ruizhen Gu, José Miguel Rojas, Donghwan Shin 0001 |
Autom. Softw. Eng. | 2 |
| 2025 | Can Test Generation and Program Repair Inform Automated Assessment of Programming Projects?abstractComputer Science educators assessing student programming assignments are typically responsible for two challenging tasks: grading and providing feedback. Producing grades that are fair and feedback that is useful to students is a goal common to most educators. In this context, automated test generation and program repair offer promising solutions for detecting bugs and suggesting corrections in students' code which could be leveraged to inform grading and feedback generation. Previous research on the applicability of these techniques to simple programming tasks (e.g., single-method algorithms) has shown promising results, but their effectiveness for more complex programming tasks remains unexplored. To fill this gap, this paper investigates the feasibility of applying existing test generation and program repair tools for assessing complex programming assignment projects. In a case study using a real-world Java programming assignment project with 296 incorrect student submissions, we found that generated tests were insufficient in detecting bugs in over 50% of cases, while full repairs could only be automatically generated for only 2.1% of submissions. Our findings indicate significant limitations in current tools for detecting bugs and repairing student submissions, highlighting the need for more advanced techniques to support automated assessment of complex assignment projects. Ruizhen Gu, José Miguel Rojas, Donghwan Shin 0001 |
ICST | 2 |
| 2025 | XRintTest: An Automated Framework for User Interaction Testing in Extended Reality ApplicationsabstractExtended Reality (XR) technologies offer immersive user experiences across diverse application domains, presenting unique testing challenges due to their spatial interaction paradigms. While existing works test XR applications through scene navigation and interaction triggering, they fail to synthesise realistic spatial input via specialised XR devices, such as 6 degrees of freedom controller gestures, that are essential for modern XR user experiences. To address this gap, we present XRintTest, an automated testing framework for Unity-based XR applications. XRintTest starts by constructing an XR User Interaction Graph that models interaction targets and required events. Leveraging this graph, it then automatically explores the XR scene under test and generates user interactions. We evaluated XRintTest on XRBench3D, a novel benchmark comprising seven XR scenes containing 367 distinct 3D user interactions. XRintTest shows great effectiveness, achieving 97% coverage of trigger and grab interactions across all scenes, 9x more effective and 5x more efficient than random exploration, while detecting runtime exceptions and functional defects. We open-sourced our tool and dataset at https://github.com/ruizhengu/XRintTest and https://github.com/ruizhengu/XRBench3D, respectively. A video demo is available on YouTube at https://youtu.be/K0Q6waE47Us. Ruizhen Gu, José Miguel Rojas, Donghwan Shin 0001 |
ASE | 2 |
| 2025 | Software testing for extended reality applications: a systematic mapping studyabstractAbstract Extended Reality (XR) is an emerging technology spanning diverse application domains and offering immersive user experiences. However, its unique characteristics, such as six degrees of freedom interactions, present significant testing challenges distinct from traditional 2D GUI applications, demanding novel testing techniques to build high-quality XR applications. This paper presents the first systematic mapping study on software testing for XR applications. We selected 34 studies focusing on techniques and empirical approaches in XR software testing for detailed examination. The studies are classified and reviewed to address the current research landscape, test facets, and evaluation methodologies in the XR testing domain. Additionally, we provide a repository summarising the mapping study, including datasets and tools referenced in the selected studies, to support future research and practical applications. Our study highlights open challenges in XR testing and proposes actionable future research directions to address the gaps and advance the field of XR software testing. Ruizhen Gu, José Miguel Rojas, Donghwan Shin 0001 |
Autom. Softw. Eng. | 2 |
| 2024 | Private-Keep Out? Understanding How Developers Account for Code Visibility in Unit TestingabstractRegression test maintenance costs can be reduced by striving to write tests that will require as few changes as possible in the future. Writing unit tests against behavior, as opposed to implementation, is one way to try to achieve this, because as long as the public API remains constant, units can be safely refactored without the need to also change the tests. However, in a study on 4,801 open-source Java projects reported in this paper, we found that 28% of projects contradict this advice, with tests that side-step the public API by directly calling non-public methods. We investigated why developers do not solely test public APIs-potentially increasing future test maintenance costs-by surveying 73 developers and conducting a systematic review of 60 StackOverflow posts dating from 2008–2023. Through numerical and thematic analyses, we uncover several findings, including (1) developers are disunited on whether to test only through public APIs or not; (2) those in favor of only testing through the public API tend to be more experienced and believe the need or desire to break with this is borne out of poor software design; while (3) those that test non-public methods directly are concerned about untested code complexity and overly intricate tests. Our findings provide multiple implications for future work, including automated developer support in the form of automated non-public method sequence replacement, and automated refactoring of production code using problematic public API-avoiding tests. Muhammad Firhard Roslan, José Miguel Rojas, Phil McMinn |
ICSME | 2 |
| 2024 | Viscount: A Direct Method Call Coverage Tool for JavaabstractWriting unit tests against implementation detail in production code, often embodied in non-public methods, is considered bad practice in formal and gray literature. This is because it leads to fragile tests that break easily when underlying implementation details change. For this reason, tests that focus on behavior are encouraged. One way to achieve this is to test units exclusively through their public API. However, our recent developer survey shows that this advice is not always followed in practice. Moreover, code coverage tools do not provide a way to determine which methods were called directly from tests, meaning there is no easy way to identify whether units make calls to non-public methods, other than through manual examination. To address this problem, we developed Viscount, a tool that can determine direct method call coverage for Java tests written in JUnit. Viscount reports the percentage of methods invoked directly from tests, according to their visibility - i.e., public or non-public (protected, package-private, or private). This can help developers and researchers identify tests that potentially need to be refactored or rewritten. In this paper, we describe Viscount's overall architecture, its core features, and how to use it. Viscount is also publicly available on GitHub: https://github.com/unittesting-nonpublic/viscount. A demo video of Viscount is available at: https://youtu.be/ZUyRtiUnbsU. Muhammad Firhard Roslan, José Miguel Rojas, Phil McMinn |
ICSME | 2 |
| 2022 | On the feasibility and challenges of synthesizing executable Espresso testsabstractSeveral tools have been proposed to automatically test Android applications, achieving outstanding results in terms of both code coverage and crash discovery. While useful for crash reproduction and bug-fixing, these tools usually do not present the generated interactions in a format that motivates developers to read and modify such tests later on. This hinders the ability of developers to add those tests to their existing test suites, or adapt them to new scenarios - common practices in modern software development where tests are maintained and evolve alongside production code. Iván Arcuschin, Juan P. Galeotti, Christian Ciccaroni, José Miguel Rojas |
AST | 4 |
| 2022 | An Empirical Comparison of EvoSuite and DSpot for Improving Developer-Written Test Suites with Respect to Mutation Score
Muhammad Firhard Roslan, José Miguel Rojas, Phil McMinn |
SSBSE | 2 |
| 2019 | Gamifying a Software Testing Course with Code DefendersabstractSoftware testing is an essential skill for software developers, but it is challenging to get students engaged in this activity. The Code Defenders game addresses this problem by letting students compete over code under test by either introducing faults ("attacking") or by writing tests ("defending") to reveal these faults. In this paper, we describe how we integrated Code Defenders as a semester-long activity of an undergraduate and graduate level university course on software testing. We complemented the regular course sessions with weekly Code Defenders sessions, addressing challenges such as selecting suitable code to test, managing games, and assessing performance. Our experience and our data show that the integration of Code Defenders was well-received by students and led them to practice testing thoroughly. Positive learning effects are evident as student performance improved steadily throughout the semester. Gordon Fraser 0001, Alessio Gambi, Marvin Kreis, José Miguel Rojas |
SIGCSE | 4 |
| 2019 | Special issue on mutation testing and analysisabstractThe effectiveness and usefulness of the proposed operators is evaluated against classical mutation operators on two case studies.The empirical evaluation indicates that the proposed operators are suitable to cause variability faults and therefore extend the capabilities of conventional mutation operators. René Just, Jens Krinke, José Miguel Rojas |
Softw. Test. Verification Reliab. | 4 |
| 2018 | Automated Accessibility Testing of Mobile AppsabstractIt is important to make mobile apps accessible, so as not to exclude users with common disabilities such as blindness, low vision, or color blindness. Even when developers are aware of these accessibility needs, the lack of tool support makes the development and assessment of accessible apps challenging. Some accessibility properties can be checked statically, but user interface widgets are often created dynamically and are not amenable to static checking. Some accessibility checking frameworks analyze accessibility properties at runtime, but have to rely on existing thorough test suites. In this paper, we introduce the idea of using automated test generation to explore the accessibility of mobile apps. We present the MATE tool (Mobile Accessibility Testing), which automatically explores apps while applying different checks for accessibility issues related to visual impairment. For each issue, MATE generates a detailed report that supports the developer in fixing the issue. Experiments on a sample of 73 apps demonstrate that MATE detects more basic accessibility problems than static analysis, and many additional types of accessibility problems that cannot be detected statically at all. Comparison with existing accessibility testing frameworks demonstrates that the independence of an existing test suite leads to the identification of many more accessibility problems. Even when enabling Android's assistive features like contrast enhancement, MATE can still find many accessibility issues. Marcelo Medeiros Eler, José Miguel Rojas, Yan Ge 0002, Gordon Fraser 0001 |
ICST | 2 |
| 2018 | How Do Automatically Generated Unit Tests Influence Software Maintenance?abstractGenerating unit tests automatically saves time over writing tests manually and can lead to higher code coverage. However, automatically generated tests are usually not based on realistic scenarios, and are therefore generally considered to be less readable. This places a question mark over their practical value: Every time a test fails, a developer has to decide whether this failure has revealed a regression fault in the program under test, or whether the test itself needs to be updated. Does the fact that automatically generated tests are harder to read outweigh the time-savings gained by their automated generation, and render them more of a hindrance than a help for software maintenance? In order to answer this question, we performed an empirical study in which participants were presented with an automatically generated or manually written failing test, and were asked to identify and fix the cause of the failure. Our experiment and two replications resulted in a total of 150 data points based on 75 participants. Whilst maintenance activities take longer when working with automatically generated tests, we found developers to be equally effective with manually written and automatically generated tests. This has implications on how automated test generation is best used in practice, and it indicates a need for research into the generation of more realistic tests. Sina Shamshiri, José Miguel Rojas, Juan P. Galeotti, Neil Walkinshaw, Gordon Fraser 0001 |
ICST | 2 |
| 2018 | Random or evolutionary search for object-oriented test suite generation?abstractSummary An important aim in software testing is constructing a test suite with high structural code coverage, that is, ensuring that most if not all of the code under test have been executed by the test cases comprising the test suite. Several search‐based techniques have proved successful at automatically generating tests that achieve high coverage. However, despite the well‐established arguments behind using evolutionary search algorithms (eg, genetic algorithms) in preference to random search, it remains an open question whether the benefits can actually be observed in practice when generating unit test suites for object‐oriented classes. In this paper, we report an empirical study on the effects of using evolutionary algorithms (including a genetic algorithm and chemical reaction optimization) to generate test suites, compared with generating test suites incrementally with random search. We apply the EVOSUITEunit test suite generator to 1000 classes randomly selected from the SF110 corpus of open‐source projects. Surprisingly, the results show that the difference is much smaller than one might expect: While evolutionary search covers more branches of the type where standard fitness functions provide guidance, we observed that, in practice, the vast majority of branches do not provide any guidance to the search. These results suggest that, although evolutionary algorithms are more effective at covering complex branches, a random search may suffice to achieve high coverage of most object‐oriented classes. Sina Shamshiri, José Miguel Rojas, Luca Gazzola, Gordon Fraser 0001, Phil McMinn, Leonardo Mariani, Andrea Arcuri |
Softw. Test. Verification Reliab. | 2 |
| 2017 | Code defenders: crowdsourcing effective tests and subtle mutants with a mutation testing gameabstractWriting good software tests is difficult and not every developer's favorite occupation. Mutation testing aims to help by seeding artificial faults (mutants) that good tests should identify, and test generation tools help by providing automatically generated tests. However, mutation tools tend to produce huge numbers of mutants, many of which are trivial, redundant, or semantically equivalent to the original program, automated test generation tools tend to produce tests that achieve good code coverage, but are otherwise weak and have no clear purpose. In this paper, we present an approach based on gamification and crowdsourcing to produce better software tests and mutants: The Code Defenders web-based game lets teams of players compete over a program, where attackers try to create subtle mutants, which the defenders try to counter by writing strong tests. Experiments in controlled and crowdsourced scenarios reveal that writing tests as part of the game is more enjoyable, and that playing Code Defenders results in stronger test suites and mutants than those produced by automated tools. José Miguel Rojas, Thomas D. White, Ben Clegg 0002, Gordon Fraser 0001 |
ICSE | 1 |
| 2017 | Generating unit tests with descriptive names or: would you name your children thing1 and thing2?abstractThe name of a unit test helps developers to understand the purpose and scenario of the test, and test names support developers when navigating amongst sets of unit tests. When unit tests are generated automatically, however, they tend to be given non-descriptive names such as “test0”, which provide none of the benefits a descriptive name can give a test. The underlying challenge is that automatically generated tests typically do not represent real scenarios and have no clear purpose other than covering code, which makes naming them di cult. In this paper, we present an automated approach which generates descriptive names for automatically generated unit tests by summarizing API-level coverage goals. The tests are optimized to be short, descriptive of the test, have a clear relation to the covered code under test, and allow developers to uniquely distinguish tests in a test suite. An empirical evaluation with 47 participants shows that developers agree with the synthesized names, and the synthesized names are equally descriptive as manually written names. Study participants were even more accurate and faster at matching code and tests with synthesized names compared to manually derived names. Ermira Daka, José Miguel Rojas, Gordon Fraser 0001 |
ISSTA | 2 |
| 2017 | A detailed investigation of the effectiveness of whole test suite generationabstractA common application of search-based software testing is to generate test cases for all goals defined by a coverage criterion (e.g., lines, branches, mutants). Rather than generating one test case at a time for each of these goals individually, whole test suite generation optimizes entire test suites towards satisfying all goals at the same time. There is evidence that the overall coverage achieved with this approach is superior to that of targeting individual coverage goals. Nevertheless, there remains some uncertainty on (a) whether the results generalize beyond branch coverage, (b) whether the whole test suite approach might be inferior to a more focused search for some particular coverage goals, and (c) whether generating whole test suites could be optimized by only targeting coverage goals not already covered. In this paper, we perform an in-depth analysis to study these questions. An empirical study on 100 Java classes using three different coverage criteria reveals that indeed there are some testing goals that are only covered by the traditional approach, although their number is only very small in comparison with those which are exclusively covered by the whole test suite approach. We find that keeping an archive of already covered goals along with the tests covering them and focusing the search on uncovered goals overcomes this small drawback on larger classes, leading to an improved overall effectiveness of whole test suite generation. José Miguel Rojas, Mattia Vivanti, Andrea Arcuri, Gordon Fraser 0001 |
Empir. Softw. Eng. | 1 |
| 2016 | Seeding strategies in search-based unit test generationabstractSummary Search‐based techniques have been applied successfully to the task of generating unit tests for object‐oriented software. However, as for any meta‐heuristic search, the efficiency heavily depends on many factors; seeding, which refers to the use of previous related knowledge to help solve the testing problem at hand, is one such factor that may strongly influence this efficiency. This paper investigates different seeding strategies for unit test generation, in particular seeding of numerical and string constants derived statically and dynamically, seeding of type information and seeding of previously generated tests. To understand the effects of these seeding strategies, the results of a large empirical analysis carried out on a large collection of open‐source projects from the SF110 corpus and the Apache Commons repository are reported. These experiments show with strong statistical confidence that, even for a testing tool already able to achieve high coverage, the use of appropriate seeding strategies can further improve performance. © 2016 The Authors. Software Testing, Verification and Reliability Published by John Wiley & Sons Ltd. José Miguel Rojas, Gordon Fraser 0001, Andrea Arcuri |
Softw. Test. Verification Reliab. | 1 |
| 2015 | Random or Genetic Algorithm Search for Object-Oriented Test Suite Generation?abstractAchieving high structural coverage is an important aim in software testing. Several search-based techniques have proved successful at automatically generating tests that achieve high coverage. However, despite the well- established arguments behind using evolutionary search algorithms (e.g., genetic algorithms) in preference to random search, it remains an open question whether the benefits can actually be observed in practice when generating unit test suites for object-oriented classes. In this paper, we report an empirical study on the effects of using a genetic algorithm (GA) to generate test suites over generating test suites incrementally with random search, by applying the EvoSuite unit test suite generator to 1,000 classes randomly selected from the SF110 corpus of open source projects. Surprisingly, the results show little difference between the coverage achieved by test suites generated with evolutionary search compared to those generated using random search. A detailed analysis reveals that the genetic algorithm covers more branches of the type where standard fitness functions provide guidance. In practice, however, we observed that the vast majority of branches in the analyzed projects provide no such guidance. Sina Shamshiri, José Miguel Rojas, Gordon Fraser 0001, Phil McMinn |
GECCO | 2 |
| 2015 | Automated unit test generation during software development: a controlled experiment and think-aloud observationsabstractAutomated unit test generation tools can produce tests that are superior to manually written ones in terms of code coverage, but are these tests helpful to developers while they are writing code? A developer would first need to know when and how to apply such a tool, and would then need to understand the resulting tests in order to provide test oracles and to diagnose and fix any faults that the tests reveal. Considering all this, does automatically generating unit tests provide any benefit over simply writing unit tests manually? We empirically investigated the effects of using an automated unit test generation tool (EvoSuite) during development. A controlled experiment with 41 students shows that using EvoSuite leads to an average branch coverage increase of +13%, and 36% less time is spent on testing compared to writing unit tests manually. However, there is no clear effect on the quality of the implementations, as it depends on how the test generation tool and the generated tests are used. In-depth analysis, using five think-aloud observations with professional programmers, confirms the necessity to increase the usability of automated unit test generation tools, to integrate them better during software development, and to educate software developers on how to best use those tools. José Miguel Rojas, Gordon Fraser 0001, Andrea Arcuri |
ISSTA | 1 |
| 2015 | Do Automatically Generated Unit Tests Find Real Faults? An Empirical Study of Effectiveness and Challenges (T)abstractRather than tediously writing unit tests manually, tools can be used to generate them automatically - sometimes even resulting in higher code coverage than manual testing. But how good are these tests at actually finding faults? To answer this question, we applied three state-of-the-art unit test generation tools for Java (Randoop, EvoSuite, and Agitar) to the 357 real faults in the Defects4J dataset and investigated how well the generated test suites perform at detecting these faults. Although the automatically generated test suites detected 55.7% of the faults overall, only 19.9% of all the individual test suites detected a fault. By studying the effectiveness and problems of the individual tools and the tests they generate, we derive insights to support the development of automated unit test generators that achieve a higher fault detection rate. These insights include 1) improving the obtained code coverage so that faulty statements are executed in the first instance, 2) improving the propagation of faulty program states to an observable output, coupled with the generation of more sensitive assertions, and 3) improving the simulation of the execution environment to detect faults that are dependent on external factors such as date and time. Sina Shamshiri, René Just, José Miguel Rojas, Gordon Fraser 0001, Phil McMinn, Andrea Arcuri |
ASE | 3 |
| 2015 | Combining Multiple Coverage Criteria in Search-Based Unit Test Generation
José Miguel Rojas, José Campos 0001, Mattia Vivanti, Gordon Fraser 0001, Andrea Arcuri |
SSBSE | 1 |
| 2013 | A CLP heap solver for test case generationabstractAbstract One of the main challenges to software testing today is to efficiently handle heap-manipulating programs. These programs often build complex, dynamically allocated data structures during execution and, to ensure reliability, the testing process needs to consider all possible shapes these data structures can take. This creates scalability issues since high (often exponential) numbers of shapes may be built due to the aliasing of references. This paper presents a novel CLP heap solver for the test case generation of heap-manipulating programs that is more scalable than previous proposals, thanks to the treatment of reference aliasing by means of disjunction, and to the use of advanced back-propagation of heap related constraints. In addition, the heap solver supports the use of heap assumptions to avoid aliasing of data that, though legal, should not be provided as input. Elvira Albert, Maria Garcia de la Banda, Miguel Gómez-Zamalloa, José Miguel Rojas, Peter J. Stuckey |
Theory Pract. Log. Program. | 4 |
| 2012 | Automated Extraction of Abstract Behavioural Models from JMS Applications
Elvira Albert, Bjarte M. Østvold, José Miguel Rojas |
FMICS | 3 |
| 2012 | A Framework for Guided Test Case Generation in Constraint Logic Programming
José Miguel Rojas, Miguel Gómez-Zamalloa |
LOPSTR | 1 |
| 2011 | Resource-Driven CLP-Based Test Case Generation
Elvira Albert, Miguel Gómez-Zamalloa, José Miguel Rojas |
LOPSTR | 3 |
| 2010 | Compositional CLP-Based Test Data Generation for Imperative Languages
Elvira Albert, Miguel Gómez-Zamalloa, José Miguel Rojas, Germán Puebla |
LOPSTR | 3 |
| 2009 | On the Solutions of NP-Complete Problems by Means of jNEP Run on Computers
Emilio del Rosal García, José Miguel Rojas, Rafael Núñez Hervás, Carlos Castañeda Marroquín, Alfonso Ortega 0002 |
ICAART | 2 |