EDBT 2026 Demo / reviewers in the wild / expert
Francisco Gomes de Oliveira Neto
dblp:56/1704 · also Francisco G. Oliveira Neto, Francisco Gomes 0001
· DBLP profile ↗
25ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0001-9226-5417ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 24 · 5 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Logs to Lessons: An Exploration of LLM-based Log Summarization for Debugging Automotive SoftwareabstractIdentifying where faults occur is an essential part of debugging, yet examining extensive system logs can be slow and mentally demanding, especially in complex software environments. One emerging strategy to enhance log analysis is to employ large language models (LLMs) to distill log information into more manageable summaries that can guide human reasoning during diagnosis. We report on a case study carried out in an automotive setting, where engineers investigated actual failures with and without support from an LLM-based summarization tool. During fault localization sessions where participants analyzed real failure logs, we collected cognitive load measurements, observed their reasoning processes, and gathered feedback on both the LLM-based summarization and the workflow through post-session interviews. Our results indicate that although the use of summaries raised certain cognitive demands, particularly related to mental effort and time pressure, participants experienced less frustration overall and considered the support helpful in focusing their attention. They also expressed a clear interest in being able to shape and refine summaries as their understanding evolved. These findings offer insights into how LLM-generated summaries influence practitioners’ diagnostic work and point toward the need for more adaptive, interactive, and workflow-aware support. Anton Ekström, Hampus Rhedin Stam, Francisco Gomes de Oliveira Neto, Gregory Gay 0002, Sabina Edenlund |
AST | 3 |
| 2026 | REST-at: An LLM-Based Tool for Automating Traceability between Requirements and Test CasesabstractLarge Language Models (LLMs) are increasingly used across Software Engineering tasks, and their ability to process natural language makes them particularly well suited for aligning requirements and testing artifacts. In this study, we investigate AI-supported traceability in software development through the creation and evaluation of REST-at, a tool that facilitates alignment between requirements engineering and software testing (REST) by automatically generating trace links between requirements and manual system-level tests. To evaluate its effectiveness, we compare five LLMs, including proprietary models such as GPT-4 and open-weight alternatives like Mixtral 8x7B and LLAMA 3 70B. The results indicate that open-weight models can achieve comparable or even superior levels of precision and recall relative to larger proprietary models. Among them, LLAMA 3 70B consistently provided robust and reliable link generation between requirements and tests. We also observe a high sensitivity to prompt design, with model performance varying according to prompt structure. Based on these insights, we provide practical recommendations for researchers and practitioners concerning hardware demands, governance and transparency, and the importance of standardized documentation when applying open-weight models to software engineering artifacts. Nicole Leon-Quinstedt, Bao Lindgren, Mert Yurdakul, Francisco Gomes de Oliveira Neto |
AST | 4 |
| 2025 | Comparative analysis of text mining and clustering techniques for assessing functional dependency between manual test casesabstractAbstract Text mining techniques, particularly those leveraging machine learning for natural language processing, have gained significant attention for qualitative data analysis in software testing. However, their complexity and lack of transparency can pose challenges, especially in safety-critical domains where simpler, interpretable solutions are often preferred unless accuracy is heavily compromised. This study investigates the trade-offs between complexity, effort, accuracy, and utility in text mining and clustering techniques, focusing on their application for detecting functional dependencies among manual integration test cases in safety-critical systems. Using empirical data from an industrial testing project at ALSTOM Sweden, we evaluate various string distance methods, NCD compressors, and machine learning approaches. The results highlight the impact of preprocessing techniques, such as tokenization, and intrinsic factors, such as text length, on algorithm performance. Findings demonstrate how text mining and clustering can be optimized for safety-critical contexts, offering actionable insights for researchers and practitioners aiming to balance simplicity and effectiveness in their testing workflows. Sahar Tahvili, Leo Hatvani, Michael Felderer, Francisco Gomes de Oliveira Neto, Wasif Afzal, Robert Feldt |
Softw. Qual. J. | 4 |
| 2025 | The Impact of Prompt Programming on Function-Level Code GenerationabstractLarge Language Models (LLMs) are increasingly used by software engineers for code generation. However, limitations of LLMs such as irrelevant or incorrect code have highlighted the need for prompt programming (or prompt engineering) where engineers apply specific prompt techniques (e.g., chain-of-thought or input-output examples) to improve the generated code. While some prompt techniques have been studied, the impact of different techniques—and their interactions— on code generation is still not fully understood. In this study, we introduce CodePromptEval, a dataset of 7072 prompts designed to evaluate five prompt techniques (few-shot, persona, chain-of-thought, function signature, list of packages) and their effect on the correctness, similarity, and quality of complete functions generated by three LLMs (GPT-4o, Llama3, and Mistral). Our findings show that while certain prompt techniques significantly influence the generated code, combining multiple techniques does not necessarily improve the outcome. Additionally, we observed a trade-off between correctness and quality when using prompt techniques. Our dataset and replication package enable future research on improving LLM-generated code and evaluating new prompt techniques. Ranim Khojah, Francisco Gomes de Oliveira Neto, Mazen Mohamad, Philipp Leitner 0001 |
IEEE Trans. Software Eng. | 2 |
| 2024 | Exploring the Role of Automation in Duplicate Bug Report Detection: An Industrial Case StudyabstractDuplicate bug reports can increase technical debt and tester work-load in long-running software projects. Many automated techniques have been proposed to detect potential duplicate reports. However, such techniques have not seen widespread industrial adoption. Our objective in this study is to better understand how automated techniques could effectively be employed within a tester's duplicate detection workflow. We are particularly interested in exploring the potential of a human-in-the-loop scenario where tools and humans work together to make duplicate determinations. Malte Götharsson, Karl Stahre, Gregory Gay 0002, Francisco Gomes de Oliveira Neto |
AST | 4 |
| 2024 | The Impact of Compiler Warnings on Code Quality in C++ ProjectsabstractModern compilers often offer a variety of warning flags, which developers can enable to get feedback on code that, while syntactically correct, may be problematic. In the case of C++, one example of such "correct but problematic" code is code that leads to undefined behavior (UB). The usage of compiler warnings has long been suspected as a way to decrease bugs and increase code quality. However, empirical evidence that supports this hypothesis is rare. In this study, we present evidence from a study of 127 open source C++ projects. We categorize their usage of compiler warnings into five groups based on which warning flags are being used, and analyse the relationship between compiler warnings and five quality metrics (bugs, critical issues, vulnerabilities, code smells, and technical debt) using Bayesian analysis. We conclude that, in general, compiler warnings indeed correlate with, and potentially cause, higher code quality, with the clearest impact being on the number of critical issues in a project. Using stricter warning flags expectantly correlates with higher code quality in our study objects. However, there are substantial differences between projects, which we attribute to the project's individual development culture. That is, while warnings matter, other factors such as quality culture, are likely to be even more important to source code quality. Albin Johansson, Carl Holmberg, Francisco Gomes de Oliveira Neto, Philipp Leitner 0001 |
ICPC | 3 |
| 2024 | Integrating Mutation Testing Into Developer Workflow: An Industrial Case StudyabstractMutation testing is a potentially effective method to assess test suite adequacy. Researchers have made mutation testing more computationally efficient, and new frameworks are regularly emerging. However, there is still limited adoption of mutation testing in industry. We hypothesize that such adoption is hindered by a lack of guidance on how to effectively and efficiently utilize mutation testing in a development workflow. To that end, we have conducted an industrial case study exploring the technical challenges of implementing mutation testing in continuous integration, what information from mutation testing is of use to developers, and how that information should be presented (in textual and visual form). Our results reveal five technical challenges of integrating mutation testing and nine key findings regarding how the results of mutation testing are used and presented. We also offer a dashboard to visualize mutation testing results, as well as 16 recommendations for making effective use of mutation testing in practice1. Stefan Alexander van Heijningen, Theo Wiik, Francisco Gomes de Oliveira Neto, Gregory Gay 0002, Kim Viggedal, David Friberg |
ASE | 3 |
| 2023 | Evaluating the Trade-offs of Text-based Diversity in Test PrioritisationabstractDiversity-based techniques (DBT) have been cost-effective by prioritizing the most dissimilar test cases to detect faults at earlier stages of test execution. Diversity is measured on test specifications to convey how different test cases are from one another. However, there is little research on the trade-off of diversity measures based on different types of text-based specification (lexicographical or semantics). Particularly because the text content in test scripts vary widely from unit (e.g., code) to system-level (e.g., natural language). This paper compares and evaluates the cost-effectiveness in coverage and failures of different text-based diversity measures for different levels of tests. We perform an experiment on the test suites of 7 open source projects on the unit level, and 2 industry projects on the integration and system level. Our results show that test suites prioritised using semantic-based diversity measures causes a small improvement in requirements coverage, as opposed to lexical diversity that showed less coverage than random for system-level artefacts. In contrast, using lexical-based measures such as Jaccard or Levenshtein to prioritise code artefacts yield better failure coverage across all levels of tests. We summarise our findings into a list of recommendations for using semantic or lexical diversity on different levels of testing. Ranim Khojah, Chi Hong Chao, Francisco Gomes de Oliveira Neto |
AST | 3 |
| 2023 | Test Maintenance for Machine Learning Systems: A Case Study in the Automotive IndustryabstractMachine Learning (ML) systems have seen widespread use for automated decision making. Testing is essential to ensure the quality of these systems, especially safety-critical autonomous systems in the automotive domain. ML systems introduce new challenges with the potential to affect test maintenance, the process of updating test cases to match the evolving system. We conducted an exploratory case study in the automotive domain to identify factors that affect test maintenance for ML systems, as well as to make recommendations to improve the maintenance process. Based on interview and artifact analysis, we identified 14 factors affecting maintenance, including five especially relevant for ML systems—with the most important relating to non-determinism and large input spaces. We also proposed ten recommendations for improving test maintenance, including four targeting ML systems—in particular, emphasizing the use of test oracles tolerant to acceptable non-determinism. The study’s findings expand our knowledge of test maintenance for an emerging class of systems, benefiting the practitioners testing these systems. Lukas Berglund, Tim Grube, Gregory Gay 0002, Francisco Gomes de Oliveira Neto, Dimitrios Platis |
ICST | 4 |
| 2023 | An Industrial Study on the Challenges and Effects of Diversity-Based Testing in Continuous IntegrationabstractMany test prioritisation techniques have been proposed in order to improve test effectiveness of Continuous Integration (CI) pipelines. Particularly, diversity-based testing (DBT) has shown promising and competitive results to improve test effectiveness. However, the technical and practical challenges of introducing test prioritisation in CI pipelines are rarely discussed, thus hindering the applicability and adoption of those proposed techniques. This research builds on our prior work in which we evaluated diversity-based techniques in an industrial setting. This work investigates the factors that influence the adoption of DBT both in connection to improvements in test cost-effectiveness, as well as the process and human related challenges to transfer and use DBT prioritisation in CI pipelines. We report on a case study considering the CI pipeline of Axis Communications in Sweden. We performed a thematic analysis of a focus group interview with senior practitioners at the company to identify the challenges and perceived benefits of using test prioritisation in their test process. Our thematic analysis reveals a list of ten challenges and seven perceived effects of introducing test prioritisation in CI cycles. For instance, our participants emphasized the importance of introducing comprehensible and transparent techniques that instill trust in its users. Moreover, practitioners prefer techniques compatible with their current test infrastructure (e.g., test framework and environments) in order to reduce instrumentation efforts and avoid disrupting their current setup. In conclusion, we have identified tradeoffs between different test prioritisation techniques pertaining to the technical, process and human aspects of regression testing in CI. We summarize those findings in a list of seven advantages that refer to specific stakeholder interests and describe the effects of adopting DBT in CI pipelines. Azeem Ahmad, Francisco Gomes de Oliveira Neto, Eduard Paul Enoiu, Kristian Sandahl, Ola Leifler |
QRS | 2 |
| 2023 | Investigating Software Testing and Maintenance of Open-Source Distributed LedgerabstractA distributed ledger is the backbone of all blockchain solutions. It provides a shared database spreading across a network of nodes. The number of DL solutions and their implementations has grown in recent years. Besides the architectural and performance promises of thesesolutions, organizations seekingto implement DL also need to consider the overall quality of the software available and its ecosystem. Particularly, previous research has identified the need to better understand the testing and maintenance practices behind these types of technologies. This paper investigates the testing and maintenance of 18 different open-source projects that implement distributed ledgers. We perform a manual inspection of test artefacts and mine the history of commits, issues and contributors of the chosen projects to understand the landscape of testing and maintenance in these projects. Our findings suggest that unit and integration tests are present in most projects, they do not follow a holistic system testing approach. Moreover, projects rely on a small team of core contributors (5 on average). While the projects are continuously maintained, larger changes are uncommon. Our results can be used for benchmarking and pinpointing areas of improvement for the development of distributed ledgers. Petya Hristova Cvitic, Felix Dobslaw, Francisco Gomes de Oliveira Neto |
SANER | 3 |
| 2022 | An Empirical Analysis of Microservices Systems Using Consumer-Driven Contract TestingabstractTesting has a prominent role in revealing faults in software based on microservices. One of the most important discussion points in MSAs is the granularity of services, often in different levels of abstraction. Similarly, the granularity of tests in MSAs is reflected in different test types. However, it is challenging to conceptualize how the overall testing architecture comes together when combining testing in different levels of abstraction for microservices. There is no empirical evidence on the overall testing architecture in such microservices implementations. Furthermore, there is a need to empirically understand how the current state of practice resonates with existing best practices on testing. In this study, we mine Github to find different candidate projects for an in-depth, qualitative assessment of their test artifacts. We analyze 16 repositories that use microservices and include various test artifacts. We focus on four projects that use consumer-driven-contract testing. Our results demonstrate how these projects cover different levels of testing. This study (i) drafts a testing architecture including activities and artifacts, and (ii) demonstrates how these align with best practices and guidelines. Our proposed architecture helps the categorization of system and test artifacts in empirical studies of microservices. Finally, we showcase a view of the boundaries between different levels of testing in systems using microservices. Hamdy Michael Ayas, Hartmut Fischer, Philipp Leitner 0001, Francisco Gomes de Oliveira Neto |
SEAA | 4 |
| 2022 | Comparing Input Prioritization Techniques for Testing Deep Learning AlgorithmsabstractDeep learning (DL) systems are becoming an essential part of software systems, so it is necessary to test them thoroughly. This is a challenging task since the test sets can grow over time as the new data is being acquired, and it becomes time-consuming. Input prioritization is necessary to reduce the testing time since prioritized test inputs are more likely to reveal the erroneous behavior of a DL system earlier during test execution. Input prioritization approaches have been rudimentary analyzed against each other, this study compares different input prioritization techniques regarding their effectiveness and efficiency. This work considers surprise adequacy, autoencoder-based, and similarity-based input prioritization approaches in the example of testing a DL image classification algorithms applied on MNIST, Fashion-MNIST, CIFAR-10, and STL-10 datasets. To measure effectiveness and efficiency, we use a modified APFD (Average Percentage of Fault Detected), and set up & execution time, respectively. We observe that the surprise adequacy is the most effective (0.785 to 0.914 APFD). The autoencoder-based and similarity-based techniques are less effective, with the performance from 0.532 to 0.744 APFD and 0.579 to 0.709 APFD, respectively. In contrast, the similarity-based and surprise adequacy-based approaches are the most and least efficient, respectively. The findings in this work demonstrate the trade-off between the considered input prioritization techniques to understanding their practical applicability for testing DL algorithms. Vasilii Mosin, Miroslaw Staron, Darko Durisic, Francisco Gomes de Oliveira Neto, Sushant Kumar Pandey, Ashok Chaitanya Koppisetty |
SEAA | 4 |
| 2022 | A Method to Assess and Argue for Practical Significance in Software EngineeringabstractA key goal of empirical research in software engineering is to assess practical significance, which answers the question whether the observed effects of some compared treatments show a relevant difference in practice in realistic scenarios. Even though plenty of standard techniques exist to assess statistical significance, connecting it to practical significance is not straightforward or routinely done; indeed, only a few empirical studies in software engineering assess practical significance in a principled and systematic way. In this paper, we argue that Bayesian data analysis provides suitable tools to assess practical significance rigorously. We demonstrate our claims in a case study comparing different test techniques. The case study's data was previously analyzed (Afzalet al., 2015) using standard techniques focusing on statistical significance. Here, we build a multilevel model of the same data, which we fit and validate using Bayesian techniques. Our method is to apply cumulative prospect theory on top of the statistical model to quantitatively connect our statistical analysis output to a practically meaningful context. This is then the basis both for assessing and arguing for practical significance. Our study demonstrates that Bayesian analysis provides a technically rigorous yet practical framework for empirical software engineering. A substantial side effect is that any uncertainty in the underlying data will be propagated through the statistical model, and its effects on practical significance are made clear. Thus, in combination with cumulative prospect theory, Bayesian analysis supports seamlessly assessing practical significance in an empirical software engineering context, thus potentially clarifying and extending the relevance of research for practitioners. Richard Torkar, Carlo A. Furia, Robert Feldt, Francisco Gomes de Oliveira Neto, Lucas Gren, Per Lenberg, Neil A. Ernst |
IEEE Trans. Software Eng. | 4 |
| 2021 | A Multi-factor Approach for Flaky Test Detection and Automated Root Cause AnalysisabstractDevelopers often spend time to determine whether test case failures are real failures or flaky. The flaky tests, also known as non-deterministic tests, switch their outcomes without any modification in the codebase, hence reducing the confidence of developers during maintenance as well as in the quality of a product. Re-running test cases to reveal flakiness is resource-consuming, unreliable and does not reveal the root causes of test flakiness. Our paper evaluates a multi-factor approach to identify flaky test executions implemented in a tool named MDF laker. The four factors are: trace-back coverage, flaky frequency, number of test smells, and test size. Based on the extracted factors, MDFlaker uses k-Nearest Neighbor (KNN) to determine whether failed test executions are flaky. We investigate MDFlaker in a case study with 2166 test executions from different open-source repositories. We evaluate the effectiveness of our flaky detection tool. We illustrate how the multi-factor approach can be used to reveal root causes for flakiness, and we conduct a qualitative comparison between MDF laker and other tools proposed in literature. Our results show that the combination of different factors can be used to identify flaky tests. Each factor has its own trade-off, e.g., trace-back leads to many true positives, while flaky frequency yields more true negatives. Therefore, specific combinations of factors enable classification for testers with limited information (e.g., not enough test history information). Azeem Ahmad, Francisco Gomes de Oliveira Neto, Zhixiang Shi, Kristian Sandahl, Ola Leifler |
APSEC | 2 |
| 2021 | Requirements engineering challenges and practices in large-scale agile system developmentabstractAgile methods have become mainstream even in large-scale systems engineering companies that need to accommodate different development cycles of hardware and software. For such companies, requirements engineering is an essential activity that involves upfront and detailed analysis which can be at odds with agile development methods. This paper presents a multiple case study with seven large-scale systems companies, reporting their challenges, together with best practices from industry. We also analyze literature about two popular large-scale agile frameworks, SAFe® and LeSS, to derive potential solutions for the challenges. Our results are based on 20 qualitative interviews, five focus groups, and eight cross-company workshops which we used to both collect and validate our results. We found 24 challenges which we grouped in six themes, then mapped to solutions from SAFe®, LeSS, and our companies, when available. In this way, we contribute a comprehensive overview of RE challenges in relation to large-scale agile system development, evaluate the degree to which they have been addressed, and outline research gaps. We expect these results to be useful for practitioners who are responsible for designing processes, methods, or tools for large scale agile development as well as guidance for researchers. Rashidah Kasauli, Eric Knauss, Jennifer Horkoff, Grischa Liebel, Francisco Gomes de Oliveira Neto |
J. Syst. Softw. | 5 |
| 2020 | Evaluating the Effects of Different Requirements Representations on Writing Test Cases
Francisco Gomes de Oliveira Neto, Jennifer Horkoff, Richard Berntsson-Svensson, David Issa Mattos, Alessia Knauss |
REFSQ | 1 |
| 2020 | An empirical study of bots in software development: characteristics and challenges from a practitioner's perspectiveabstractSoftware engineering bots – automated tools that handle tedious tasks – are increasingly used by industrial and open source projects to improve developer productivity. Current research in this area is held back by a lack of consensus of what software engineering bots (DevBots) actually are, what characteristics distinguish them from other tools, and what benefits and challenges are associated with DevBot usage. In this paper we report on a mixed-method empirical study of DevBot usage in industrial practice. We report on findings from interviewing 21 and surveying a total of 111 developers. We identify three different personas among DevBot users (focusing on autonomy, chat interfaces, and “smartness”), each with different definitions of what a DevBot is, why developers use them, and what they struggle with.We conclude that future DevBot research should situate their work within our framework, to clearly identify what type of bot the work targets, and what advantages practitioners can expect. Further, we find that there currently is a lack of general purpose “smart” bots that go beyond simple automation tools or chat interfaces. This is problematic, as we have seen that such bots, if available, can have a transformative effect on the projects that use them. Linda Erlenhov, Francisco Gomes de Oliveira Neto, Philipp Leitner 0001 |
ESEC/SIGSOFT FSE | 2 |
| 2019 | Estimating Return on Investment for GUI Test Automation FrameworksabstractAutomated graphical user interface (GUI) tests can reduce manual testing activities and increase test frequency. This motivates the conversion of manual test cases into automated GUI tests. However, it is not clear whether such automation is cost-effective given that GUI automation scripts add to the code base and demand maintenance as a system evolves. In this paper, we introduce a method for estimating maintenance cost and Return on Investment (ROI) for Automated GUI Testing (AGT). The method utilizes the existing source code change history and has the potential to be used for the evaluation of other testing or quality assurance automation technologies. We evaluate the method for a real-world, industrial software system and compare two fundamentally different AGT frameworks, namely Selenium and EyeAutomate, to estimate and compare their ROI. We also report on their defect-finding capabilities and usability. The quantitative data is complemented by interviews with employees at the company the study has been conducted at. The method was successfully applied, and estimated maintenance cost and ROI for both frameworks are reported. Overall, the study supports earlier results showing that implementation time is the leading cost for introducing AGT. The findings further suggest that, while EyeAutomate tests are significantly faster to implement, Selenium tests require more of a programming background but less maintenance. Felix Dobslaw, Robert Feldt, David Michaelsson, Patrick Haar, Francisco Gomes de Oliveira Neto, Richard Torkar |
ISSRE | 5 |
| 2019 | Evolution of statistical analysis in empirical software engineering research: Current state and steps forward
Francisco Gomes de Oliveira Neto, Richard Torkar, Robert Feldt, Lucas Gren, Carlo A. Furia |
J. Syst. Softw. | 1 |
| 2018 | Visualizing Test Diversity to Support Test OptimisationabstractDiversity has been used as an effective criteria to optimise test suites for cost-effective testing. Particularly, diversity-based (alternatively referred to as similarity-based) techniques have the benefit of being generic and applicable across different Systems Under Test (SUT), and have been used to automatically select or prioritise large sets of test cases. However, there is a challenge in how to present diversity information to developers and testers since results are typically many-dimensional. Furthermore, the generality of diversity-based approaches makes it harder to choose when and where to apply them. In this paper we address these challenges by investigating: i) what are the trade-offs in using different sources of diversity (e.g., diversity of test requirements or test scripts) to optimise large test suites, and ii) how visualisation of test diversity data can assist testers for test optimisation and improvement. We perform a case study on three industrial projects and present quantitative results on the fault detection capabilities and redundancy levels of different sets of test cases. Our key result is that test similarity maps, based on pair-wise diversity calculations, helped industrial practitioners identify issues with their test repositories and decide on actions to improve. We conclude that the visualisation of diversity information can assist testers in their maintenance and optimisation activities. Francisco Gomes de Oliveira Neto, Robert Feldt, Linda Erlenhov, José Benardi de Souza Nunes |
APSEC | 1 |
| 2016 | Full modification coverage through automatic similarity-based test case selection
Francisco Gomes de Oliveira Neto, Richard Torkar, Patrícia Duarte de Lima Machado |
Inf. Softw. Technol. | 1 |
| 2015 | An Initiative to Improve Reproducibility and Empirical Evaluation of Software Testing TechniquesabstractThe current concern regarding quality of evaluation performed in existing studies reveals the need for methods and tools to assist in the definition and execution of empirical studies and experiments. However, when trying to apply general methods from empirical software engineering in specific fields, such as evaluation of software testing techniques, new obstacles and threats to validity appears, hindering researchers' use of empirical methods. This paper discusses those issues specific for evaluation of software testing techniques and proposes an initiative for a collaborative effort to encourage reproducibility of experiments evaluating software testing techniques (STT). We also propose the development of a tool that enables automatic execution and analysis of experiments producing a reproducible research compendia as output that is, in turn, shared among researchers. There are many expected benefits from this Endeavour, such as providing a foundation for evaluation of existing and upcoming STT, and allowing researchers to devise and publish better experiments. Francisco Gomes de Oliveira Neto, Richard Torkar, Patrícia Duarte de Lima Machado |
ICSE (2) | 1 |
| 2011 | On the use of a similarity function for test case selection in the context of model-based testingabstractAbstract Test case selection in model‐based testing is discussed focusing on the use of a similarity function. Automatically generated test suites usually have redundant test cases. The reason is that test generation algorithms are usually based on structural coverage criteria that are applied exhaustively. These criteria may not be helpful to detect redundant test cases as well as the suites are usually impractical due to the huge number of test cases that can be generated. Both problems are addressed by applying a similarity function. The idea is to keep in the suite the less similar test cases according to a goal that is defined in terms of the intended size of the test suite. The strategy presented is compared with random selection by considering transition‐based and fault‐based coverage. The results show that, in most of the cases, similarity‐based selection can be more effective than random selection when applied to automatically generated test suites. Copyright © 2009 John Wiley & Sons, Ltd. Emanuela Gadelha Cartaxo, Patrícia Duarte de Lima Machado, Francisco Gomes de Oliveira Neto |
Softw. Test. Verification Reliab. | 3 |
| 2007 | Test case generation by means of UML sequence diagrams and labeled transition systemsabstractWe present a systematic procedure of functional test case generation for feature testing of mobile phone applications. A feature is an increment of functionality, usually with a coherent purpose that is added on top of a basic system. Feature are usually developed and tested separately from the basic system as independent modules. The procedure is based on model-based testing techniques with test cases generated from UML sequence diagrams translated into labeled transition systems (LTSs). A case study is presented to illustrate the application of the procedure. The work is part of a research initiative for automation of test case generation, selection and evaluation of Motorola mobile phone applications. Emanuela Gadelha Cartaxo, Francisco Gomes de Oliveira Neto, Patrícia Duarte de Lima Machado |
SMC | 2 |