EDBT 2026 Demo / reviewers in the wild / expert
Sergio Segura
dblp:97/1263 · also Sergio Segura-Ramos
· DBLP profile ↗
49ranked-venue papers
11as first author
18since 2021 · last 2026
0000-0001-8816-6213ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 44 · 10 first-author · 17 since 2021Artificial intelligence and machine learning · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When AI Joins the Team: Understanding Human-Agent Collaboration in Pull Requests
Miguel Romero-Arjona, Saman A. Barakat, Alberto Martin-Lopez, Ana Belén Sánchez, Gabriele Bavota, Sergio Segura |
COMPSAC | 6 |
| 2026 | Metamorphic Testing of Vision-Language Action-Enabled Robots
Sergio Segura, Shaukat Ali 0001, Aitor Arrieta |
ICST | 2 |
| 2026 | Meta-Fair: AI-assisted fairness testing of large language modelsabstractFairness—the absence of unjustified bias—is a core principle in the development of Artificial Intelligence (AI) systems, yet it remains difficult to assess and enforce. Current approaches to fairness testing in large language models (LLMs) often rely on manual evaluation, fixed templates, deterministic heuristics, and curated datasets, making them resource-intensive and difficult to scale. This work aims to lay the groundwork for a novel, automated method for testing fairness in LLMs, reducing the dependence on domain-specific resources and broadening the applicability of current approaches. Our approach, Meta-Fair, is based on two key ideas. First, we adopt metamorphic testing to uncover bias by examining how model outputs vary in response to controlled modifications of input prompts, defined by metamorphic relations (MRs). Second, we propose exploiting the potential of LLMs for both test case generation and output evaluation, leveraging their capability to generate diverse inputs and classify outputs effectively. The proposal is complemented by three open-source tools supporting LLM-driven generation, execution, and evaluation of test cases. We report the findings of several experiments involving 12 pre-trained LLMs, 14 MRs, 5 bias dimensions, and 7.9K automatically generated test cases. The results show that Meta-Fair is effective in uncovering bias in LLMs, achieving an average precision of 92% and revealing biased behaviour in 29% of executions. Additionally, LLMs prove to be reliable and consistent evaluators, with the best-performing models achieving F1-scores of up to 0.79. Although non-determinism affects consistency, these effects can be mitigated through careful MR design. This work highlights the feasibility and potential of integrating metamorphic testing with LLM-driven test generation and assessment. While challenges remain to ensure broader applicability, the results indicate a promising path towards an unprecedented level of automation in LLM testing. Miguel Romero-Arjona, José Antonio Parejo, Juan C. Alonso, Ana Belén Sánchez, Aitor Arrieta, Sergio Segura |
Inf. Softw. Technol. | 6 |
| 2026 | Test Oracle Generation for REST APIsabstractThe number and complexity of test case generation tools for REST APIs have significantly increased in recent years. These tools excel in automating input generation but are limited by their test oracles, which can only detect crashes, regressions, and violations of API specifications or design best practices. This article introduces AGORA+, an approach for generating test oracles for REST APIs through the detection of invariants—output properties that should always hold. AGORA+ learns the expected behavior of an API by analyzing API requests and their corresponding responses. We enhanced the Daikon tool for dynamic detection of likely invariants, adding new invariant types and creating a front-end called Beet. Beet translates any OpenAPI specification and a set of API requests and responses into Daikon inputs. AGORA+ can detect 106 different types of invariants in REST APIs. We also developed PostmanAssertify, which converts the invariants identified by AGORA+ into executable JavaScript assertions. AGORA+ achieved a precision of 80% on 25 operations from 20 industrial APIs. It also identified 48% of errors systematically seeded in the outputs of the APIs under test. AGORA+ uncovered 32 bugs in popular APIs, including Amadeus, Deutschebahn, GitHub, Marvel, NYTimesBooks, and YouTube, leading to fixes and documentation updates. Juan C. Alonso, Michael D. Ernst, Sergio Segura, Antonio Ruiz Cortés |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2026 | More Code, Less Understanding? On the Impact of AI Assistants on Developers' Productivity and Code OwnershipabstractArtificial Intelligence (AI) is transforming many domains, including software engineering. AI-based tools are gaining popularity and are increasingly being integrated into software development workflows, automating complex tasks such as code writing and reviewing. When it comes to coding tasks, some evidence suggests that tools like Copilot boost developers’ productivity (e.g.,developers can handle a larger number of pull requests per week). However, it remains unclear whether this comes at the expense of code ownership (i.e.,the developer’ ability to argue about their implementation choices). To partially address this gap, we present an experiment aimed at investigating the impact of AI-based assistants on developers’ productivity and behavior in the context of code writing (e.g.,developing a program from scratch or evolving an existing code). Our focus is on the interplay between productivity and code ownership. We asked 69 participants (34 BSc and 13 MSc students, 8 researchers, and 14 professional developers) to perform two code writing tasks, one with the support of AI and one without. Then, we compared the two treatments in terms of: (i) time spent on the coding task and percentage of the task completeness—both being productivity proxies; and (ii) ability of the participants to answer questions about the code they implemented—code ownership proxy. While the time-based analyses did not provide strong evidence on the impact of AI on the investigated dependent variables, participants using AI achieved a much higher task completeness (>2× in terms of median), confirming a positive impact on their productivity. However, such a boost did not come for free. Indeed, we also observed a loss in code ownership when participants used AI, with lower ability to answer technical questions (–12.5%). Alberto Martin-Lopez, Rosalia Tufano, Emanuela Guglielmi, Ana Belén Sánchez, Ana Lourdes Sanz, Simone Scalabrino, Rocco Oliveto, Sergio Segura, Gabriele Bavota |
IEEE Trans. Software Eng. | 8 |
| 2025 | ASTRAL: Automated Safety Testing of Large Language ModelsabstractLarge Language Models (LLMs) have recently gained significant attention due to their ability to understand and generate sophisticated human-like content. However, ensuring their safety is paramount as they might provide harmful and unsafe responses. Existing LLM testing frameworks address various safety-related concerns (e.g., drugs, terrorism, animal abuse) but often face challenges due to unbalanced and obsolete datasets. In this paper, we present ASTRAL, a tool that automates the generation and execution of test cases (i.e., prompts) for testing the safety of LLMs. First, we introduce a novel black-box coverage criterion to generate balanced and diverse unsafe test inputs across a diverse set of safety categories as well as linguistic writing characteristics (i.e., different style and persuasive writing techniques). Second, we propose an LLM-based approach that leverages Retrieval Augmented Generation (RAG), few-shot prompting strategies and web browsing to generate up-to-date test inputs. Lastly, similar to current LLM test automation techniques, we leverage LLMs as test oracles to distinguish between safe and unsafe test outputs, allowing a fully automated testing approach. We conduct an extensive evaluation on well-known LLMs, revealing the following key findings: i) GPT3.5 outperforms other LLMs when acting as the test oracle, accurately detecting unsafe responses, and even surpassing more recent LLMs (e.g., GPT-4), as well as LLMs that are specifically tailored to detect unsafe LLM outputs (e.g., LlamaGuard); ii) the results confirm that our approach can uncover nearly twice as many unsafe LLM behaviors with the same number of test inputs compared to currently used static datasets; and iii) our black-box coverage criterion combined with web browsing can effectively guide the LLM on generating up-to-date unsafe test inputs, significantly increasing the number of unsafe LLM behaviors. Miriam Ugarte Querejeta, José Antonio Parejo, Sergio Segura, Aitor Arrieta |
AST | 4 |
| 2025 | SATORI: Static Test Oracle Generation for REST APIsabstractREST API test case generation tools are evolving rapidly, with growing capabilities for the automated generation of complex tests. However, despite their strengths in test data generation, these tools are constrained by the types of test oracles they support, often limited to crashes, regressions, and noncompliance with API specifications or design standards. This paper introduces SATORI (Static API Test ORacle Inference), a black-box approach for generating test oracles for REST APIs by analyzing their OpenAPI Specification. SATORI uses large language models to infer the expected behavior of an API by analyzing the properties of the response fields of its operations, such as their name and descriptions. To foster its adoption, we extended the PostmanAssertify tool to automatically convert the test oracles reported by SATORI into executable assertions. Evaluation results on 17 operations from 12 industrial APIs show that SATORI can automatically generate up to hundreds of valid test oracles per operation. SATORI achieved an F1-score of 74.3%, outperforming the state-of-the-art dynamic approach AGORA+(69.3%)—which requires executing the API—when generating comparable oracle types. Moreover, our findings show that static and dynamic oracle inference methods are complementary: together, SATORI and AGORA+ found 90% of the oracles in our annotated ground-truth dataset. Notably, SATORI uncovered 18 bugs in popular APIs (Amadeus Hotel, Deutschebahn, FDIC, GitLab, Marvel, OMDb and Vimeo) leading to documentation updates by the API maintainers. Juan C. Alonso, Alberto Martin-Lopez, Sergio Segura, Gabriele Bavota, Antonio Ruiz Cortés |
ASE | 3 |
| 2024 | Mutation Testing in Practice: Insights From Open-Source Software DevelopersabstractMutation testing drives the creation and improvement of test cases by evaluating their ability to identify synthetic faults. Over the past decades, the technique has gained popularity in academic circles. In practice, however, little is known about its adoption and use. While there are some pilot studies applying mutation testing in industry, the overall usage of mutation testing among developers remains largely unexplored. To fill this gap, this paper presents the results of a qualitative study among open-source developers on the use of mutation testing. Specifically, we report the results of a survey of 104 contributors to open-source projects using a variety of mutation testing tools. The findings of our study provide helpful insights into the use of mutation testing in practice, including its main benefits and limitations. Overall, we observe a high degree of satisfaction with mutation testing across different programming languages and mutation testing tools. Developers find the technique helpful for improving the quality of test suites, detecting bugs, and improving code maintainability. Popularity, usability, and configurability emerge as key factors for the adoption of mutation tools, whereas performance stands overwhelmingly as their main limitation. These results lay the groundwork for new research contributions and tools that meet the needs of developers and boost the widespread adoption of mutation testing. Ana Belén Sánchez, José Antonio Parejo, Sergio Segura, Amador Durán Toro, Mike Papadakis |
IEEE Trans. Software Eng. | 3 |
| 2023 | IDLGen: Automated Code Generation for Inter-parameter Dependencies in Web APIs
Saman A. Barakat, Ana Belén Sánchez, Sergio Segura |
ICSOC (1) | 3 |
| 2023 | AGORA: Automated Generation of Test Oracles for REST APIsabstractTest case generation tools for REST APIs have grown in number and complexity in recent years. However, their advanced capabilities for automated input generation contrast with the simplicity of their test oracles, which limit the types of failures they can detect to crashes, regressions, and violations of the API specification or design best practices. In this paper, we present AGORA, an approach for the automated generation of test oracles for REST APIs through the detection of invariants—properties of the output that should always hold. In practice, AGORA aims to learn the expected behavior of an API by analyzing previous API requests and their corresponding responses. For this, we extended the Daikon tool for dynamic detection of likely invariants, including the definition of new types of invariants and the implementation of an instrumenter called Beet. Beet converts any OpenAPI specification and a collection of API requests and responses to a format processable by Daikon. As a result, AGORA currently supports the detection of up to 105 different types of invariants in REST APIs. AGORA achieved a total precision of 81.2% when tested on a dataset of 11 operations from 7 industrial APIs. More importantly, the test oracles generated by AGORA detected 6 out of every 10 errors systematically seeded in the outputs of the APIs under test. Additionally, AGORA revealed 11 bugs in APIs with millions of users: Amadeus, GitHub, Marvel, OMDb and YouTube. Our reports have guided developers in improving their APIs, including bug fixes and documentation updates in GitHub. Since it operates in black-box mode, AGORA can be seamlessly integrated into existing API testing tools. Juan C. Alonso, Sergio Segura, Antonio Ruiz Cortés |
ISSTA | 2 |
| 2023 | Performance-Driven Metamorphic Testing of Cyber-Physical SystemsabstractCyber-physical systems(CPSs) are a new generation of systems, which integrate software with physical processes. The increasing complexity of these systems, combined with the uncertainty in their interactions with the physical world, makes the definition of effective test oracles especially challenging, facing the well-knowntest oracle problem. Metamorphic testing has shown great potential to alleviate the test oracle problem by exploiting the relations among the inputs and outputs of different executions of the system, so-calledmetamorphic relations(MRs). In this article, we propose an MR pattern called PV for the identification of performance-driven MRs, and we show its applicability in two CPSs from different domains, which are automated navigation systems and elevator control systems. For the evaluation, we assessed the effectiveness of this approach for detecting failures in an open-source simulation-based autonomous navigation system, as well as in an industrial case study from the elevation domain. We derive concrete MRs based on the PV pattern for both case studies, and we evaluate their effectiveness with seeded faults. Results show that the approach is effective at detecting over 88% of the seeded faults, while keeping the ratio of FPs at 4% or lower. Jon Ayerdi, Sergio Segura, Aitor Arrieta, Goiuria Sagardui Mendieta, Maite Arratibel |
IEEE Trans. Reliab. | 3 |
| 2023 | ARTE: Automated Generation of Realistic Test Inputs for Web APIsabstractAutomated test case generation for web APIs is a thriving research topic, where test cases are frequently derived from the API specification. However, this process is only partially automated since testers are usually obliged to manually set meaningful valid test inputs for each input parameter. In this article, we present ARTE, an approach for the automated extraction of realistic test data for web APIs from knowledge bases like DBpedia. Specifically, ARTE leverages the specification of the API parameters to automatically search for realistic test inputs using natural language processing, search-based, and knowledge extraction techniques. ARTE has been integrated into RESTest, an open-source testing framework for RESTful APIs, fully automating the test case generation process. Evaluation results on 140 operations from 48 real-world web APIs show that ARTE can efficiently generate realistic test inputs for 64.9% of the target parameters, outperforming the state-of-the-art approach SAIGEN (31.8%). More importantly, ARTE supported the generation of over twice as many valid API calls (57.3%) as random generation (20%) and SAIGEN (26%), leading to a higher failure detection capability and uncovering several real-world bugs. These results show the potential of ARTE for enhancing existing web API testing tools, achieving an unprecedented level of automation. Juan C. Alonso, Alberto Martin-Lopez, Sergio Segura, José María García, Antonio Ruiz Cortés |
IEEE Trans. Software Eng. | 3 |
| 2022 | Online testing of RESTful APIs: promises and challengesabstractOnline testing of web APIs—testing APIs in production—is gaining traction in industry. Platforms such as RapidAPI and Sauce Labs provide online testing and monitoring services of web APIs 24/7, typically by re-executing manually designed test cases on the target APIs on a regular basis. In parallel, research on the automated generation of test cases for RESTful APIs has seen significant advances in recent years. However, despite their promising results in the lab, it is unclear whether research tools would scale to industrial-size settings and, more importantly, how they would perform in an online testing setup, increasingly common in practice. In this paper, we report the results of an empirical study on the use of automated test case generation methods for online testing of RESTful APIs. Specifically, we used the RESTest framework to automatically generate and execute test cases in 13 industrial APIs for 15 days non-stop, resulting in over one million test cases. To scale at this level, we had to transition from a monolithic tool approach to a multi-bot architecture with over 200 bots working cooperatively in tasks like test generation and reporting. As a result, we uncovered about 390K failures, which we conservatively triaged into 254 bugs, 65 of which have been acknowledged or fixed by developers to date. Among others, we identified confirmed faults in the APIs of Amadeus, Foursquare, Yelp, and YouTube, accessed by millions of applications worldwide. More importantly, our reports have guided developers on improving their APIs, including bug fixes and documentation updates in the APIs of Amadeus and YouTube. Our results show the potential of online testing of RESTful APIs as the next must-have feature in industry, but also some of the key challenges to overcome for its full adoption in practice. Alberto Martin-Lopez, Sergio Segura, Antonio Ruiz Cortés |
ESEC/SIGSOFT FSE | 2 |
| 2022 | Mutation testing in the wild: findings from GitHubabstractAbstract Mutation testing exploits artificial faults to measure the adequacy of test suites and guide their improvement. It has become an extremely popular testing technique as evidenced by the vast literature, numerous tools, and research events on the topic. Previous survey papers have successfully compiled the state of research, its evolution, problems, and challenges. However, the use of mutation testing in practice is still largely unexplored. In this paper, we report the results of a thorough study on the use of mutation testing in GitHub projects. Specifically, we first performed a search for mutation testing tools, 127 in total, and we automatically searched the GitHub repositories including evidence of their use. Then, we focused on the top ten most widely used tools, based on the previous results, and manually revised and classified over 3.5K GitHub active repositories importing them. Among other findings, we observed a recent upturn in interest and activity, with Infection (PHP), PIT (Java) and Humbug (PHP) being the most widely used mutation tools in recent years. The predominant use of mutation testing is development, followed by teaching and learning, and research projects, although with significant differences among mutation tools found in the literature—less adopted and largely used in teaching and research—and those found in GitHub only—more popular and more widely used in development. Our work provides a new and encouraging perspective on the state of practice of mutation testing. Ana Belén Sánchez, Pedro Delgado-Pérez, Inmaculada Medina-Bulo, Sergio Segura |
Empir. Softw. Eng. | 4 |
| 2022 | Specification and Automated Analysis of Inter-Parameter Dependencies in Web APIsabstractWeb services often impose inter-parameter dependencies that restrict the way in which two or more input parameters can be combined to form valid calls to the service. Unfortunately, current specification languages for web services like the OpenAPI Specification (OAS) provide no support for the formal description of such dependencies, which makes it hardly possible to automatically discover and interact with services without human intervention. In this article, we present an approach for the specification and automated analysis of inter-parameter dependencies in web APIs. We first present a domain-specific language, calledInter-parameter Dependency Language(IDL), for the specification of dependencies among input parameters in web services. Then, we propose a mapping to translate an IDL document into a constraint satisfaction problem (CSP), enabling the automated analysis of IDL specifications using standard CSP-based reasoning operations. Specifically, we present a catalogue of seven analysis operations on IDL documents allowing to compute, for example, whether a given request satisfies all the dependencies of the service. Finally, we present a tool suite including an editor, a parser, an OAS extension, a constraint programming-aided library, and a test suite supporting IDL specifications and their analyses. Together, these contributions pave the way for a new range of specification-driven applications in areas such as code generation and testing. Alberto Martin-Lopez, Sergio Segura, Carlos Müller, Antonio Ruiz Cortés |
IEEE Trans. Serv. Comput. | 2 |
| 2021 | Black-Box and White-Box Test Case Generation for RESTful APIs: Enemies or Allies?abstractAutomated test case generation for RESTful APIs is a thriving research topic due to their critical role in software integration. Testing approaches can be divided into black-box and white-box. Black-box approaches exploit the API specification for the generation of test cases, while white-box approaches can also leverage the source code. Both strategies have shown great promise, but they have not been fully compared yet, hindering the selection of the right tool for the job. In this paper, we report on our experience comparing black-box and white-box test case generation for RESTful APIs using the state-of-the-art tools RESTest (black-box) and EvoMaster (white-box). Also, we propose integrating both approaches by using black-box test cases as the seed for white-box search-based test case generation. Evaluation results on four RESTful APIs involving over 40 million API calls show that there is no one-size-fits-all strategy. More importantly, the combination of black-box and white- box yielded the best results in most case studies in terms of code coverage and fault finding, paving the way for better tools integrating the best of both perspectives. As a result of our work, we provide lessons learned and open challenges for guiding the use and further development of current tool support. Alberto Martin-Lopez, Andrea Arcuri, Sergio Segura, Antonio Ruiz Cortés |
ISSRE | 3 |
| 2021 | RESTest: automated black-box testing of RESTful web APIsabstractTesting RESTful APIs thoroughly is critical due to their key role in software integration. Existing tools for the automated generation of test cases in this domain have shown great promise, but their applicability is limited as they mostly rely on random inputs, i.e., fuzzing. In this paper, we present RESTest, an open source black-box testing framework for RESTful web APIs. Based on the API specification, RESTest supports the generation of test cases using different testing techniques such as fuzzing and constraint-based testing, among others. RESTest is developed as a framework and can be easily extended with new test case generators and test writers for different programming languages. We evaluate the tool in two scenarios: offline and online testing. In the former, we show how RESTest can efficiently generate realistic test cases (test inputs and test oracles) that uncover bugs in real-world APIs. In the latter, we show RESTest's capabilities as a continuous testing and monitoring framework. Demo video: https://youtu.be/1f_tjdkaCKo. Alberto Martin-Lopez, Sergio Segura, Antonio Ruiz Cortés |
ISSTA | 2 |
| 2021 | Performance mutation testingabstractSummary Performance bugs are known to be a major threat to the success of software products. Performance tests aim to detect performance bugs by executing the program through test cases and checking whether it exhibits a noticeable performance degradation. The principles of mutation testing, a well‐established testing technique for the assessment of test suites through the injection of artificial faults, could be exploited to evaluate and improve the detection power of performance tests. However, the application of mutation testing to assess performance tests, henceforth called performance mutation testing (PMT), is a novel research topic with numerous open challenges. In previous papers, we identified some key challenges related to PMT. In this work, we go a step further and explore the feasibility of applying PMT at the source‐code level in general‐purpose languages. To do so, we revisit concepts associated with classical mutation testing and design seven novel mutation operators to model known bug‐inducing patterns. As a proof of concept, we applied traditional mutation operators as well as performance mutation operators to open‐source C++ programs. The results reveal the potential of the new performance‐mutants to help assess and enhance performance tests when compared with traditional mutants. A review of live mutants in these programs suggests that they can induce the design of special test inputs. In addition to these promising results, our work brings a whole new set of challenges related to PMT, which will hopefully serve as a starting point for new contributions in the area. Pedro Delgado-Pérez, Ana Belén Sánchez, Sergio Segura, Inmaculada Medina-Bulo |
Softw. Test. Verification Reliab. | 3 |
| 2020 | RESTest: Black-Box Constraint-Based Testing of RESTful Web APIs
Alberto Martin-Lopez, Sergio Segura, Antonio Ruiz Cortés |
ICSOC | 2 |
| 2020 | QoS-aware Metamorphic Testing: An Elevation Case StudyabstractElevators are among the oldest and most widespread transportation systems, yet their complexity increases rapidly to satisfy customization demands and to meet quality of service requirements. Verification and validation tasks in this context are costly, since they rely on the manual intervention of domain experts at some points of the process. This is mainly due to the difficulty to assess whether the elevators behave as expected in the different test scenarios, the so-called test oracle problem. Metamorphic testing is a thriving testing technique that alleviates the oracle problem by reasoning on the relations among multiple executions of the system under test, the so-called metamorphic relations. In this practical experience paper, we report on the application of metamorphic testing to verify an industrial elevator dispatcher. Together with domain experts from the elevation sector, we defined multiple metamorphic relations that consider domain-specific quality of service measures. Evaluation results with seeded faults show that the approach is effective at detecting faults automatically. Jon Ayerdi, Sergio Segura, Aitor Arrieta, Goiuria Sagardui Mendieta, Maite Arratibel |
ISSRE | 2 |
| 2020 | Many-Objective Test Suite Generation for Software Product LinesabstractA Software Product Line (SPL) is a set of products built from a number of features, the set of valid products being defined by a feature model. Typically, it does not make sense to test all products defined by an SPL and one instead chooses a set of products to test (test selection) and, ideally, derives a good order in which to test them (test prioritisation). Since one cannot know in advance which products will reveal faults, test selection and prioritisation are normally based on objective functions that are known to relate to likely effectiveness or cost. This article introduces a new technique, the grid-based evolution strategy (GrES), which considers several objective functions that assess a selection or prioritisation and aims to optimise on all of these. The problem is thus a many-objective optimisation problem. We use a new approach, in which all of the objective functions are considered but one (pairwise coverage) is seen as the most important. We also derive a novel evolution strategy based on domain knowledge. The results of the evaluation, on randomly generated and realistic feature models, were promising, with GrES outperforming previously proposed techniques and a range of many-objective optimisation algorithms. Robert M. Hierons, Miqing Li, Xiaohui Liu 0001, José Antonio Parejo, Sergio Segura, Xin Yao 0001 |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2019 | An Extended Abstract of "Metamorphic Testing: Testing the Untestable"abstractThis document is an extended abstract of an IEEE Software paper, "Metamorphic Testing: Testing the Untestable," presented as a J1C2 (Journal publication first, Conference presentation following) at the IEEE Computer Society signature conference on Computers, Software and Applications (COMPSAC 2019), hosted by Marquette University, Milwaukee, Wisconsin, USA. Sergio Segura, Dave Towey, Zhiquan Zhou 0001, Tsong Yueh Chen |
COMPSAC (1) | 1 |
| 2019 | A Catalogue of Inter-parameter Dependencies in RESTful Web APIs
Alberto Martin-Lopez, Sergio Segura, Antonio Ruiz Cortés |
ICSOC | 2 |
| 2018 | Metamorphic testing of RESTful web APIsabstractWeb Application Programming Interfaces (APIs) specify how to access services and data over the network, typically using Web services. Web APIs are rapidly proliferating as a key element to foster reusability, integration, and innovation, enabling new consumption models such as mobile or smart TV apps. Companies such as Facebook, Twitter, Google, eBay or Netflix receive billions of API calls every day from thousands of different third-party applications and devices, which constitutes more than half of their total traffic. Sergio Segura, José Antonio Parejo, Javier Troya, Antonio Ruiz Cortés |
ICSE | 1 |
| 2018 | Spectrum-based fault localization in software product lines
Aitor Arrieta, Sergio Segura, Urtzi Markiegi, Goiuria Sagardui Mendieta, Leire Etxeberria Elorza |
Inf. Softw. Technol. | 2 |
| 2018 | Performance mutation testing: Hypothesis and open questions
Ana Belén Sánchez, Pedro Delgado-Pérez, Sergio Segura, Inmaculada Medina-Bulo |
Inf. Softw. Technol. | 3 |
| 2018 | Performance metamorphic testing: A Proof of concept
Sergio Segura, Javier Troya, Amador Durán Toro, Antonio Ruiz Cortés |
Inf. Softw. Technol. | 1 |
| 2018 | Automated inference of likely metamorphic relations for model transformations
Javier Troya, Sergio Segura, Antonio Ruiz Cortés |
J. Syst. Softw. | 2 |
| 2018 | Spectrum-Based Fault Localization in Model TransformationsabstractModel transformations play a cornerstone role in Model-Driven Engineering (MDE), as they provide the essential mechanisms for manipulating and transforming models. The correctness of software built using MDE techniques greatly relies on the correctness of model transformations. However, it is challenging and error prone to debug them, and the situation gets more critical as the size and complexity of model transformations grow, where manual debugging is no longer possible. Spectrum-Based Fault Localization (SBFL) uses the results of test cases and their corresponding code coverage information to estimate the likelihood of each program component (e.g., statements) of being faulty. In this article we present an approach to apply SBFL for locating the faulty rules in model transformations. We evaluate the feasibility and accuracy of the approach by comparing the effectiveness of 18 different state-of-the-art SBFL techniques at locating faults in model transformations. Evaluation results revealed that the best techniques, namely Kulcynski2 , Mountford , Ochiai , and Zoltar , lead the debugger to inspect a maximum of three rules to locate the bug in around 74% of the cases. Furthermore, we compare our approach with a static approach for fault localization in model transformations, observing a clear superiority of the proposed SBFL-based method. Javier Troya, Sergio Segura, José Antonio Parejo, Antonio Ruiz Cortés |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2018 | Metamorphic Testing of RESTful Web APIsabstractWeb Application Programming Interfaces (APIs) allow systems to interact with each other over the network. Modern Web APIs often adhere to the REST architectural style, being referred to as RESTful Web APIs. RESTful Web APIs are decomposed into multiple resources (e.g., a video in the YouTube API) that clients can manipulate through HTTP interactions. Testing Web APIs is critical but challenging due to the difficulty to assess the correctness of API responses, i.e., the oracle problem. Metamorphic testing alleviates the oracle problem by exploiting relations (so-called metamorphic relations) among multiple executions of the program under test. In this paper, we present a metamorphic testing approach for the detection of faults in RESTful Web APIs. We first propose six abstract relations that capture the shape of many of the metamorphic relations found in RESTful Web APIs, we call these Metamorphic Relation Output Patterns (MROPs). Each MROP can then be instantiated into one or more concrete metamorphic relations. The approach was evaluated using both automatically seeded and real faults in six subject Web APIs. Among other results, we identified 60 metamorphic relations (instances of the proposed MROPs) in the Web APIs of Spotify and YouTube. Each metamorphic relation was implemented using both random and manual test data, running over 4.7K automated tests. As a result, 11 issues were detected (3 in Spotify and 8 in YouTube), 10 of them confirmed by the API developers or reproduced by other users, supporting the effectiveness of the approach. Sergio Segura, José Antonio Parejo, Javier Troya, Antonio Ruiz Cortés |
IEEE Trans. Software Eng. | 1 |
| 2017 | Evolutionary composition of QoS-aware web services: A many-objective perspective
Aurora Ramírez 0001, José Antonio Parejo, José Raúl Romero, Sergio Segura, Antonio Ruiz Cortés |
Expert Syst. Appl. | 4 |
| 2017 | FLAME: a formal framework for the automated analysis of software product lines validated by automated specification testing
Amador Durán Toro, David Benavides 0001, Sergio Segura, Pablo Trinidad Martín-Arroyo, Antonio Ruiz Cortés |
Softw. Syst. Model. | 3 |
| 2017 | Variability testing in the wild: the Drupal case study
Ana Belén Sánchez, Sergio Segura, José Antonio Parejo, Antonio Ruiz Cortés |
Softw. Syst. Model. | 2 |
| 2017 | Assessment of C++ object-oriented mutation operators: A selective mutation approachabstractSummary Mutation testing is an effective but costly testing technique. Several studies have observed that some mutants can be redundant and therefore removed without affecting its effectiveness. Similarly, some mutants may be more effective than others in guiding the tester on the creation of high‐quality test cases. On the basis of these findings, we present an assessment of C++ class mutation operators by classifying them into 2 rankings: the first ranking sorts the operators on the basis of their degree of redundancy and the second regarding the quality of the tests they help to design. Both rankings are used in a selective mutation study analysing the trade‐off between the reduction achieved and the effectiveness when using a subset of mutants. Experimental results consistently show that leveraging the operators at the top of the 2 rankings, which are different, lead to a significant reduction in the number of mutants with a minimum loss of effectiveness. Pedro Delgado-Pérez, Sergio Segura, Inmaculada Medina-Bulo |
Softw. Test. Verification Reliab. | 2 |
| 2016 | Multi-objective test case prioritization in highly configurable systems: A case study
José Antonio Parejo, Ana Belén Sánchez, Sergio Segura, Antonio Ruiz Cortés, Roberto Erick Lopez-Herrejon, Alexander Egyed |
J. Syst. Softw. | 3 |
| 2016 | SIP: Optimal Product Selection from Feature Models Using Many-Objective Evolutionary OptimizationabstractA feature model specifies the sets of features that define valid products in a software product line. Recent work has considered the problem of choosing optimal products from a feature model based on a set of user preferences, with this being represented as a many-objective optimization problem. This problem has been found to be difficult for a purely search-based approach, leading to classical many-objective optimization algorithms being enhanced either by adding in a valid product as a seed or by introducing additional mutation and replacement operators that use an SAT solver. In this article, we instead enhance the search in two ways: by providing a novel representation and by optimizing first on the number of constraints that hold and only then on the other objectives. In the evaluation, we also used feature models with realistic attributes, in contrast to previous work that used randomly generated attribute values. The results of experiments were promising, with the proposed (SIP) method returning valid products with six published feature models and a randomly generated feature model with 10,000 features. For the model with 10,000 features, the search took only a few minutes. Robert M. Hierons, Miqing Li, Xiaohui Liu 0001, Sergio Segura, Wei Zheng 0006 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2016 | A Survey on Metamorphic TestingabstractA test oracle determines whether a test execution reveals a fault, often by comparing the observed program output to the expected output. This is not always practical, for example when a program's input-output relation is complex and difficult to capture formally. Metamorphic testing provides an alternative, where correctness is not determined by checking an individual concrete output, but by applying a transformation to a test input and observing how the program output “morphs” into a different one as a result. Since the introduction of such metamorphic relations in 1998, many contributions on metamorphic testing have been made, and the technique has seen successful applications in a variety of domains, ranging from web services to computer graphics. This article provides a comprehensive survey on metamorphic testing: It summarises the research results and application areas, and analyses common practice in empirical studies of metamorphic testing as well as the main open challenges. Sergio Segura, Gordon Fraser 0001, Ana Belén Sánchez, Antonio Ruiz Cortés |
IEEE Trans. Software Eng. | 1 |
| 2015 | An assessment of search-based techniques for reverse engineering feature models
Roberto Erick Lopez-Herrejon, Lukas Linsbauer, José A. Galindo, José Antonio Parejo, David Benavides 0001, Sergio Segura, Alexander Egyed |
J. Syst. Softw. | 6 |
| 2015 | Automated metamorphic testing of variability analysis toolsabstractSummary Variability determines the capability of software applications to be configured and customized. A common need during the development of variability‐intensive systems is the automated analysis of their underlying variability models, for example, detecting contradictory configuration options. The analysis operations that are performed on variability models are often very complex, which hinders the testing of the corresponding analysis tools and makes difficult, often infeasible, to determine the correctness of their outputs, that is, the well‐knownoracle problemin software testing. In this article, we present a generic approach for the automated detection of faults in variability analysis tools overcoming the oracle problem. Our work enables the generation of random variability models together with the exact set of valid configurations represented by these models. These test data are generated from scratch using stepwise transformations and assuring that certain constraints (a.k.a.metamorphic relations) hold at each step. To show the feasibility and generalizability of our approach, it has been used to automatically test several analysis tools in three variability domains: feature models, common upgradeability description format documents and Boolean formulas. Among other results, we detected 19 real bugs in 7 out of the 15 tools under test. Copyright © 2015 John Wiley & Sons, Ltd. Sergio Segura, Amador Durán Toro, Ana Belén Sánchez, Daniel Le Berre, Emmanuel Lonca, Antonio Ruiz Cortés |
Softw. Test. Verification Reliab. | 1 |
| 2014 | A Comparison of Test Case Prioritization Criteria for Software Product LinesabstractSoftware Product Line (SPL) testing is challenging due to the potentially huge number of derivable products. To alleviate this problem, numerous contributions have been proposed to reduce the number of products to be tested while still having a good coverage. However, not much attention has been paid to the order in which the products are tested. Test case prioritization techniques reorder test cases to meet a certain performance goal. For instance, testers may wish to order their test cases in order to detect faults as soon as possible, which would translate in faster feedback and earlier fault correction. In this paper, we explore the applicability of test case prioritization techniques to SPL testing. We propose five different prioritization criteria based on common metrics of feature models and we compare their effectiveness in increasing the rate of early fault detection, i.e. a measure of how quickly faults are detected. The results show that different orderings of the same SPL suite may lead to significant differences in the rate of early fault detection. They also show that our approach may contribute to accelerate the detection of faults of SPL test suites based on combinatorial testing. Ana Belén Sánchez, Sergio Segura, Antonio Ruiz Cortés |
ICST | 2 |
| 2014 | Automated variability analysis and testing of an E-commerce site.: an experience reportabstractIn this paper, we report on our experience on the development of La Hilandera, an e-commerce site selling haberdashery products and craft supplies in Europe. The store has a huge input space where customers can place almost three millions of different orders which made testing an extremely difficult task. To address the challenge, we explored the applicability of some of the practices for variability management in software product lines. First, we used a feature model to represent the store input space which provided us with a variability view easy to understand, share and discuss with all the stakeholders. Second, we used techniques for the automated analysis of feature models for the detection and repair of inconsistent and missing configuration settings. Finally, we used test selection and prioritization techniques for the generation of a manageable and effective set of test cases. Our findings, summarized in a set of lessons learnt, suggest that variability techniques could successfully address many of the challenges found when developing e-commerce sites. Sergio Segura, Ana Belén Sánchez, Antonio Ruiz Cortés |
ASE | 1 |
| 2014 | QoS-aware web services composition using GRASP with Path Relinking
José Antonio Parejo, Sergio Segura, Pablo Fernandez 0001, Antonio Ruiz Cortés |
Expert Syst. Appl. | 2 |
| 2014 | Automated generation of computationally hard feature models using evolutionary algorithms
Sergio Segura, José Antonio Parejo, Robert M. Hierons, David Benavides 0001, Antonio Ruiz Cortés |
Expert Syst. Appl. | 1 |
| 2012 | Reverse Engineering Feature Models with Evolutionary Algorithms: An Exploratory Study
Roberto Erick Lopez-Herrejon, José A. Galindo, David Benavides 0001, Sergio Segura, Alexander Egyed |
SSBSE | 4 |
| 2011 | Automated metamorphic testing on the analyses of feature models
Sergio Segura, Robert M. Hierons, David Benavides 0001, Antonio Ruiz Cortés |
Inf. Softw. Technol. | 1 |
| 2011 | Mutation testing on an object-oriented framework: An experience report
Sergio Segura, Robert M. Hierons, David Benavides 0001, Antonio Ruiz Cortés |
Inf. Softw. Technol. | 1 |
| 2010 | Automated Test Data Generation on the Analyses of Feature Models: A Metamorphic Testing ApproachabstractA Feature Model (FM) is a compact representation of all the products of a software product line. The automated extraction of information from FMs is a thriving research topic involving a number of analysis operations, algorithms, paradigms and tools. Implementing these operations is far from trivial and easily leads to errors and defects in analysis solutions. Current testing methods in this context mainly rely on the ability of the tester to decide whether the output of an analysis is correct. However, this is acknowledged to be time-consuming, error-prone and in most cases infeasible due to the combinatorial complexity of the analyses. In this paper, we present a set of relations (so-called metamorphic relations) between input FMs and their set of products and a test data generator relying on them. Given an FM and its known set of products, a set of neighbour FMs together with their corresponding set of products are automatically generated and used for testing different analyses. Complex FMs representing millions of products can be efficiently created applying this process iteratively. The evaluation of our approach using mutation testing as well as real faults and tools reveals that most faults can be automatically detected within a few seconds. Sergio Segura, Robert M. Hierons, David Benavides 0001, Antonio Ruiz Cortés |
ICST | 1 |
| 2010 | Automated analysis of feature models 20 years later: A literature review
David Benavides 0001, Sergio Segura, Antonio Ruiz Cortés |
Inf. Syst. | 2 |
| 2008 | FAMA FrameworkabstractFAMA Framework (FAMA FW) is a tool for the automated analysis of variability models (VM). Its main objective is providing an extensible framework where current research on VM automated analysis might be developed and easily integrated into a final product. FAMA FW is built following the SPL paradigm supporting different variability metamodels, reasoners or solvers, analysis questions and reasoner selectors, easing the production of customized VM analysis tools. FAMA FW is written in Java and distributed under LGPL License. Pablo Trinidad Martín-Arroyo, David Benavides 0001, Antonio Ruiz Cortés, Sergio Segura, Alberto Jimenez |
SPLC | 4 |