Pablo C. Cañizares

dblp:161/0410 · DBLP profile ↗
← Back
22ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0002-2084-1558ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 15 · 7 first-author · 12 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards metamorphic testing with LLM-based workflows: Metamorphic relation inference and follow-up test case generation
abstract
Context: Metamorphic testing (MT) is a well-known approach to address the oracle problem in software testing. It does so by using expected relations – called metamorphic relations (MRs) – between the inputs and outputs of multiple system executions as test oracles. This approach involves identifying meaningful MRs for the domain, defining them in an executable language, and creating suitable system inputs to check their satisfaction. However, these tasks are manual and demand substantial effort from testers. Objective: To reduce the effort of applying MT, we propose automating some tasks in the MT process with intelligent assistants, focussing on MR inference and generation of follow-up test cases. Methods: We propose a taxonomy of MT tasks amenable to LLM-based assistance, and an extensible architecture for their automation. The MT tasks are accessible via a conversational assistant, and may encompass validation and fixing cycles to improve the quality of the assistance. We have integrated the architecture with the MT framework Gotten . Results: We have evaluated our approach on the two considered MT tasks across three domains (data centres, autonomous vehicles, finite automata). We found the assistant effective at generating new MRs (between 2 and 6 depending on the domain) and follow-up test cases (with over 99.6% correctness). We also measured diversity and redundancy of the generated MRs, the benefits of fixing cycles on different LLMs when generating follow-up test cases, and performance degradation and computational cost as input complexity increases. Conclusions: Our architecture effectively automates two central MT tasks, namely, inferring MRs and generating follow-up test cases, showing potential to reduce the effort from MT practitioners and to automate other MT tasks.
Pablo C. Cañizares, Pablo Gómez-Abajo, Esther Guerra, Juan de Lara
Inf. Softw. Technol.1
2026 Multi-objective optimization of cloud systems
abstract
Currently, enormous amounts of data are continuously processed to support our daily activities, such as managing bank accounts, streaming movies, or interacting on social networks. In recent years, cloud infrastructures have proven to be a reliable solution, not only for processing this data but also for enabling users worldwide to access it remotely. However, this processing demands vast computing resources, leading to significant energy consumption. In this paper, we present a strategy to address this problem by combining multi-objective optimization techniques with Metamorphic Testing (MT) and simulation tools to optimize cloud systems, focusing on both performance and energy consumption. To achieve this, several multi-objective genetic algorithms (MOGAs) have been integrated into the MT-EA4Cloud framework, a solution that previously applied single-objective evolutionary algorithms with MT. To determine the suitability of the proposed approach, an empirical study was conducted to analyze the behavior of the different MOGAs included in the framework. In this study, various test sets and two distinct workloads – inspired by big data analytics operations – were created to represent multiple cloud scenarios. The results clearly demonstrate that MOGAs can be effectively combined with MT to optimize cloud systems while considering multiple objectives – in this case, performance and energy consumption. A careful analysis of the results indicates that increasing the mutation rate leads to the best outcomes. In general, the NSGA-II algorithm has produced the best results in the experiments conducted in this study.
Miguel Pérez 0002, Pablo C. Cañizares, Alberto Nuñez
Sci. Comput. Program.2
2025 A language-parametric test amplification framework for executable domain-specific languages
Faezeh Khorram, Erwan Bousse, Jean-Marie Mottu, Gerson Sunyé, Djamel Eddine Khelladi, Pablo Gómez-Abajo, Pablo C. Cañizares, Esther Guerra, Juan de Lara
Softw. Syst. Model.7
2024 Coverage-based Strategies for the Automated Synthesis of Test Scenarios for Conversational Agents
abstract
Conversational agents - or chatbots - are increasingly used as the user interface to many software services. While open-domain chatbots like ChatGPT excel in their ability to chat about any topic, task-oriented conversational agents are designed to perform goal-oriented tasks (e.g., booking or shopping) guided by a dialogue-based user interaction, which is explicitly designed. Like any kind of software system, task-oriented conversational agents need to be properly tested to ensure their quality. For this purpose, some tools permit defining and executing conversation test cases. However, there are currently no established means to assess the coverage of the design of a task-oriented agent by a test suite, or mechanisms to automate quality test case generation ensuring the agent coverage.
Pablo C. Cañizares, Romulo Daniel Avila Ortiz, Sara Pérez-Soler, Esther Guerra, Juan de Lara
AST1
2024 Mutation Testing for Task-Oriented Chatbots
abstract
Conversational agents, or chatbots, are increasingly used to access all sorts of services using natural language. While open-domain chatbots – like ChatGPT – can converse on any topic, task-oriented chatbots – the focus of this paper – are designed for specific tasks, like booking a flight, obtaining customer support, or setting an appointment. Like any other software, task-oriented chatbots need to be properly tested, usually by defining and executing test scenarios (i.e., sequences of user-chatbot interactions). However, there is currently a lack of methods to quantify the completeness and strength of such test scenarios, which can lead to low-quality tests, and hence to buggy chatbots.
Pablo Gómez-Abajo, Sara Pérez-Soler, Pablo C. Cañizares, Esther Guerra, Juan de Lara
EASE3
2024 Measuring and Clustering Heterogeneous Chatbot Designs
abstract
Conversational agents, or chatbots, have become popular to access all kind of software services. They provide an intuitive natural language interface for interaction, available from a wide range of channels including social networks, web pages, intelligent speakers or cars. In response to this demand, many chatbot development platforms and tools have emerged. However, they typically lack support to statically measure properties of the chatbots being built, as indicators of their size, complexity, quality or usability. Similarly, there are hardly any mechanisms to compare and cluster chatbots developed with heterogeneous technologies. To overcome this limitation, we propose a suite of 21 metrics for chatbot designs, as well as two clustering methods that help in grouping chatbots along their conversation topics and design features. Both the metrics and the clustering methods are defined on a neutral chatbot design language, becoming independent of the implementation platform. We provide automatic translations of chatbots defined on some major platforms into this neutral notation to perform the measurement and clustering. The approach is supported by our tool Asymob , which we have used to evaluate the metrics and the clustering methods over a set of 259 Dialogflow and Rasa chatbots from open-source repositories. The results open the door to incorporating the metrics within chatbot development processes for the early detection of quality issues, and to exploit clustering to organise large collections of chatbots into significant groups to ease chatbot comprehension, search and comparison.
Pablo C. Cañizares, Jose María López-Morales, Sara Pérez-Soler, Esther Guerra, Juan de Lara
ACM Trans. Softw. Eng. Methodol.1
2023 Automated engineering of domain-specific metamorphic testing environments
abstract
Testing is essential to improve the correctness of software systems. Metamorphic testing (MT) is an approach especially suited when the system under test lacks oracles, or they are expensive to compute. However, building an MT environment for a particular domain (e.g., cloud simulation, model transformation, machine learning) requires substantial effort. Our goal is to facilitate the construction of MT environments for specific domains. We propose a model-driven engineering approach to automate the construction of MT environments. Starting from a meta-model capturing the domain concepts, and a description of the domain execution environment, our approach produces an MT environment featuring comprehensive support for the MT process. This includes the definition of domain-specific metamorphic relations, their evaluation, detailed reporting of the testing results, and the automated search-based generation of follow-up test cases. Our method is supported by an extensible platform for Eclipse, called Gotten. We demonstrate its effectiveness by creating an MT environment for simulation-based testing of data centres and comparing with existing tools; its suitability to conduct MT processes by replicating previous experiments; and its generality by building another MT environment for video streaming APIs. Gotten is the first platform targeted at reducing the development effort of domain-specific MT environments. The environments created with Gotten facilitate the specification of metamorphic relations, their evaluation, and the generation of new test cases.
Pablo Gómez-Abajo, Pablo C. Cañizares, Alberto Nuñez, Esther Guerra, Juan de Lara
Inf. Softw. Technol.2
2022 Automatic test amplification for executable models
abstract
Behavioral models are important assets that must be thoroughly verified early in the design process. This can be achieved with manually-written test cases that embed carefully hand-picked domain-specific input data. However, such test cases may not always reach the desired level of quality, such as high coverage or being able to localize faults efficiently. Test amplification is an interesting emergent approach to improve a test suite by automatically generating new test cases out of existing manually-written ones. Yet, while ad-hoc test amplification solutions have been proposed for a few programming languages, no solution currently exists for amplifying the test cases of behavioral models.
Faezeh Khorram, Erwan Bousse, Jean-Marie Mottu, Gerson Sunyé, Pablo Gómez-Abajo, Pablo C. Cañizares, Esther Guerra, Juan de Lara
MoDELS6
2022 SINPA: SupportINg the automation of construction PlAnning
Pablo C. Cañizares, Sonia Estévez Martín, Manuel Núñez 0001
Expert Syst. Appl.1
2022 CloudExpert: An intelligent system for selecting cloud system simulators
Alberto Nuñez, Pablo C. Cañizares, Juan de Lara
Expert Syst. Appl.2
2022 Chaos as a Software Product Line - A platform for improving open hybrid-cloud systems resiliency
abstract
Abstract Nowadays, cloud‐native software architectures have a significant relevance due to the speed and agility they provide. These properties lead relevant organizations in different industries, like video streaming (Netflix), car‐sharing (Uber, Cabify), banking (BBVA, HSBC), and governmental agencies (NASA, FBI, CERN, ESA) to heavily rely on cloud‐native software to run their business‐critical applications. Additionally, including fault injection actions in the production infrastructure allows companies to have consistent environments, to improve applications dependability against unexpected failures, to provide better user experience, and to improve the overall system quality. Thus, cloud computing technologies allow development teams to rapidly create complex systems and to continuously deploy them, at a global scale. This work describes Pystol, a novel fault injection platform—represented as a Software Product Line—to analyze the effects caused by a wide spectrum of adverse conditions. Pystol is designed to be executed on top of cloud‐native environments, either in private or public clouds. The proposed architecture shows a way for representing feature models based on Unified Model Language (in short, UML) component diagrams. Furthermore, we present a thorough empirical study carried out in real‐world environments, providing promising results.
Carlos Camacho, Pablo C. Cañizares, Luis Llana, Alberto Nuñez
Softw. Pract. Exp.2
2022 Evaluating cloud interactions with costs and SLAs
abstract
Abstract In this paper, we investigate how to improve the profits in cloud infrastructures by using price schemes and analyzing the user interactions with the cloud provider. For this purpose, we consider two different types of client behavior, namely regular and high-priority users. Regular users do not require a continuous service, and they can wait to be attended to. In contrast, high-priority users require a continuous service, e.g., a 24/7 service, and usually need an immediate answer to any request. A complete framework has been implemented, which includes a UML profile that allows us to define specific cloud scenarios and the automatic transformations to produce the code for the cloud simulations in the Simcan2Cloud simulator. The engine of Simcan2Cloud has also been modified by adding specific SLAs and price schemes. Finally, we present a thorough experimental study to analyze the performance results obtained from the simulations, thus making it possible to draw conclusions about how to improve the cloud profit for the cloud studied by adjusting the different parameters and resource configuration.
Adrian Bernal, María-Emilia Cambronero, Alberto Nuñez, Pablo C. Cañizares, Valentín Valero Ruiz
J. Supercomput.4
2021 Studying the Impact of the User Subscription Times in Different Cloud Configurations
abstract
In this paper, we model cloud systems and the user interactions with the cloud provider using the UML2Cloud profile.In general, users request virtual machines according to their needs, but they can also subscribe to the cloud provider and wait to be notified when the requested resources are not available.In this case, users indicate a maximum subscription time, so once this time elapses without being notified, users leave the system unattended.In this paper, then, we present an exhaustive research study to measure how the user subscription times affect the overall system responsiveness.In this study, three different cloud configurations are analyzed.Each cloud processes several workloads, which are generated using two distribution functions for the user arrivals, namely a normal and a cyclic normal distribution.The purpose of this study is to find out the inflection point for the waiting time of the users, from which the cloud responsiveness and its performance do not improve.The obtained information is therefore useful for the cloud provider to improve the configuration of the cloud.
Hernán-Indibil de la Cruz, María-Emilia Cambronero, Valentín Valero Ruiz, Pablo C. Cañizares, Adrian Bernal, Alberto Nuñez
SEKE4
2021 New ideas: automated engineering of metamorphic testing environments for domain-specific languages
abstract
Two crucial aspects for the trustworthy utilization of domain-specific languages (DSLs) are their semantic correctness, and proper testing support for their users. Testing is frequently used to verify correctness, but is often done informally -- which may yield unreliable results -- and requires substantial effort for creating suitable test cases and oracles.
Pablo C. Cañizares, Pablo Gómez-Abajo, Alberto Nuñez, Esther Guerra, Juan de Lara
SLE1
2021 Analyzing the Cloud Performance Using Different User Subscription Times
abstract
Cloud providers face the challenge of managing large amounts of heterogeneous resources in real time. It is usually very costly to conduct experiments with real cloud systems. Therefore, tools to analyze and evaluate cloud scenarios and experimental studies are very useful for them. In this paper, we model cloud systems and the user interactions with the cloud provider using the UML2Cloud profile. In general, users request virtual machines according to their needs, but they can also subscribe to the cloud provider and wait to be notified when the requested resources are not available. In this case, users indicate a maximum subscription time, so once this time elapses without being notified, users leave the system unattended. Thus, we present an exhaustive experimental study to measure how the user subscription times affect the overall system responsiveness. To this end, four different cloud configurations are analyzed, and the workloads for these studies are produced by using three distribution functions for the user arrivals, namely, a uniform, a normal, and a cyclic normal distribution. Furthermore, we also analyze the cloud performance with a workload obtained from a real trace. The purpose of this study is to find out the inflection point for the waiting time of the users, from which the cloud responsiveness and its performance do not improve. The obtained information is, therefore, useful for the cloud provider to improve the configuration of the cloud.
Adrian Bernal, María-Emilia Cambronero, Pablo C. Cañizares, Alberto Nuñez, Valentín Valero Ruiz, Hernán-Indibil de la Cruz
Int. J. Softw. Eng. Knowl. Eng.3
2021 TEA-Cloud: A Formal Framework for Testing Cloud Computing Systems
abstract
The validation of a cloud system can be complicated by the size of the system, the number of users that can concurrently request services, and the virtualization used to give the illusion of using dedicated machines. Unfortunately, it is not feasible to use conventional testing methods with cloud systems. This article proposes a framework, called TEA-Cloud, that integrates simulation with testing methods for validating cloud system designs. Testing is applied on both functional and nonfunctional aspects of the cloud, like performance and cost. The aim of the framework is to provide a complete methodology to help users to model both software and hardware parts of cloud systems and automatically test the validity of these clouds using a cost-effective approach. Metamorphic testing is used to overcome the lack of an oracle that checks whether the behavior observed in testing is allowed. Metamorphic testing is based on metamorphic relations (MRs). We define three families of MRs, which target issues such as performance, resource provisioning, and cost. TEA-Cloud was evaluated through an empirical study that used fault seeding (mutation) and ten MRs for testing different cloud configurations. The results were promising, with TEA-Cloud finding all seeded faults.
Alberto Nuñez, Pablo C. Cañizares, Manuel Núñez 0001, Robert M. Hierons
IEEE Trans. Reliab.2
2020 MT-EA4Cloud: A Methodology For testing and optimising energy-aware cloud systems
Pablo C. Cañizares, Alberto Nuñez, Juan de Lara, Luis Llana
J. Syst. Softw.1
2019 An expert system for checking the correctness of memory systems using simulation and metamorphic testing
Pablo C. Cañizares, Alberto Nuñez, Juan de Lara
Expert Syst. Appl.1
2019 Improving cloud architectures using UML profiles and M2T transformation techniques
Adrian Bernal, María-Emilia Cambronero, Alberto Nuñez, Pablo C. Cañizares, Valentín Valero Ruiz
J. Supercomput.4
2018 Mutomvo: Mutation testing framework for simulated cloud and HPC environments
Pablo C. Cañizares, Alberto Nuñez, Mercedes G. Merayo
J. Syst. Softw.1
2017 MAGICIAN: Model-based design for optimizing the configuration of data-centers
abstract
Designing data-centers that provide an acceptable costperformance ratio is challenging.Generally, a wide spectrum of components must be previously analyzed, such as the kind of applications to be executed in the data-center, computing/storage requirements and the network topology, among others.Since each one of these components has a direct impact on the overall system performance, the design process is complex and difficult, which usually requires the intervention of an expert.We propose a model-based approach to design datacenters.For this purpose, we have created a meta-model that describes the structure of data-center models.Then, a set of expert rules can be used to detect sub-optimal configurations, and (in some cases) correct the design.Datacenter models can be simulated, to assess their performance and scalability, for which we use a code generator into the SIMCAN tool.We have implemented our approach as an Eclipse plugin, and illustrate the usefulness of some expert rules by showing the efficiency and scalability gains of the optimized model with respect to the original one.
Pablo C. Cañizares, Alberto Nuñez, Juan de Lara
SEKE1
2016 FARTHEST: FormAl distRibuTed scHema to dEtect Suspicious arTefacts
Pablo C. Cañizares, Mercedes G. Merayo, Alberto Nuñez
ACIIDS (1)1