Miriam Ugarte Querejeta

dblp:188/4307 · also Miriam Ugarte · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-0395-5131ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 ASTRAL: Automated Safety Testing of Large Language Models
abstract
Large Language Models (LLMs) have recently gained significant attention due to their ability to understand and generate sophisticated human-like content. However, ensuring their safety is paramount as they might provide harmful and unsafe responses. Existing LLM testing frameworks address various safety-related concerns (e.g., drugs, terrorism, animal abuse) but often face challenges due to unbalanced and obsolete datasets. In this paper, we present ASTRAL, a tool that automates the generation and execution of test cases (i.e., prompts) for testing the safety of LLMs. First, we introduce a novel black-box coverage criterion to generate balanced and diverse unsafe test inputs across a diverse set of safety categories as well as linguistic writing characteristics (i.e., different style and persuasive writing techniques). Second, we propose an LLM-based approach that leverages Retrieval Augmented Generation (RAG), few-shot prompting strategies and web browsing to generate up-to-date test inputs. Lastly, similar to current LLM test automation techniques, we leverage LLMs as test oracles to distinguish between safe and unsafe test outputs, allowing a fully automated testing approach. We conduct an extensive evaluation on well-known LLMs, revealing the following key findings: i) GPT3.5 outperforms other LLMs when acting as the test oracle, accurately detecting unsafe responses, and even surpassing more recent LLMs (e.g., GPT-4), as well as LLMs that are specifically tailored to detect unsafe LLM outputs (e.g., LlamaGuard); ii) the results confirm that our approach can uncover nearly twice as many unsafe LLM behaviors with the same number of test inputs compared to currently used static datasets; and iii) our black-box coverage criterion combined with web browsing can effectively guide the LLM on generating up-to-date unsafe test inputs, significantly increasing the number of unsafe LLM behaviors.
Miriam Ugarte Querejeta, José Antonio Parejo, Sergio Segura, Aitor Arrieta
AST1
2025 Enhancing multi-objective test case selection through the mutation operator
Miriam Ugarte Querejeta, Miren Illarramendi Rezabal, Aitor Arrieta
Autom. Softw. Eng.1
2025 Development of a runtime-condition model for proactive intelligent products using knowledge graphs and embedding
abstract
Modern manufacturing processes’ increasing complexity and variability demand advanced systems capable of real-time monitoring, adaptability, and data-driven decision-making. This paper introduces a novel runtime condition model to enhance interoperability, data integration, and decision support within intelligent manufacturing environments. The model encapsulates key manufacturing elements, including asset management, relationships, key performance indicators (KPIs), capabilities, data structures, constraints, and configurations. A key innovation is the integration of a knowledge graph enriched with embedding techniques, enabling the inference of missing relationships, dynamic reasoning, and predictive analytics. The proposed model was validated through a case study conducted in collaboration with TQC Automation Ltd., using their MicroApplication Leak Test System (MALT). A dataset of over 9,000 unique test configurations demonstrated the model’s capabilities in representing runtime conditions, managing operational parameters, and optimising test configurations. The enriched knowledge graph facilitated advanced analyses, providing actionable insights into test outcomes and enabling proactive decision-making. Empirical results showcase the model’s ability to harmonise diverse data sources, infer missing connections, and improve runtime adaptability. This study highlights the potential of combining runtime modelling with knowledge graphs to address the challenges of modern manufacturing. Future research will explore the model’s application to additional domains, integration with larger datasets, and the use of machine learning for enhanced predictive capabilities.
Hamood Ur Rehman, Miriam Ugarte Querejeta, Angela Carrera-Rivera, Sylvia Nathaly Rea Minango, Fabio Marco Monetti, Antonio Maffei, Jack C. Chaplin
Knowl. Based Syst.3
2023 Search-based Test Case Selection for PLC Systems using Functional Block Diagram Programs
abstract
Programmable Logic Controllers (PLCs) are the core unit of the production system, which frequently need to implement new processes to address customer needs. These changes must be fully tested to ensure the reliability of the PLC code, which is commonly programmed through Functional Block Diagrams (FBDs). This is a tedious task that requires considerable time and effort given the manual nature of the process involved in PLC testing. Hence, we present a cost-effective test selection approach to test FBD programs in dynamic environments. The proposed method uses a search-based multi-objective test case selection algorithm as a regression technique to test recently modified FBD programs. Specifically, we derived a total of 7 fitness function combinations, by combining different cost and quality-based fitness functions. We carried out an empirical evaluation, by employing fitness metrics in the wellknown NSGA-II algorithm to determine the best configuration setup for testing FBD programs. Furthermore, we benchmarked the performance of the NSGA-II with the baseline Random Search (RS). The study was carried out with three case studies of a reactor protection system, and evaluated with two sets of mutants. The results demonstrated that the proposed approach significantly reduces time, while keeping high the overall fault detection capability.
Miriam Ugarte Querejeta, Eunkyoung Jee, Lingjun Liu, Aitor Arrieta, Miren Illarramendi Rezabal
ISSRE1