Beatriz Pérez Lamancha

dblp:93/4641 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
2since 2021 · last 2026
0009-0001-0672-8081ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 13 · 7 first-author · 2 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Beyond the hype: Enabling informed LLM adoption in industry through systematic evaluation
abstract
The adoption of Large Language Models (LLMs) in software development has accelerated substantially, yet organizations lack systematic frameworks for continuous evaluation of LLM-powered development tools, platforms such as GitHub Copilot that leverage LLMs to automate code generation, testing, and refactoring. Unlike traditional software dependencies with predictable versioning, these tools evolve continuously as providers update models and architectures, making point-in-time assessments insufficient for informed adoption decisions. We present a 20-month longitudinal study of LLM-based test generation at LKS Next, a technology consultancy, conducted across six evaluation cycles from March 2024 through October 2025. Our framework systematically assessed test quality through objective metrics (compilation, coverage, code quality) and expert evaluation of generated tests, tracking multiple models (GPT, Claude, Gemini) and tool configurations over time. Our findings reveal the volatile nature of the LLM tool ecosystem: models achieving over 90% quality scores experienced unexpected regressions in subsequent cycles, GitHub Copilot architectural changes affected all models despite unchanged prompts, and high-performing models became unavailable. Custom-prompted agents outperformed generic tools by 20-90% across different quality metrics. These temporal patterns, invisible in point-in-time evaluations, show that continuous monitoring is necessary for industrial adoption. Building on these insights, we generalize our approach into a domain-independent framework adapting Goal-Question-Metric to LLM-specific challenges including rapid evolution, prompt engineering, and continuous tracking. We present applicability through test generation and code refactoring evaluations and present Tetrics, a research prototype showing that systematic evaluation is actionable in practice. Our work provides evidence that informed LLM adoption requires continuous, organization-specific evaluation frameworks.
Eneko Pizarro, Maider Azanza, Beatriz Pérez Lamancha
Sci. Comput. Program.3
2025 Tracking the Moving Target: A Framework for Continuous Evaluation of LLM Test Generation in Industry
abstract
Large Language Models (LLMs) have shown great potential in automating software testing tasks, including test generation. However, their rapid evolution poses a critical challenge for companies implementing DevSecOps - evaluations of their effectiveness quickly become outdated, making it difficult to assess their reliability for production use. While academic research has extensively studied LLM-based test generation, evaluations typically provide point-in-time analyses using academic benchmarks. Such evaluations do not address the practical needs of companies who must continuously assess tool reliability and integration with existing development practices. This work presents a measurement framework for the continuous evaluation of commercial LLM test generators in industrial environments. We demonstrate its effectiveness through a longitudinal study at LKS Next. The framework integrates with industry-standard tools like SonarQube and provides metrics that evaluate both technical adequacy (e.g., test coverage) and practical considerations (e.g., maintainability or expert assessment). Our methodology incorporates strategies for test case selection, prompt engineering, and measurement infrastructure, addressing challenges such as data leakage and reproducibility. Results highlight both the rapid evolution of LLM capabilities and critical factors for successful industrial adoption, offering practical guidance for companies seeking to integrate these technologies into their development pipelines.
Maider Azanza, Beatriz Pérez Lamancha, Eneko Pizarro
EASE2
2015 PROW: A Pairwise algorithm with constRaints, Order and Weight
Beatriz Pérez Lamancha, Macario Polo, Mario Piattini
J. Syst. Softw.1
2014 Extending UML Testing Profile Towards Non-functional Test Modeling
abstract
The research community has broadly recognized the importance of the validation of non-functional properties including performance and dependability requirements. However, the results of a systematic survey we carried out evidenced the lack of a standard notation for designing non-functional test cases. For some time, the greatest attention of Model-Based Testing (MBT) research has focused on functional aspects. The only exception is represented by the UML Testing Profile (UML-TP) that is a lightweight extension of UML to support the design of testing artifacts, but it only provides limited support for non-functional testing. In this paper we provide a first attempt to extend UML-TP for improving the design of non-functional tests. The proposed extension deals with some important concepts of non-functional testing such as the workload and the global verdicts. As a proof of concept we show how the extended UML-TP can be used for modeling non-functional test cases of an application example.
Federico Toledo Rodríguez, Francesca Lonetti, Antonia Bertolino, Macario Polo, Beatriz Pérez Lamancha
MODELSWARD5
2013 Automated generation of test oracles using a model-driven approach
Beatriz Pérez Lamancha, Macario Polo, Danilo Caivano, Mario Piattini, Giuseppe Visaggio
Inf. Softw. Technol.1
2012 Reduction of Test Suites Using Mutation
Macario Polo, Pedro Reales Mateo, Beatriz Pérez Lamancha
FASE3
2012 Towards a Framework for Information System Testing - A Model-driven Testing Approach
Federico Toledo Rodríguez, Beatriz Pérez Lamancha, Macario Polo
ICSOFT2
2011 Model-driven Testing - Transformations from Test Models to Test Code
Beatriz Pérez Lamancha, Pedro Reales Mateo, Macario Polo, Danilo Caivano
ENASE1
2010 An Automated Model-driven Testing Framework - For Model-Driven Development and Software Product Lines
Beatriz Pérez Lamancha, Macario Polo, Mario Piattini
ENASE1
2010 Testing Product Generation in Software Product Lines Using Pairwise for Features Coverage
Beatriz Pérez Lamancha, Macario Polo
ICTSS1
2009 Model-driven testing in software product lines
abstract
This paper describes a model-based method for the automatic generation of test cases in software product lines. The approach is completely based on well-known standards, such as UML 2.0, the UML Testing Profile and the QVT Language. To reach the goal, it is essential to maintain the traceability between different abstraction levels, as well as between domain and product engineering levels. The method is achieved using the Orthogonal Variability Model as an UML Profile.
Beatriz Pérez Lamancha, Macario Polo, Ignacio García Rodríguez de Guzmán
ICSM1
2009 Software Product Line Testing - A Systematic Review
Beatriz Pérez Lamancha, Macario Polo, Mario Piattini
ICSOFT (1)1
2009 Some Experiments on Test Case Tracebaility
Macario Polo, Beatriz Pérez Lamancha, Pedro Reales Mateo
SEKE2