VLDB 2026 Research / reviewers in the wild / expert
Maider Azanza
dblp:72/6498
· DBLP profile ↗
15ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-4537-1572ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond the hype: Enabling informed LLM adoption in industry through systematic evaluationabstractThe adoption of Large Language Models (LLMs) in software development has accelerated substantially, yet organizations lack systematic frameworks for continuous evaluation of LLM-powered development tools, platforms such as GitHub Copilot that leverage LLMs to automate code generation, testing, and refactoring. Unlike traditional software dependencies with predictable versioning, these tools evolve continuously as providers update models and architectures, making point-in-time assessments insufficient for informed adoption decisions. We present a 20-month longitudinal study of LLM-based test generation at LKS Next, a technology consultancy, conducted across six evaluation cycles from March 2024 through October 2025. Our framework systematically assessed test quality through objective metrics (compilation, coverage, code quality) and expert evaluation of generated tests, tracking multiple models (GPT, Claude, Gemini) and tool configurations over time. Our findings reveal the volatile nature of the LLM tool ecosystem: models achieving over 90% quality scores experienced unexpected regressions in subsequent cycles, GitHub Copilot architectural changes affected all models despite unchanged prompts, and high-performing models became unavailable. Custom-prompted agents outperformed generic tools by 20-90% across different quality metrics. These temporal patterns, invisible in point-in-time evaluations, show that continuous monitoring is necessary for industrial adoption. Building on these insights, we generalize our approach into a domain-independent framework adapting Goal-Question-Metric to LLM-specific challenges including rapid evolution, prompt engineering, and continuous tracking. We present applicability through test generation and code refactoring evaluations and present Tetrics, a research prototype showing that systematic evaluation is actionable in practice. Our work provides evidence that informed LLM adoption requires continuous, organization-specific evaluation frameworks. Eneko Pizarro, Maider Azanza, Beatriz Pérez Lamancha |
Sci. Comput. Program. | 2 |
| 2025 | Tracking the Moving Target: A Framework for Continuous Evaluation of LLM Test Generation in IndustryabstractLarge Language Models (LLMs) have shown great potential in automating software testing tasks, including test generation. However, their rapid evolution poses a critical challenge for companies implementing DevSecOps - evaluations of their effectiveness quickly become outdated, making it difficult to assess their reliability for production use. While academic research has extensively studied LLM-based test generation, evaluations typically provide point-in-time analyses using academic benchmarks. Such evaluations do not address the practical needs of companies who must continuously assess tool reliability and integration with existing development practices. This work presents a measurement framework for the continuous evaluation of commercial LLM test generators in industrial environments. We demonstrate its effectiveness through a longitudinal study at LKS Next. The framework integrates with industry-standard tools like SonarQube and provides metrics that evaluate both technical adequacy (e.g., test coverage) and practical considerations (e.g., maintainability or expert assessment). Our methodology incorporates strategies for test case selection, prompt engineering, and measurement infrastructure, addressing challenges such as data leakage and reproducibility. Results highlight both the rapid evolution of LLM capabilities and critical factors for successful industrial adoption, offering practical guidance for companies seeking to integrate these technologies into their development pipelines. Maider Azanza, Beatriz Pérez Lamancha, Eneko Pizarro |
EASE | 1 |
| 2024 | Can LLMs Facilitate Onboarding Software Developers? An Ongoing Industrial Case StudyabstractOnboarding new software developers presents per-sistent challenges for teams, with newcomers facing steep learning curves and senior staff burdened by providing training and mentoring. This research explores leveraging Large Language Models (LLMs) to streamline onboarding, mitigating productivity losses from disrupted workflows. We collaborated with LKS Next, an IT consulting firm that has experienced first hand the challenges onboarding new developers brings. This paper presents an ongoing case study within LKS Next, conducted using action design research methodology. Two cycles have been completed, gathering feedback from newcomers, team leaders, and managers on issues like the need to prioritize self-directed learning resources over “burdening” mentors, and considerations around using third-party LLMs. The third cycle will focus on evaluating open source LLMs to maintain control within the company. Our envisioned goal is to develop an LLM-powered conversational agent, delivering tailored onboarding support while avoiding privacy risks. Maider Azanza, Juanan Pereira, Arantza Irastorza, Aritz Galdos |
CSEE&T | 1 |
| 2024 | Leveraging Open Source LLMs for Software Engineering Education and TrainingabstractGenerative AI, particularly Large Language Models (LLMs), presents innovative opportunities to enhance software engineering education. Open source LLMs such as LLaMA and Mistral leverage the potential of generative AI offering distinct advantages over proprietary options including transparency, customizability, collaboration, and cost savings. This paper de-velops a catalog of LLM prompt examples tailored for software engineering training, mapped to knowledge areas from the Soft-ware Engineering Body of Knowledge (SWEBoK) framework. Example prompts demonstrate LLMs' capabilities in eliciting requirements, diagram generation, API simulation, effort esti-mation through role-playing, and other areas. The methodology involves evaluating prompt responses from ChatGPT, Mistral, and LLaMA on representative tasks. Quantitative and qualitative analysis assesses quality, usefulness, and correctness. Findings show ChatGPT and Mistral outperforming LLaMA overall, but no model perfectly executes complex interactions. We examine implications and challenges of integrating open source LLMs into classrooms, emphasizing the need for oversight, verification, and prompt design aligned with pedagogical objectives. Juanan Pereira, Juan Miguel López 0001, Xabier Garmendia 0001, Maider Azanza |
CSEE&T | 4 |
| 2022 | Visualizing the customization endeavor in product-based-evolving software product lines: a case of action design researchabstractAbstract Software Product Lines (SPLs) aim at systematically reusing software assets, and deriving products (a.k.a., variants) out of those assets. However, it is not always possible to handle SPL evolution directly through these reusable assets. Time-to-market pressure, expedited bug fixes, or product specifics lead to the evolution to first happen at the product level, and to be later merged back into the SPL platform where the core assets reside. This is referred to as product-based evolution. In this scenario, deciding when and what should go into the next SPL release is far from trivial. Distinct questions arise. How much effort are developers spending on product customization? Which are the most customized core assets? To which extent is the core asset code being reused for a given product? We refer to this endeavor as Customization Analysis, i.e., understanding the functional increments in adjusting products from the last SPL platform release. The scale of the SPLs’ code-base calls for customization analysis to be conducted through Visual Analytics tools. This work addresses the design principles for such tools through a joint effort between academia and industry, specifically, Danfoss Drives, a company division in charge of the P400 SPL. Accordingly, we adopt an Action Design Research approach where answers are sought by interacting with the practitioners in the studied situations. We contribute by providing informed goals for customization analysis as well as an intervention in terms of a visual analytics tool. We conclude by discussing to what extent this experience can be generalized to product-based evolving SPL organizations other than Danfoss Drives. Oscar Díaz 0001, Leticia Montalvillo-Mendizabal, Raul Medeiros, Maider Azanza, Thomas Fogdal |
Empir. Softw. Eng. | 4 |
| 2017 | Teaching model-driven engineering from a relational database perspective
Don S. Batory, Maider Azanza |
Softw. Syst. Model. | 2 |
| 2016 | Improving refactoring speed by 10XabstractRefactoring engines are standard tools in today's Integrated Development Environments (IDEs). They allow programmers to perform one refactoring at a time, but programmers need more. Most design patterns in the Gang-of-Four text can be written as a refactoring script -- a programmatic sequence of refactorings. In this paper, we present R3, a new Java refactoring engine that supports refactoring scripts. It builds a main-memory, non-persistent database to encode Java entity declarations (e.g., packages, classes, methods), their containment relationships, and language features such as inheritance and modifiers. Unlike classical refactoring engines that modify Abstract Syntax Trees (ASTs), R3 refactorings modify only the database; refactored code is produced only when pretty-printing ASTs that reference database changes. R3 performs comparable precondition checks to those of the Eclipse Java Development Tools (JDT) but R3's codebase is about half the size of the JDT refactoring engine and runs an order of magnitude faster. Further, a user study shows that R3 improved the success rate of retrofitting design patterns by 25% up to 50%. Jongwook Kim, Don S. Batory, Danny Dig, Maider Azanza |
ICSE | 4 |
| 2015 | Editing Anxiety in Corporate Wikis: From Private Drafting to Public Edits
Cristóbal Arellano, Oscar Díaz 0001, Maider Azanza |
CAiSE | 3 |
| 2014 | Embedded software product lines: domain and application engineering model-based analysis processesabstractSUMMARY Nowadays, embedded systems are gaining importance. At the same time, the development of their software is increasing its complexity, having to deal with quality, cost, and time‐to‐market issues among others. With stringent quality requirements such as performance, early verification and validation become critical in these systems. In this regard, advanced development paradigms such as model‐driven engineering and software product line engineering bring considerable benefits to the development and validation of embedded system software. However, these benefits come at the cost of increasing process complexity. This work presents a process based on UML and MARTE for the analysis of embedded model‐driven product lines. It specifies the tasks, the involved roles, and the workproducts that form the process and how it is integrated in the more general development process. Existing tools that support the tasks to be performed in the process are also described. A classification of such tools and a study of traceability among them are provided, allowing engineering teams to choose the most adequate chain of tools to support the process. Copyright © 2012 John Wiley & Sons, Ltd. Lorea Belategi, Goiuria Sagardui Mendieta, Leire Etxeberria Elorza, Maider Azanza |
J. Softw. Evol. Process. | 4 |
| 2013 | Teaching Model Driven Engineering from a Relational Database Perspective
Don S. Batory, Eric Latimer, Maider Azanza |
MoDELS | 3 |
| 2013 | A language for end-user web augmentation: Caring for producers and consumers alikeabstractWeb augmentation is to the Web what augmented reality is to the physical world: layering relevant content/layout/navigation over the existing Web to customize the user experience. This is achieved through JavaScript (JS) using browser weavers (e.g., Greasemonkey). To date, over 43 million of downloads of Greasemonkey scripts ground the vitality of this movement. However, Web augmentation is hindered by being programming intensive and prone to malware. This prevents end-users from participating as both producers and consumers of scripts: producers need to know JS, consumers need to trust JS. This article aims at promoting end-user participation in both roles. The vision is for end-users to prosume (the act of simultaneously caring for producing and consuming) scripts as easily as they currently prosume their pictures or videos. Encouraging production requires more “natural” and abstract constructs. Promoting consumption calls for augmentation scripts to be easier to understand, share, and trust upon. To this end, we explore the use of Domain-Specific Languages (DSLs) by introducing Sticklet . Sticklet is an internal DSL on JS, where JS generality is reduced for the sake of learnability and reliability. Specifically, Web augmentation is conceived as fixing in existing web sites (i.e., the wall ) HTML fragments extracted from either other sites or Web services (i.e., the stickers ). Sticklet targets hobby programmers as producers, and computer literates as consumers. From a producer perspective, benefits are threefold. As a restricted grammar on top of JS, Sticklet expressions are domain oriented and more declarative than their JS counterparts, hence speeding up development. As syntactically correct JS expressions, Sticklet scripts can be installed as traditional scripts and hence, programmers can continue using existing JS tools. As declarative expressions, they are easier to maintain, and amenable for optimization. From a consumer perspective, domain specificity brings understandability (due to declarativeness), reliability (due to built-in security), and “consumability” (i.e., installation/enactment/sharing of Sticklet expressions are tuned to the shortage of time and skills of the target audience). Preliminary evaluations indicate that 77% of the subjects were able to develop new Sticklet scripts in less than thirty minutes while 84% were able to consume these scripts in less than ten minutes. Sticklet is available to download as a Mozilla add-on. Oscar Díaz 0001, Cristóbal Arellano, Maider Azanza |
ACM Trans. Web | 3 |
| 2012 | Model Transformation Co-evolution: A Semi-automatic Approach
Jokin García, Oscar Díaz 0001, Maider Azanza |
SLE | 3 |
| 2008 | The Objects and Arrows of Computational Design
Don S. Batory, Maider Azanza, João Saraiva |
MoDELS | 2 |
| 2007 | Generative metaprogrammingabstractRecent advances in Software Engineering have reduced the cost of coding programs at the expense of increasing the complexity of program synthesis, i.e. metaprograms, which when executed, will synthesize a target program. The traditional cycle of configuring-linking-compiling, now needs to be supplemented with additional transformation steps that refine and enhance an initial specification until the target program is obtained. So far, these synthesis processes are based on error-prone, hand-crafted scripting. To depart from this situation, this paper addresses generative metaprogramming, i.e. the generation of program-synthesis metaprograms from declarative specifications. To this end, we explore (i) the (meta) primitives for program synthesis, (ii) the architecture that dictates how these primitives can be intertwined, and (iii) the declarative specification of the metaprogram from which the code counterpart is generated. Salvador Trujillo, Maider Azanza, Oscar Díaz 0001 |
GPCE | 2 |
| 2006 | Modeling Portlet Aggregation Through Statecharts
Oscar Díaz 0001, Arantza Irastorza, Maider Azanza, Felipe M. Villoria |
WISE | 3 |