Beatriz Cabrero-Daniel

dblp:264/8812 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0001-5275-8372ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 12 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Recommendations for efficient and responsible LLM adoption within industrial software development
abstract
Context: Large language models (LLMs) are observed to have a significant positive impact on various software engineering (SE) activities. With improved accessibility, the adoption of powerful LLMs in industry has surged recently. However, there is a lack of actionable best practices for the efficient and responsible adoption of LLMs within industrial software settings. Objectives: We developed seven actionable recommendations to address this research gap. Methods: We conducted a multi-case study with three organisations that use LLMs within their SE activities and synthesised seven recommendations through qualitative thematic analysis. We conducted a complementary online survey with software practitioners from various industries to evaluate the perceived relevance of our recommendations. Results: Our results and recommendations focus on (i) users’ preference to use LLMs as AI assistants, (ii) the importance of relevant stakeholders’ satisfaction in the LLM-output evaluation, (iii) scoping the applicability of LLMs within SE tasks, (iv) the effect of LLMs on SE workflows, (v) the necessity and directions for developing human oversight mechanisms, and (vi) the necessary skills for practitioners for leveraging LLMs within SE. The online survey indicates a high level of agreement from the participants regarding the perceived relevance of the recommendations. Conclusion: We outline future research directions, including mapping the seven recommendations to the principles of the EU AI Act (AIA) in order to examine how they relate to the current regulatory compliance frameworks.
Krishna Ronanki, Beatriz Cabrero-Daniel, Tomas Herda, Stefan Sitkovich, Jennifer Horkoff, Christian Berger 0001
Inf. Softw. Technol.2
2025 On Simulation-Guided LLM-based Code Generation for Safe Autonomous Driving Software
abstract
Automated Driving System (ADS) is a safety-critical software system responsible for the interpretation of the vehicle’s environment and making decisions accordingly. The unbounded complexity of the driving context, including unforeseeable events, necessitate continuous improvement, often achieved through iterative DevOps processes. However, DevOps processes are themselves complex, making these improvements both time- and resource-intensive. Automation in code generation for ADS using Large Language Models (LLM) is one potential approach to address this challenge. Nevertheless, the development of ADS requires rigorous processes to verify, validate, assess, and qualify the code before it can be deployed in the vehicle and used. In this study, we developed and evaluated a prototype for automatic code generation and assessment using a designed pipeline of a LLM-based agent, simulation model, and rule-based feedback generator in an industrial setup. The LLM-generated code is evaluated automatically in a simulation model against multiple critical traffic scenarios, and an assessment report is provided as feedback to the LLM for modification or bug fixing. We report about the experimental results of the prototype employing Codellama:34b, DeepSeek (r1:32b and Coder:33b), CodeGemma:7b, Mistral:7b, and GPT4 for Adaptive Cruise Control (ACC) and Unsupervised Collision Avoidance by Evasive Manoeuvre (CAEM). We finally assessed the tool with 11 experts at two Original Equipment Manufacturers (OEMs) by conducting an interview study.
Ali Nouri, Johan Andersson, Kailash De Jesus Hornig, Zhennan Fei, Emil Knabe, Håkan Sivencrona, Beatriz Cabrero-Daniel, Christian Berger 0001
EASE7
2025 Challenges of Virtual Validation and Verification for Automotive Functions
Beatriz Cabrero-Daniel, Mazen Mohamad
SEAA1
2025 Large Language Models in Code Co-generation for Safe Autonomous Vehicles
Ali Nouri, Beatriz Cabrero-Daniel, Zhennan Fei, Krishna Ronanki, Håkan Sivencrona, Christian Berger 0001
SAFECOMP2
2025 The DevSafeOps dilemma: A systematic literature review on rapidity in safe autonomous driving development and operation
abstract
Developing autonomous driving (AD) systems is challenging due to the complexity of the systems and the need to assure their safe and reliable operation. The widely adopted approach of DevOps seems promising to support the continuous technological progress in AI and the demand for fast reaction to incidents, which necessitate continuous development, deployment, and monitoring. We present a systematic literature review meant to identify, analyse, and synthesise a broad range of existing literature related to usage of DevOps in autonomous driving development. Our results provide a structured overview of challenges and solutions, arising from applying DevOps to safety-related AI-enabled functions. Our results indicate that there are still several open topics to be addressed to enable safe DevOps for the development of safe AD. • Applying DevOps to autonomous driving presents several open topics to be addressed. • DevSafeOps is introduced, adding safety-related activities into DevOps iterative loops. • Our systematic literature review led to 11 challenges in the DevSafeOps loop. • Potential solutions are identified and mapped to challenges in DevSafeOps.
Ali Nouri, Beatriz Cabrero-Daniel, Fredrik Törner, Christian Berger 0001
J. Syst. Softw.2
2025 Generative Artificial Intelligence for Software Engineering - A Research Agenda
abstract
ABSTRACT Context Generative artificial intelligence (GenAI) tools have become increasingly prevalent in software development, offering assistance to various managerial and technical project activities. Notable examples of these tools include OpenAI's ChatGPT, GitHub Copilot, and Amazon CodeWhisperer. Objective Although many recent publications have explored and evaluated the application of GenAI, a comprehensive understanding of the current development, applications, limitations, and open challenges remains unclear to many. Particularly, we do not have an overall picture of the current state of GenAI technology in practical software engineering usage scenarios. Method We conducted a literature review and focus groups for a duration of five months to develop a research agenda on GenAI for software engineering. Results We identified 78 open research questions (RQs) in 11 areas of software engineering. Our results show that it is possible to explore the adoption of GenAI in partial automation and support decision‐making in all software development activities. While the current literature is skewed toward software implementation, quality assurance and software maintenance, other areas, such as requirements engineering, software design, and software engineering education, would need further research attention. Common considerations when implementing GenAI include industry‐level assessment, dependability and accuracy, data accessibility, transparency, and sustainability aspects associated with the technology. Conclusions GenAI is bringing significant changes to the field of software engineering. Nevertheless, the state of research on the topic still remains immature. We believe that this research agenda holds significance and practical value for informing both researchers and practitioners about current applications and guiding future research.
Anh Nguyen-Duc 0001, Beatriz Cabrero-Daniel, Adam Przybylek, Chetan Arora 0002, Dron Khanna, Tomas Herda, Usman Rafiq, Jorge Melegati, Eduardo Guerra 0001, Kai-Kristian Kemell, Mika Saari, Zheying Zhang, Thanh Tho Quan, Pekka Abrahamsson
Softw. Pract. Exp.2
2024 Welcome Your New AI Teammate: On Safety Analysis by Leashing Large Language Models
abstract
DevOps is a necessity in many industries, including the development of Autonomous Vehicles. In those settings, there are iterative activities that reduce the speed of SafetyOps cycles. One of these activities is "Hazard Analysis & Risk Assessment" (HARA), which is an essential step to start the safety requirements specification. As a potential approach to increase the speed of this step in SafetyOps, we have delved into the capabilities of Large Language Models (LLMs). Our objective is to systematically assess their potential for application in the field of safety engineering. To that end, we propose a framework to support a higher degree of automation of HARA with LLMs. Despite our endeavors to automate as much of the process as possible, expert review remains crucial to ensure the validity and correctness of the analysis results, with necessary modifications made accordingly.
Ali Nouri, Beatriz Cabrero-Daniel, Fredrik Törner, Håkan Sivencrona, Christian Berger 0001
CAIN2
2024 Prompt Smells: An Omen for Undesirable Generative AI Outputs
abstract
Recent trends in the world of Generative Artificial Intelligence (GenAI) focus on developing deep learning (DL)-based models capable of learning structures and temporal patterns from supplied training data to generate content in different formats like text, images, or sound. GenAI models have been widely used in various applications, including creating stories, illustrations, poems, articles, computer code, music compositions, and videos [5, 11].
Krishna Ronanki, Beatriz Cabrero-Daniel, Christian Berger 0001
CAIN2
2024 Digital Twins for Early Verification and Validation of Autonomous Driving Features: Open-source Tools and Standard Formats
abstract
Getting real data from safety critical systems for specific, rare situations (e.g., edge cases) is challenging. Moreover, data gathered during operations (e.g., crash reports) are not often publicly accessible and reports might be incomplete. However, covering all scenarios is very important for Verification and Validation (V&V) of safety-critical systems such as AVs. Therefore, synthetic data could be used for V&V to fill that gap. Synthetic data generation, labelling, and validation are open challenges, though. To the best of the authors’ knowledge, no standard methods for integrating synthetic data into V&V are shared across automotive-domain companies. Therefore, this study (i) gathers expert knowledge on current practices for Digital Twins for V&V development, (ii) proposes a general 6-stage pipeline for synthetic data usage within an early V&V process, and (iii) discusses open source tools and formats standardisation of synthetic data use within V&V. The open-source tools and format standardisation may facilitate the integration of synthetic data into the V&V process. The proposed pipeline and mapping study, provide a foundation for future research on synthetic data use within V&V.
Beatriz Cabrero-Daniel, Ahmed Yasser Abdelkarim, Axel Broberg
IV1
2024 LLMs Can Check Their Own Results to Mitigate Hallucinations in Traffic Understanding Tasks
Malsha Ashani Mahawatta Dona, Beatriz Cabrero-Daniel, Yinan Yu, Christian Berger 0001
ICTSS2
2024 Automating Requirements Review in the Automotive Sector: A Tailored AI Approach
abstract
Requirements serve as the foundation for defining what a software product should accomplish, highlighting the importance of clear and well-written specifications [1]. Deficient requirements often lead to defects in delivered software, which can be challenging and costly to rectify [2].
Sivajeet Chand, Cristina Martinez Montes, Beatriz Cabrero-Daniel, Jennifer Horkoff
RE4
2024 Engineering Safety Requirements for Autonomous Driving with Large Language Models
abstract
Changes and updates in the requirement artifacts, which can be frequent in the automotive domain, are a challenge for SafetyOps. Large Language Models (LLMs), with their impressive natural language understanding and generating capabilities, can play a key role in automatically refining and decomposing requirements after each update. In this study, we propose a prototype of a pipeline of prompts and LLMs that receives an item definition and outputs solutions in the form of safety requirements. This pipeline also performs a review of the requirement dataset and identifies redundant or contradictory requirements. We first identified the necessary characteristics for performing HARA and then defined tests to assess an LLM's capability in meeting these criteria. We used design science with multiple iterations and let experts from different companies evaluate each cycle quantitatively and qualitatively. Finally, the prototype was implemented at a case company and the responsible team evaluated its efficiency.
Ali Nouri, Beatriz Cabrero-Daniel, Fredrik Törner, Håkan Sivencrona, Christian Berger 0001
RE2
2024 Exploring Human-AI Collaboration in Agile: Customised LLM Meeting Assistants
abstract
Abstract This action research study focuses on the integration of “AI assistants” in two Agile software development meetings: the Daily Scrum and a feature refinement, a planning meeting that is part of an in-house Scaled Agile framework. We discuss the critical drivers of success, and establish a link between the use of AI and team collaboration dynamics. We conclude with a list of lessons learnt during the interventions in an industrial context, and provide a assessment checklist for companies and teams to reflect on their readiness level. This paper is thus a road-map to facilitate the integration of AI tools in Agile setups.
Beatriz Cabrero-Daniel, Tomas Herda, Victoria Pichler, Martin Eder
XP1
2024 Generative AI in Software Engineering Must Be Human-Centered: The Copenhagen Manifesto
Daniel Russo 0002, Sebastian Baltes, Niels van Berkel, Paris Avgeriou, Fabio Calefato, Beatriz Cabrero-Daniel, Gemma Catolino, Jürgen Cito, Neil A. Ernst, Thomas Fritz 0001, Hideaki Hata, Reid Holmes, Maliheh Izadi, Foutse Khomh, Mikkel Baun Kjærgaard, Grischa Liebel, Alberto Lluch-Lafuente, Stefano Lambiase, Walid Maalej, Gail C. Murphy, Nils Brede Moe, Gabrielle O'Brien, Elda Paja, Mauro Pezzè, John Stouby Persson, Rafael Prikladnicki, Paul Ralph, Martin P. Robillard, Thiago Rocha Silva, Klaas-Jan Stol, Margaret-Anne D. Storey, Viktoria Stray, Paolo Tell, Christoph Treude, Bogdan Vasilescu
J. Syst. Softw.6
2022 Dynamic Combination of Crowd Steering Policies Based on Context
abstract
Abstract Simulating crowds requires controlling a very large number of trajectories of characters and is usually performed using crowd steering algorithms. The question of choosing the right algorithm with the right parameter values is of crucial importance given the large impact on the quality of results. In this paper, we study the performance of a number of steering policies (i.e., simulation algorithm and its parameters) in a variety of contexts, resorting to an existing quality function able to automatically evaluate simulation results. This analysis allows us to map contexts to the performance of steering policies. Based on this mapping, we demonstrate that distributing the best performing policies among characters improves the resulting simulations. Furthermore, we also propose a solution to dynamically adjust the policies, for each agent independently and while the simulation is running, based on the local context each agent is currently in. We demonstrate significant improvements of simulation results compared to previous work that would optimize parameters once for the whole simulation, or pick an optimized, but unique and static, policy for a given global simulation context.
Beatriz Cabrero-Daniel, Ricardo Marques, Ludovic Hoyet, Julien Pettré, Josep Blat
Comput. Graph. Forum1
2020 Generalized Microscropic Crowd Simulation using Costs in Velocity Space
abstract
To simulate the low-level (‘microscopic’) behavior of human crowds, a local navigation algorithm computes how a single person (‘agent’) should move based on its surroundings. Many algorithms for this purpose have been proposed, each using different principles and implementation details that are difficult to compare.
Wouter van Toll, Fabien Grzeskowiak, Axel López-Gandía, Javad Amirian, Florian Berton, Julien Bruneau 0002, Beatriz Cabrero-Daniel, Alberto Jovane, Julien Pettré
I3D7