Tomas Herda

dblp:359/5638 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0005-2912-380XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Recommendations for efficient and responsible LLM adoption within industrial software development
abstract
Context: Large language models (LLMs) are observed to have a significant positive impact on various software engineering (SE) activities. With improved accessibility, the adoption of powerful LLMs in industry has surged recently. However, there is a lack of actionable best practices for the efficient and responsible adoption of LLMs within industrial software settings. Objectives: We developed seven actionable recommendations to address this research gap. Methods: We conducted a multi-case study with three organisations that use LLMs within their SE activities and synthesised seven recommendations through qualitative thematic analysis. We conducted a complementary online survey with software practitioners from various industries to evaluate the perceived relevance of our recommendations. Results: Our results and recommendations focus on (i) users’ preference to use LLMs as AI assistants, (ii) the importance of relevant stakeholders’ satisfaction in the LLM-output evaluation, (iii) scoping the applicability of LLMs within SE tasks, (iv) the effect of LLMs on SE workflows, (v) the necessity and directions for developing human oversight mechanisms, and (vi) the necessary skills for practitioners for leveraging LLMs within SE. The online survey indicates a high level of agreement from the participants regarding the perceived relevance of the recommendations. Conclusion: We outline future research directions, including mapping the seven recommendations to the principles of the EU AI Act (AIA) in order to examine how they relate to the current regulatory compliance frameworks.
Krishna Ronanki, Beatriz Cabrero-Daniel, Tomas Herda, Stefan Sitkovich, Jennifer Horkoff, Christian Berger 0001
Inf. Softw. Technol.3
2025 A Multi-agent LLM System for Automated Requirements Analysis: A Study on User Story Generation and Prioritization
Malik Abdul Sami, Zheying Zhang, Muhammad Waseem 0011, Kai-Kristian Kemell, Zeeshan Rasheed 0001, Tomas Herda, Md Toufique Hasan, Jussi Rasku, Pekka Abrahamsson
SEAA (2)6
2025 Generative Artificial Intelligence for Software Engineering - A Research Agenda
abstract
ABSTRACT Context Generative artificial intelligence (GenAI) tools have become increasingly prevalent in software development, offering assistance to various managerial and technical project activities. Notable examples of these tools include OpenAI's ChatGPT, GitHub Copilot, and Amazon CodeWhisperer. Objective Although many recent publications have explored and evaluated the application of GenAI, a comprehensive understanding of the current development, applications, limitations, and open challenges remains unclear to many. Particularly, we do not have an overall picture of the current state of GenAI technology in practical software engineering usage scenarios. Method We conducted a literature review and focus groups for a duration of five months to develop a research agenda on GenAI for software engineering. Results We identified 78 open research questions (RQs) in 11 areas of software engineering. Our results show that it is possible to explore the adoption of GenAI in partial automation and support decision‐making in all software development activities. While the current literature is skewed toward software implementation, quality assurance and software maintenance, other areas, such as requirements engineering, software design, and software engineering education, would need further research attention. Common considerations when implementing GenAI include industry‐level assessment, dependability and accuracy, data accessibility, transparency, and sustainability aspects associated with the technology. Conclusions GenAI is bringing significant changes to the field of software engineering. Nevertheless, the state of research on the topic still remains immature. We believe that this research agenda holds significance and practical value for informing both researchers and practitioners about current applications and guiding future research.
Anh Nguyen-Duc 0001, Beatriz Cabrero-Daniel, Adam Przybylek, Chetan Arora 0002, Dron Khanna, Tomas Herda, Usman Rafiq, Jorge Melegati, Eduardo Guerra 0001, Kai-Kristian Kemell, Mika Saari, Zheying Zhang, Thanh Tho Quan, Pekka Abrahamsson
Softw. Pract. Exp.6
2024 Generating Test Scenarios from NL Requirements Using Retrieval-Augmented LLMs: An Industrial Study
abstract
Test scenarios are specific instances of test cases that describe a sequence of actions to validate a particular software functionality. By outlining the conditions under which the software operates and the expected outcomes, test scenarios ensure that the software functionality is tested in an integrated manner. Test scenarios are crucial for systematically testing an application under various conditions, including edge cases, to identify potential issues and guarantee overall performance and reliability. Manually specifying test scenarios is tedious and requires a deep understanding of software functionality and the underlying domain. It further demands substantial effort and investment from already time- and budget-constrained requirements engineers and testing teams. This paper presents an automated approach (RAGTAG) for test scenario generation using Retrieval-Augmented Generation (RAG) with Large Language Models (LLMs). RAG allows the integration of specific domain knowledge with LLMs' generation capabilities. We evaluate RAGTAG on two industrial projects from Austrian Post with bilingual requirements in German and English. Our results from an interview survey conducted with four experts on five dimensions – relevance, coverage, correctness, coherence and feasibility, affirm the potential of RAGTAG in automating test scenario generation. Specifically, our results indicate that, despite the difficult task of analyzing bilingual requirements, RAGTAG is able to produce scenarios that are well-aligned with the underlying requirements and provide coverage of different aspects of the intended functionality. The generated scenarios are easily understandable to experts and feasible for testing in the project environment. The overall correctness is deemed satisfactory; however, gaps in capturing exact action sequences and domain nuances remain, underscoring the need for domain expertise when applying LLMs.
Chetan Arora 0002, Tomas Herda, Verena Homm
RE2
2024 Exploring Human-AI Collaboration in Agile: Customised LLM Meeting Assistants
abstract
Abstract This action research study focuses on the integration of “AI assistants” in two Agile software development meetings: the Daily Scrum and a feature refinement, a planning meeting that is part of an in-house Scaled Agile framework. We discuss the critical drivers of success, and establish a link between the use of AI and team collaboration dynamics. We conclude with a list of lessons learnt during the interventions in an industrial context, and provide a assessment checklist for companies and teams to reflect on their readiness level. This paper is thus a road-map to facilitate the integration of AI tools in Agile setups.
Beatriz Cabrero-Daniel, Tomas Herda, Victoria Pichler, Martin Eder
XP2
2024 LLM-Based Agents for Automating the Enhancement of User Story Quality: An Early Report
abstract
Abstract In agile software development, maintaining high-quality user stories is crucial, but also challenging. This study explores the application of large language models (LLMs) to improve the quality of user stories within the agile teams of Austrian Post Group IT. We developed an Autonomous LLM-based Agent System (ALAS) and evaluated its impact on user story quality with 11 participants from six agile teams. Our findings reveal the potential of LLMs in improving user story quality, provide a practical example, and lay the foundation for future research into the broad application of LLMs in a variety of industry settings.
Zheying Zhang, Maruf Rayhan, Tomas Herda, Manuel Goisauf, Pekka Abrahamsson
XP3