Panagiota Chatzipetrou

dblp:09/10404 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-0311-1502ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 23 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
YearPublicationVenuePosition
2026 Evaluating the quality of GenAI applications in software engineering: a multi-case study
abstract
Abstract Context Generative AI (GenAI) is increasingly adopted in software development for tasks such as document generation, data analysis, and code generation. However, evaluating the quality of GenAI applications becomes challenging, as traditional quality measurements may not be fully applicable. Objective In this study, we explore how practitioners evaluate the quality of GenAI applications and investigate quality evaluation techniques. Method We conducted a multi-case study in three industrial projects from software development companies. We examined four GenAI application domains: document generation, data analysis and insight generation, customer service, and code generation. Data were collected through three workshops and 23 semi-structured interviews with industrial practitioners. Results We identified fourteen GenAI use cases and 28 metrics currently used to evaluate the quality of GenAI applications’ outputs. We synthesized the identified metrics’ usage patterns and challenges based on the collected data. Conclusions This study presents practical insights into using metrics to measure GenAI-based system qualities in real industrial settings. Our findings indicate that practitioners use custom-built and context-specific metrics; combining these with academic metrics can strengthen GenAI system quality evaluation.
Emil Alégroth, Panagiota Chatzipetrou, Tony Gorschek
Empir. Softw. Eng.3
2026 A Framework for Evaluating GenAI Adoption and Use in Software Engineering
abstract
Generative Artificial Intelligence (GenAI) is increasingly integrated into software products to enable new features and user capabilities, from early exploration to operational deployment. GenAI adoption as a component within a software system introduces quality risks because GenAI outputs are probabilistic, prompt-sensitive, and may drift after release. Organizations, therefore, need to decide what to evaluate, when to evaluate, and who owns quality evaluation activities across software design, development, and operations. ISO/IEC 25059 standard distinguishes between software product quality (e.g., usability) and quality-in-use (e.g., satisfaction) for AI-enabled software, yet it provides limited operational guidance for these evaluation activities. We therefore investigate how industrial software teams adopt and use GenAI models in the software systems they build and operate, and how they evaluate system qualities when deciding to adopt GenAI during development and after deployment. We do not benchmark the underlying GenAI model itself. In this study, we conducted 19 semi-structured interviews in two software development companies. We triangulated the interviews with archival data (15 internal documents and 184 internal wiki/web pages) to capture GenAI adoption steps, quality concerns, evaluation practices, and role responsibilities. Our findings describe a three-phase adoption process – Ideation, Development, and Operation – highlighting where quality evaluations occur, which criteria are used, and how evaluation responsibilities are distributed. Based on observed practices and using ISO/IEC 25059 as an organizing lens, we synthesize a process-oriented quality evaluation framework. This framework maps metrics to explicit gatekeeping, validation, and monitoring checkpoints, bridging abstract ISO quality characteristics with engineering workflows. We applied the framework in a GenAI-enabled software product (SE4AI) use case and reported how it supported structured evaluation activities. We also observed that quality evaluations span legal, security, development, QA, and operations, but ownership is fragmented across phases. We therefore propose a GenAI Quality Lead responsibility (often assignable to an existing senior role) to coordinate criteria, evidence, and traceability across quality evaluation activities. The results contribute to Software Engineering for AI (SE4AI) by clarifying how teams can measure qualities when building software that adopts and uses GenAI.
Emil Alégroth, Panagiota Chatzipetrou, Tony Gorschek
IEEE Trans. Software Eng.3
2025 Trust vs. Control: Comparing Flexible and Restrictive Hybrid Work Policies in Two Software Companies
abstract
The pandemic experiences of forced work from home (WFH) in the tech industry turned out better than expected. As a result, modern workplaces that employ software engineers have become increasingly hybrid allowing employees to alternate days spent in the office with days spent working remotely. Yet, approaches to regulate the hybrid work arrangements vary. Some companies implemented strict policies with controlled office presence while others rely on recommendations and permit greater locational flexibility. In this paper, we evaluate how different degrees of locational flexibility influence individual work arrangements and satisfaction in two comparative cases: a company with high degree of flexibility (FinCo) and a company with mandatory office presence (TelCo). Through a survey of 547 practitioners, our findings reveal that flexible policies can achieve higher voluntary office attendance than mandatory requirements. To our surprise, we found that the number of employees visiting the office at least 2-3 days per week in the company with greater flexibility was higher (68%) compared to the company with mandatory attendance (58%). The study also identifies key factors influencing work location choices, including commute time and role. The type of tasks and dependencies with colleagues also matter - employees with more WFH days tended to have more individual tasks, while those with more onsite work days engaged in more collaborative tasks and had colleagues who depended on them. Our results suggest that trust-based approaches and creating attractive office environments may be more effective than strict attendance policies in maintaining desired office presence while supporting employee satisfaction. These findings contribute practical insights for organizations seeking to establish effective post-pandemic work policies for software engineers.
Darja Smite, Nils Brede Moe, Panagiota Chatzipetrou, Povilas Godliauskas, Per Kristian Helland, Anastasiia Tkalich
EASE3
2025 Language Models to Support Multi-Label Classification of Industrial Data
abstract
Background: Multi-label requirements classification is an inherently challenging task, especially when dealing with numerous classes at varying levels of abstraction. The task becomes even more difficult when a limited number of requirements is available to train a supervised classifier. Zero-shot learning does not require training data and can potentially address this problem. Objective: This paper investigates the performance of zero-shot classifiers on a multi-label industrial dataset. The study focuses on classifying requirements according to a hierarchical taxonomy designed to support requirements tracing. Method: We compare multiple variants of zero-shot classifiers using different embeddings, including 9 language models (LMs) with a reduced number of parameters (up to 3B), e.g., BERT, and 5 large LMs (LLMs) with a large number of parameters (up to 70B), e.g., Llama. Our ground truth includes 377 requirements and 1968 labels from 6 output spaces. For the evaluation, we adopt traditional metrics, i.e., precision, recall,$F_{1}$, and$F_{\beta}$, as well as a novel label distance metric$D_{n}$. This aims to better capture the classification's hierarchical nature and to provide a more nuanced evaluation of how far the results are from the ground truth. Results: 1) The top-performing model on 5 out of$\mathbf{6}$output spaces is TS-xl, with maximum$F_{\beta}=0.78$and$D_{n}=0.04$, while BERT base outperformed the other models in one case, with maximum$F_{\beta}=0.83$and$D_{n}=0.04.2$) LMs with smaller parameter size produce the best classification results compared to LLMs. Thus, addressing the problem in practice is feasible as limited computing power is needed. 3) The model architecture (auto encoding, autoregression, and sentence-to-sentence) significantly affects the classifier's performance. Contribution: We conclude that using zero-shot learning for multi-label requirements classification offers promising results. We also present a novel metric that can be used to select the top-performing model for this problem.
Waleed Abdeen, Michael Unterkalmsteiner, Krzysztof Wnuk, Alessio Ferrari 0001, Panagiota Chatzipetrou
SANER5
2025 Measuring the quality of generative AI systems: Mapping metrics to quality characteristics - Snowballing literature review
abstract
Context : Generative Artificial Intelligence (GenAI) and the use of Large Language Models (LLMs) have revolutionized tasks that previously required significant human effort, which has attracted considerable interest from industry stakeholders. This growing interest has accelerated the integration of AI models into various industrial applications. However, the model integration introduces challenges to product quality, as conventional quality measuring methods may fail to assess GenAI systems. Consequently, evaluation techniques for GenAI systems need to be adapted and refined. Examining the current state and applicability of evaluation techniques for the GenAI system outputs is essential. Objective : This study aims to explore the current metrics, methods, and processes for assessing the outputs of GenAI systems and the potential of risky outputs. Method : We performed a snowballing literature review to identify metrics, evaluation methods, and evaluation processes from 43 selected papers. Results : We identified 28 metrics and mapped these metrics to four quality characteristics defined by the ISO/IEC 25023 standard for software systems. Additionally, we discovered three types of evaluation methods to measure the quality of system outputs and a three-step process to assess faulty system outputs. Based on these insights, we suggested a five-step framework for measuring system quality while utilizing GenAI models. Conclusion : Our findings present a mapping that visualizes candidate metrics to be selected for measuring quality characteristics of GenAI systems, accompanied by step-by-step processes to assist practitioners in conducting quality assessments.
Emil Alégroth, Panagiota Chatzipetrou, Tony Gorschek
Inf. Softw. Technol.3
2025 Quality attributes of test cases and test suites - importance & challenges from practitioners' perspectives
abstract
Abstract The quality of the test suites and the constituent test cases significantly impacts confidence in software testing. While research has identified several quality attributes of test cases and test suites, there is a need for a better understanding of their relative importance in practice. We investigate practitioners’ perceptions regarding the relative importance of quality attributes of test cases and test suites and the challenges that they face in ensuring the perceived important quality attributes. To capture the practitioners’ perceptions, we conducted an industrial survey using a questionnaire based on the quality attributes identified in an extensive literature review. We used a sampling strategy that leverages LinkedIn to draw a large and heterogeneous sample of professionals with experience in software testing. We collected 354 responses from practitioners with a wide range of experience (from less than one year to 42 years of experience). We found that the majority of practitioners rated Fault Detection, Usability, Maintainability, Reliability, and Coverage to be the most important quality attributes. Resource Efficiency, Reusability, and Simplicity received the most divergent opinions, which, according to our analysis, depend on the software-testing contexts. Also, we identified common challenges that apply to the important attributes, namely inadequate definition, lack of useful metrics, lack of an established review process, and lack of external support. The findings point out where practitioners actually need further support with respect to achieving high-quality test cases and test suites under different software testing contexts. Hence, the findings can serve as a guideline for academic researchers when looking for research directions on the topic. Furthermore, the findings can be used to encourage companies to provide more support to practitioners to achieve high-quality test cases and test suites.
Huynh Khanh Vi Tran, Nauman Bin Ali, Michael Unterkalmsteiner, Jürgen Börstler, Panagiota Chatzipetrou
Softw. Qual. J.5
2024 Interest in Working Remotely: Is Gender a Factor?
Panagiota Chatzipetrou, Darja Smite, Anastasiia Tkalich, Nils Brede Moe, Eriks Klotins
PROFES1
2024 Experience with Large Language Model Applications for Information Retrieval from Enterprise Proprietary Data
Emil Alégroth, Panagiota Chatzipetrou, Tony Gorschek
PROFES3
2023 Automated NFR testing in continuous integration environments: a multi-case study of Nordic companies
abstract
Abstract Context Non-functional requirements (NFRs) (also referred to as system qualities) are essential for developing high-quality software. Notwithstanding its importance, NFR testing remains challenging, especially in terms of automation. Compared to manual verification, automated testing shows the potential to improve the efficiency and effectiveness of quality assurance, especially in the context of Continuous Integration (CI). However, studies on how companies manage automated NFR testing through CI are limited. Objective This study examines how automated NFR testing can be enabled and supported using CI environments in software development companies. Method We performed a multi-case study at four companies by conducting 22 semi-structured interviews with industrial practitioners. Results Maintainability,reliability,performance,securityandscalability, were found to be evaluated with automated tests in CI environments. Testing practices, quality metrics, and challenges for measuring NFRs were reported. Conclusions This study presents an empirically derived model that shows how data produced by CI environments can be used for evaluation and monitoring of implemented NFR quality. Additionally, the manuscript presents explicit metrics, CI components, tools, and challenges that shall be considered while performing NFR testing in practice.
Emil Alégroth, Panagiota Chatzipetrou, Tony Gorschek
Empir. Softw. Eng.3
2023 The state-of-practice in requirements specification: an extended interview study at 12 companies
abstract
Abstract Requirements specification is a core activity in the requirements engineering phase of a software development project. Researchers have contributed extensively to the field of requirements specification, but the extent to which their proposals have been adopted in practice remains unclear. We gathered evidence about the state of practice in requirements specification by focussing on the artefacts used in this activity, the application of templates or guidelines, how requirements are structured in the specification document, what tools practitioners use to specify requirements, and what challenges they face. We conducted an interview-based survey study involving 24 practitioners from 12 different Swedish IT companies. We recorded the interviews and analysed these recordings, primarily by using qualitative methods. Natural language constitutes the main specification artefact but is usually accompanied by some other type of instrument. Most requirements specifications use templates or guidelines, although they seldom follow any fixed standard. Requirements are always structured in the document according to the main functionalities of the system or to project areas or system parts. Different types of tools, including MS Office tools, are used, either individually or combined, in the compilation of requirements specifications. We also note that challenges related to the use of natural language (dealing with ambiguity, inconsistency, and incompleteness) are the most frequent challenges that practitioners face in the compilation of requirements specifications. These findings are contextualized in terms of demographic factors related to the individual interviewees, the organization they are affiliated with, and the project they selected to discuss during our interviews. A number of our findings have been previously reported in related studies. These findings show that, in spite of the large number of notations, models and tools proposed from academia for improving requirements specification, practitioners still mainly rely on plain natural language and general-purpose tool support. We expect more empirical studies in this area in order to better understand the reason of this low adoption of research results.
Xavier Franch, Cristina Palomares, Carme Quer, Panagiota Chatzipetrou, Tony Gorschek
Requir. Eng.4
2021 The state-of-practice in requirements elicitation: an extended interview study at 12 companies
Cristina Palomares, Xavier Franch, Carme Quer, Panagiota Chatzipetrou, Lidia López 0001, Tony Gorschek
Requir. Eng.4
2021 A Progression Model of Software Engineering Goals, Challenges, and Practices in Start-Ups
abstract
Context: Software start-ups are emerging as suppliers of innovation and software-intensive products. However, traditional software engineering practices are not evaluated in the context, nor adopted to goals and challenges of start-ups. As a result, there is insufficient support for software engineering in the start-up context. Objective: We aim to collect data related to engineering goals, challenges, and practices in start-up companies to ascertain trends and patterns characterizing engineering work in start-ups. Such data allows researchers to understand better how goals and challenges are related to practices. This understanding can then inform future studies aimed at designing solutions addressing those goals and challenges. Besides, these trends and patterns can be useful for practitioners to make more informed decisions in their engineering practice. Method: We use a case survey method to gather first-hand, in-depth experiences from a large sample of software start-ups. We use open coding and cross-case analysis to describe and identify patterns, and corroborate the findings with statistical analysis. Results: We analyze 84 start-up cases and identify 16 goals, 9 challenges, and 16 engineering practices that are common among start-ups. We have mapped these goals, challenges, and practices to start-up life-cycle stages (inception, stabilization, growth, and maturity). Thus, creating the progression model guiding software engineering efforts in start-ups. Conclusions: We conclude that start-ups to a large extent face the same challenges and use the same practices as established companies. However, the primary software engineering challenge in start-ups is to evolve multiple process areas at once, with a little margin for serious errors.
Eriks Klotins, Michael Unterkalmsteiner, Panagiota Chatzipetrou, Tony Gorschek, Rafael Prikladnicki, Nirnaya Tripathi, Leandro Bento Pompermaier
IEEE Trans. Software Eng.3
2020 Utilising CI environment for efficient and effective testing of NFRs
Emil Alégroth, Panagiota Chatzipetrou, Tony Gorschek
Inf. Softw. Technol.3
2020 Component attributes and their importance in decisions and component selection
abstract
Component-based software engineering is a common approach in the development and evolution of contemporary software systems. Different component sourcing options are available, such as: (1) Software developed internally (in-house) , (2) Software developed outsourced , (3) Commercial off-the-shelf software , and (4) Open-Source Software . However, there is little available research on what attributes of a component are the most important ones when selecting new components. The objective of this study is to investigate what matters the most to industry practitioners when they decide to select a component. We conducted a cross-domain anonymous survey with industry practitioners involved in component selection. First, the practitioners selected the most important attributes from a list. Next, they prioritized their selection using the Hundred-Dollar ($100) test. We analyzed the results using compositional data analysis. The results of this exploratory analysis showed that cost was clearly considered to be the most important attribute for component selection. Other important attributes for the practitioners were: support of the component , longevity prediction , and level of off-the-shelf fit to product . Moreover, several practitioners still consider in-house software development to be the sole option when adding or replacing a component. On the other hand, there is a trend to complement it with other component sourcing options and, apart from cost, different attributes factor into their decision. Furthermore, in our analysis, nonparametric tests and biplots were used to further investigate the practitioners’ inherent characteristics. It seems that smaller and larger organizations have different views on what attributes are the most important, and the most surprising finding is their contrasting views on the cost attribute: larger organizations with mature products are considerably more cost aware.
Panagiota Chatzipetrou, Efi Papatheocharous, Krzysztof Wnuk, Markus Borg, Emil Alégroth, Tony Gorschek
Softw. Qual. J.1
2019 Requirements' Characteristics: How do they Impact on Project Budget in a Systems Engineering Context?
abstract
Background: Requirements engineering is of a principal importance when starting a new project. However, the number of the requirements involved in a single project can reach up to thousands. Controlling and assuring the quality of natural language requirements (NLRs), in these quantities, is challenging. Aims: In a field study, we investigated with the Swedish Transportation Agency (STA) to what extent the characteristics of requirements had an influence on change requests and budget changes in the project. Method: We choose the following models to characterize system requirements formulated in natural language: Concern-based Model of Requirements (CMR), Requirements Abstractions Model (RAM) and Software-Hardware model (SHM). The classification of the NLRs was conducted by the three authors. The robust statistical measure Fleiss’ Kappa was used to verify the reliability of the results. We used descriptive statistics, contingency tables, results from the Chi-Square test of association along with post hoc tests. Finally, a multivariate statistical technique, Correspondence analysis was used in order to provide a means of displaying a set of requirements in two-dimensional graphical form. Results: The results showed that software requirements are associated with less budget cost than hardware requirements. Moreover, software requirements tend to stay open for a longer period indicating that they are ”harder” to handle. Finally, the more discussion or interaction on a change request can lower the actual estimated change request cost. Conclusions: The results lead us to a need to further investigate the reasons why the software requirements are treated differently from the hardware requirements, interview the project managers, understand better the way those requirements are formulated and propose effective ways of Software management.
Panagiota Chatzipetrou, Michael Unterkalmsteiner, Tony Gorschek
SEAA1
2019 Selecting component sourcing options: A survey of software engineering's broader make-or-buy decisions
Markus Borg, Panagiota Chatzipetrou, Krzysztof Wnuk, Emil Alégroth, Tony Gorschek, Efi Papatheocharous, Syed Muhammad Ali Shah, Jakob Axelsson
Inf. Softw. Technol.2
2019 Understanding the order of agile practice introduction: Comparing agile maturity models and practitioners' experience
Indira Nurdiani, Jürgen Börstler, Samuel Fricker, Kai Petersen, Panagiota Chatzipetrou
J. Syst. Softw.5
2018 When and who leaves matters: emerging results from an empirical study of employee turnover
abstract
Background: Employee turnover in GSD is an extremely important issue, especially in Western companies offshoring to emerging nations. Aims: In this case study we investigated an offshore vendor company and in particular whether the employees' retention is related with their experience. Moreover, we studied whether we can identify a threshold associated with the employees' tendency to leave the particular company. Method: We used a case study, applied and presented descriptive statistics, contingency tables, results from Chi-Square test of association and post hoc tests. Results: The emerging results showed that employee retention and company experience are associated. In particular, almost 90% of the employees are leaving the company within the first year, where the percentage within the second year is 50-50%. Thus, there is an indication that the 2 years' time is the retention threshold for the investigated offshore vendor company. Conclusions: The results are preliminary and lead us to the need for building a prediction model which should include more inherent characteristics of the projects to aid the companies avoiding massive turnover waves.
Panagiota Chatzipetrou, Darja Smite, Rini van Solingen
ESEM1
2018 Component Selection in Software Engineering - Which Attributes are the Most Important in the Decision Process?
abstract
Component-based software engineering is a common approach to develop and evolve contemporary software systems where different component sourcing options are available: 1)Software developed internally (in-house), 2)Software developed outsourced, 3)Commercial of the shelf software, and 4) Open Source Software. However, there is little available research on what attributes of a component are the most important ones when selecting new components. The object of the present study is to investigate what matters the most to industry practitioners during component selection. We conducted a cross-domain anonymous survey with industry practitioners involved in component selection. First, the practitioners selected the most important attributes from a list. Next, they prioritized their selection using the Hundred-Dollar ($100) test. We analyzed the results using Compositional Data Analysis. The descriptive results showed that Cost was clearly considered the most important attribute during the component selection. Other important attributes for the practitioners were: Support of the component, Longevity prediction, and Level of off-the-shelf fit to product. Next, an exploratory analysis was conducted based on the practitioners' inherent characteristics. Nonparametric tests and biplots were used. It seems that smaller organizations and more immature products focus on different attributes than bigger organizations and mature products which focus more on Cost.
Panagiota Chatzipetrou, Emil Alégroth, Efi Papatheocharous, Markus Borg, Tony Gorschek, Krzysztof Wnuk
SEAA1
2015 A multivariate statistical framework for the analysis of software effort phase distribution
Panagiota Chatzipetrou, Efi Papatheocharous, Lefteris Angelis, Andreas S. Andreou
Inf. Softw. Technol.1
2015 An experience-based framework for evaluating alignment of software quality goals
Panagiota Chatzipetrou, Lefteris Angelis, Sebastian Barney, Claes Wohlin
Softw. Qual. J.1
2014 Software quality across borders: Three case studies on company internal alignment
Sebastian Barney, Varun Mohankumar, Panagiota Chatzipetrou, Aybüke Aurum, Claes Wohlin, Lefteris Angelis
Inf. Softw. Technol.3
2011 Offshore Insourcing: A Case Study on Software Quality Alignment
abstract
Background: Software quality issues are commonly reported when off shoring software development. Value-based software engineering addresses this by ensuring key stakeholders have a common understanding of quality. Aim: This work seeks to understand the levels of alignment between key stakeholders on aspects of software quality for two products developed as part of an offshore in sourcing arrangement. The study further aims to explain the levels of alignment identified. Method: Representatives of key stakeholder groups for both products ranked aspects of software quality. The results were discussed with the groups to gain a deeper understanding. Results: Low levels of alignment were found between the groups studied. This is associated with insufficiently defined quality requirements, a culture that does not question management and conflicting temporal reflections on the product's quality. Conclusion: The work emphasizes the need for greater support to align success-critical stakeholder groups in their understanding of quality when off shoring software development.
Sebastian Barney, Claes Wohlin, Panagiota Chatzipetrou, Lefteris Angelis
ICGSE3