Andreas Vogelsang

dblp:19/8725 · DBLP profile ↗
← Back
59ranked-venue papers
10as first author
22since 2021 · last 2026
0000-0003-1041-0815ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 50 · 10 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 6 · 5 since 2021Artificial intelligence and machine learning · 2Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Opportunities and Limitations of GenAI in RE: Viewpoints from Practice
Anne Hess, Andreas Vogelsang, Xavier Franch, Andrea Herrmann 0001, Sylwia Kopczynska, Alexander Rachmann
REFSQ2
2026 Explanation strategies in smart environments: Effects on usability, system understanding, and task performance
abstract
• First empirical comparison of explanation strategies in smart environments: We conduct an empirical study (N = 159) to systematically compare three explanation: no explanation, static explanations (fixed and identical across users and contexts), and context-adapted explanations (dynamically adjusting content based on user preferences and situational factors). To the best of our knowledge, this is the first empirical study in the rule-based smart environments domain to directly contrast static and context-adapted explanations in terms of their impact on task performance, system understanding, usability, and perceived explanation quality. • Explanations boost performance and understanding: We found that providing explanations, whether static or context-adapted, significantly improves user task performance and comprehension of system behavior compared to offering no explanations. • Context-adaptivity is not always superior: We revealed that context-adapted explanations are not universally better than static ones, with their benefits depending on factors such as task complexity and user preferences. The increasing complexity of interactive smart environments presents a significant engineering challenge: their automated, context-aware decisions are often opaque, undermining system usability. While explainability is becoming known as a promising remedy, systematically evaluating different explanation strategies for these systems remains an open problem. This paper presents a rigorous empirical evaluation to address this gap. We conducted a controlled experiment (N = 159) to compare three approaches: no explanation, static explanations, and context-adapted explanations. Our results quantify the significant benefits of explainability on task performance and user understanding of system behavior. Furthermore, we identify a key engineering tradeoff: context-adapted explanations are not universally superior to simpler static implementations, suggesting that a one-size-fits-all approach is suboptimal in this domain. To support our evaluation, we developed a web-based testbed simulating a smart home environment with light gamification elements, enabling reliable human-grounded assessment. Based on our findings, we offer concrete insights and recommendations to guide the design of explainable interactive systems. Our study underscores the importance of tailoring explanations to user needs and contextual factors, contributing to more transparent and user-friendly smart environments.
Mersedeh Sadeghi, Simon Scholz, Anna Trapp, Max Unterbusch, Andreas Vogelsang
J. Syst. Softw.5
2025 Prompts as Software Engineering Artifacts: A Research Agenda and Preliminary Findings
Hugo Villamizar, Jannik Fischbach, Alexander Korn, Andreas Vogelsang, Daniel Méndez 0001
PROFES4
2025 LLMREI: Automating Requirements Elicitation Interviews with LLMs
abstract
Requirements elicitation interviews are crucial for gathering system requirements but heavily depend on skilled analysts, making them resource-intensive, susceptible to human biases, and prone to miscommunication. Recent advancements in Large Language Models present new opportunities for automating parts of this process. This study introduces LLMREI, a chat bot designed to conduct requirements elicitation interviews with minimal human intervention, aiming to reduce common interviewer errors and improve the scalability of requirements elicitation. We explored two main approaches, zero-shot prompting and least-to-most prompting, to optimize LLMREI for requirements elicitation and evaluated its performance in 33 simulated stakeholder interviews. A third approach, fine-tuning, was initially considered but abandoned due to poor performance in preliminary trials. Our study assesses the chat bot’s effectiveness in three key areas: minimizing common interview errors, extracting relevant requirements, and adapting its questioning based on interview context and user responses. Our findings indicate that LLMREI makes a similar number of errors compared to human interviewers, is capable of extracting a large portion of requirements, and demonstrates a notable ability to generate highly context-dependent questions. We envision the greatest benefit of LLMREI in automating interviews with a large number of stakeholders.
Alexander Korn, Samuel Gorsch, Andreas Vogelsang
RE3
2025 From Requirements to Code: Understanding Developer Practices in LLM-Assisted Software Engineering
abstract
With the advent of generative LLMs and their advanced code generation capabilities, some people already envision the end of traditional software engineering, as LLMs may be able to produce high-quality code based solely on the requirements a domain expert feeds into the system. The feasibility of this vision can be assessed by understanding how developers currently incorporate requirements when using LLMs for code generation—a topic that remains largely unexplored. We interviewed 18 practitioners from 14 companies to understand how they (re)use information from requirements and other design artifacts to feed LLMs when generating code. Based on our findings, we propose a theory that explains the processes developers employ and the artifacts they rely on. Our theory suggests that requirements, as typically documented, are too abstract for direct input into LLMs. Instead, they must first be manually decomposed into programming tasks, which are then enriched with design decisions and architectural constraints before being used in prompts. Our study highlights that fundamental RE work is still necessary when LLMs are used to generate code. Our theory is important for contextualizing scientific approaches to automating requirements-centric SE tasks.
Jonathan Ullrich, Matthias Koch, Andreas Vogelsang
RE3
2024 SmartEx: A Framework for Generating User-Centric Explanations in Smart Environments
abstract
Explainability is crucial for complex systems like pervasive smart environments, as they collect and analyze data from various sensors, follow multiple rules, and control different devices resulting in behavior that is not trivial and, thus, should be explained to the users. The current approaches, however, offer flat, static, and algorithm-focused explanations. User-centric explanations, on the other hand, consider the recipient and context, providing personalized and context-aware explanations. To address this gap, we propose an approach to incorporate user-centric explanations into smart environments. We introduce a conceptual model and a reference architecture for characterizing and generating such explanations. Our work is the first technical solution for generating context-aware and granular explanations in smart environments. Our architecture implementation demonstrates the feasibility of our approach through various scenarios.
Mersedeh Sadeghi, Lars Herbold, Max Unterbusch, Andreas Vogelsang
PerCom4
2024 Requirements Engineering for Research Software: A Vision
abstract
Modern science is relying on software more than ever. The behavior and outcomes of this software shape the scientific and public discourse on important topics like climate change, economic growth, or the spread of infections. Most researchers creating software for scientific purposes are not trained in Software Engineering. As a consequence, research software is often developed ad hoc without following stringent processes. With this paper, we want to characterize research software as a new application domain that needs attention from the Requirements Engineering community. We conducted an exploratory study based on 8 interviews with 12 researchers who develop software. We describe how researchers elicit, document, and analyze requirements for research software and what processes they follow. From this, we derive specific challenges and describe a vision of Requirements Engineering for research software.
Adrian Bajraktari, Michelle Binder, Andreas Vogelsang
RE3
2024 Explaining the Unexplainable: The Impact of Misleading Explanations on Trust in Unreliable Predictions for Hardly Assessable Tasks
abstract
To increase trust in systems, engineers strive to create explanations that are as accurate as possible. However, if the system’s accuracy is compromised, providing explanations for its incorrect behavior may inadvertently lead to misleading explanations. This concern is particularly pertinent when the correctness of the system is difficult for users to judge. In an online survey experiment with 162 participants, we analyze the impact of misleading explanations on users’ perceived and demonstrated trust in a system that performs a hardly assessable task in an unreliable manner. Participants who used a system that provided potentially misleading explanations rated their trust significantly higher than participants who saw the system’s prediction alone. They also aligned their initial prediction with the system’s prediction significantly more often. Our findings underscore the importance of exercising caution when generating explanations, especially in tasks that are inherently difficult to evaluate. The paper and supplementary materials are available at https://doi.org/10.17605/osf.io/azu72
Mersedeh Sadeghi, Daniel Pöttgen, Patrick Ebel 0001, Andreas Vogelsang
UMAP4
2024 Interoperability of heterogeneous Systems of Systems: from requirements to a reference architecture
abstract
Abstract Interoperability stands as a critical hurdle in developing and overseeing distributed and collaborative systems. Thus, it becomes imperative to gain a deep comprehension of the primary obstacles hindering interoperability and the essential criteria that systems must satisfy to achieve it. In light of this objective, in the initial phase of this research, we conducted a survey questionnaire involving stakeholders and practitioners engaged in distributed and collaborative systems. This effort resulted in the identification of eight essential interoperability requirements, along with their corresponding challenges. Then, the second part of our study encompassed a critical review of the literature to assess the effectiveness of prevailing conceptual approaches and associated technologies in addressing the identified requirements. This analysis led to the identification of a set of components that promise to deliver the desired interoperability by addressing the requirements identified earlier. These elements subsequently form the foundation for the third part of our study, a reference architecture for interoperability-fostering frameworks that is proposed in this paper. The results of our research can significantly impact the software engineering of interoperable systems by introducing their fundamental requirements and the best practices to address them, but also by identifying the key elements of a framework facilitating interoperability in Systems of Systems.
Mersedeh Sadeghi, Alessio Carenini, Óscar Corcho, Matteo G. Rossi, Riccardo Santoro, Andreas Vogelsang
J. Supercomput.6
2023 Exploring Millions of User Interactions with ICEBOAT: Big Data Analytics for Automotive User Interfaces
abstract
User Experience (UX) professionals need to be able to analyze large amounts of usage data on their own to make evidence-based design decisions. However, the design process for In-Vehicle Information Systems (IVISs) lacks data-driven support and effective tools for visualizing and analyzing user interaction data. Therefore, we propose ICEBOAT1, an interactive visualization tool tailored to the needs of automotive UX experts to effectively and efficiently evaluate driver interactions with IVISs. ICEBOAT visualizes telematics data collected from production line vehicles, allowing UX experts to perform task-specific analyses. Following a mixed methods User-Centered Design (UCD) approach, we conducted an interview study (N=4) to extract the domain specific information and interaction needs of automotive UX experts and used a co-design approach (N=4) to develop an interactive analysis tool. Our evaluation (N=12) shows that ICEBOAT enables UX experts to efficiently generate knowledge that facilitates data-driven design decisions.
Patrick Ebel 0001, Kim Julian Gülle, Christoph Lingenfelder, Andreas Vogelsang
AutomotiveUI4
2023 Automatically Classifying Kano Model Factors in App Reviews
Michelle Binder, Annika Vogt, Adrian Bajraktari, Andreas Vogelsang
REFSQ4
2023 Multitasking While Driving: How Drivers Self-Regulate Their Interaction with In-Vehicle Touchscreens in Automated Driving
abstract
Driver assistance systems are designed to increase comfort and safety by automating parts of the driving task. At the same time, modern in-vehicle information systems with large touchscreens provide the driver with numerous options for entertainment, information, or communication, and are a potential source of distraction. However, little is known about how driving automation affects how drivers interact with the center stack touchscreen, i.e., how drivers self-regulate their behavior in response to different levels of driving automation. To investigate this, we apply multilevel models to a real-world driving dataset consisting of 31,378 sequences. Our results show significant differences in drivers’ interaction and glance behavior in response to different levels of driving automation, vehicle speed, and road curvature. During automated driving, drivers perform more interactions per touchscreen sequence and increase the time spent looking at the center stack touchscreen. Specifically, at higher levels of driving automation (level 2), the mean glance duration toward the center stack touchscreen increases by 36% and the mean number of interactions per sequence increases by 17% compared to manual driving. Furthermore, partially automated driving has a strong impact on the use of more complex UI elements (e.g., maps) and touch gestures (e.g., multitouch). We also show that the effect of driving automation on drivers’ self-regulation is greater than that of vehicle speed and road curvature. The derived knowledge can inform the design and evaluation of touch-based infotainment systems and the development of context-aware driver monitoring systems.
Patrick Ebel 0001, Christoph Lingenfelder, Andreas Vogelsang
Int. J. Hum. Comput. Interact.3
2023 Introduction to the Special Section on the Best Papers from REFSQ 2022
Vincenzo Gervasi, Andreas Vogelsang
Inf. Softw. Technol.2
2023 Automatic creation of acceptance tests by extracting conditionals from requirements: NLP approach and case study
Jannik Fischbach, Julian Frattini, Andreas Vogelsang, Daniel Méndez 0001, Michael Unterkalmsteiner, Andreas Wehrle, Pablo Restrepo Henao, Parisa Yousefi, Tedi Juricic, Jeannette Radduenz, Carsten Wiecher
J. Syst. Softw.3
2023 Causality in requirements artifacts: prevalence, detection, and impact
abstract
Abstract Causal relations in natural language (NL) requirements convey strong, semantic information. Automatically extracting such causal information enables multiple use cases, such as test case generation, but it also requires to reliably detect causal relations in the first place. Currently, this is still a cumbersome task as causality in NL requirements is still barely understood and, thus, barely detectable. In our empirically informed research, we aim at better understanding the notion of causality and supporting the automatic extraction of causal relations in NL requirements. In a first case study, we investigate 14.983 sentences from 53 requirements documents to understand the extent and form in which causality occurs. Second, we present and evaluate a tool-supported approach, called CiRA, for causality detection. We conclude with a second case study where we demonstrate the applicability of our tool and investigate the impact of causality on NL requirements. The first case study shows that causality constitutes around 28 % of all NL requirements sentences. We then demonstrate that our detection tool achieves a macro- $$\hbox {F}_{1}$$ F1 score of 82 % on real-world data and that it outperforms related approaches with an average gain of 11.06 % in macro-Recall and 11.43 % in macro-Precision. Finally, our second case study corroborates the positive correlations of causality with features of NL requirements. The results strengthen our confidence in the eligibility of causal relations for downstream reuse, while our tool and publicly available data constitute a first step in the ongoing endeavors of utilizing causality in RE and beyond.
Julian Frattini, Jannik Fischbach, Daniel Méndez 0001, Michael Unterkalmsteiner, Andreas Vogelsang, Krzysztof Wnuk
Requir. Eng.5
2022 How do Practitioners Perceive the Relevance of Requirements Engineering Research?
abstract
Context: The relevance of Requirements Engineering (RE) research to practitioners is vital for a long-term dissemination of research results to everyday practice. Some authors have speculated about a mismatch between research and practice in the RE discipline. However, there is not much evidence to support or refute this perception.Objective: This article presents the results of a study aimed at gathering evidence from practitioners about their perception of the relevance of RE research and at understanding the factors that influence that perception.Method: We conducted a questionnaire-based survey of industry practitioners with expertise in RE. The participants rated the perceived relevance of 435 scientific papers presented at five top RE-related conferences.Results: The 153 participants provided a total of 2,164 ratings. The practitioners rated RE research as essential or worthwhile in a majority of cases. However, the percentage of non-positive ratings is still higher than we would like. Among the factors that affect the perception of relevance are the research's links to industry, the research method used, and respondents’ roles. The reasons for positive perceptions were primarily related to the relevance of the problem and the soundness of the solution, while the causes for negative perceptions were more varied. The respondents also provided suggestions for future research, including topics researchers have studied for decades, like elicitation or requirement quality criteria.Conclusions: The study is valuable for both researchers and practitioners. Researchers can use the reasons respondents gave for positive and negative perceptions and the suggested research topics to help make their research more appealing to practitioners and thus more prone to industry adoption. Practitioners can benefit from the overall view of contemporary RE research by learning about research topics that they may not be familiar with, and compare their perception with those of their colleagues to self-assess their positioning towards more academic research.
Xavier Franch, Daniel Méndez 0001, Andreas Vogelsang, Rogardt Heldal, Eric Knauss, Marc Oriol, Guilherme Horta Travassos, Jeffrey C. Carver, Thomas Zimmermann 0001
IEEE Trans. Software Eng.3
2021 Visualizing Event Sequence Data for User Behavior Evaluation of In-Vehicle Information Systems
abstract
With modern In-Vehicle Information Systems (IVISs) becoming more capable and complex than ever, their evaluation becomes increasingly difficult. The analysis of large amounts of user behavior data can help to cope with this complexity and can support UX experts in designing IVISs that serve customer needs and are safe to operate while driving. We, therefore, propose a Multi-level User Behavior Visualization Framework providing effective visualizations of user behavior data that is collected via telematics from production vehicles. Our approach visualizes user behavior data on three different levels: (1) The Task Level View aggregates event sequence data generated through touchscreen interactions to visualize user flows. (2) The Flow Level View allows comparing the individual flows based on a chosen metric. (3) The Sequence Level View provides detailed insights into touch interactions, glance, and driving behavior. Our case study proves that UX experts consider our approach a useful addition to their design process.
Patrick Ebel 0001, Christoph Lingenfelder, Andreas Vogelsang
AutomotiveUI3
2021 Integrated and Iterative Requirements Analysis and Test Specification: A Case Study at Kostal
abstract
Currently, practitioners follow a top-down approach in automotive development projects. However, recent studies have shown that this top-down approach is not suitable for the implementation and testing of modern automotive systems. Specifically, practitioners increasingly fail to specify requirements and tests for systems with complex component interactions (e.g., e-mobility systems). In this paper, we address this research gap and propose an integrated and iterative scenario-based technique for the specification of requirements and test scenarios. Our idea is to combine both a top-down and a bottom-up integration strategy. For the top-down approach, we use a behavior-driven development (BDD) technique to drive the modeling of high-level system interactions from the user's perspective. For the bottom-up approach, we discovered that natural language processing (NLP) techniques are suited to make textual specifications of existing components accessible to our technique. To integrate both directions, we support the joint execution and automated analysis of system-level interactions and component-level behavior. We demonstrate the feasibility of our approach by conducting a case study at Kostal (Tierl supplier). The case study corroborates, among other things, that our approach supports practitioners in improving requirements and test specifications for integrated system behavior.
Carsten Wiecher, Jannik Fischbach, Joel Greenyer, Andreas Vogelsang, Carsten Wolff, Roman Dumitrescu
MoDELS4
2021 How Do Practitioners Interpret Conditionals in Requirements?
Jannik Fischbach, Julian Frattini, Daniel Méndez 0001, Michael Unterkalmsteiner, Henning Femmer, Andreas Vogelsang
PROFES6
2021 Automatic Detection of Causality in Requirement Artifacts: The CiRA Approach
Jannik Fischbach, Julian Frattini, Arjen Spaans, Maximilian Kummeth, Andreas Vogelsang, Daniel Méndez 0001, Michael Unterkalmsteiner
REFSQ5
2021 Improving Trace Link Recovery Using Semantic Relation Graphs and Spreading Activation
Aaron Schlutter, Andreas Vogelsang
REFSQ2
2021 Characteristics, potentials, and limitations of open-source Simulink projects for empirical research
abstract
Abstract Simulink is an example of a successful application of the paradigm of model-based development into industrial practice. Numerous companies create and maintain Simulink projects for modeling software-intensive embedded systems, aiming at early validation and automated code generation. However, Simulink projects are not as easily available as code-based ones, which profit from large publicly accessible open-source repositories, thus curbing empirical research. In this paper, we investigate a set of 1734 freely available Simulink models from 194 projects and analyze their suitability for empirical research. We analyze the projects considering (1) their development context, (2) their complexity in terms of size and organization within projects, and (3) their evolution over time. Our results show that there are both limitations and potentials for empirical research. On the one hand, some application domains dominate the development context, and there is a large number of models that can be considered toy examples of limited practical relevance. These often stem from an academic context, consist of only a few Simulink blocks, and are no longer (or have never been) under active development or maintenance. On the other hand, we found that a subset of the analyzed models is of considerable size and complexity. There are models comprising several thousands of blocks, some of them highly modularized by hierarchically organized Simulink subsystems. Likewise, some of the models expose an active maintenance span of several years, which indicates that they are used as primary development artifacts throughout a project’s lifecycle. According to a discussion of our results with a domain expert, many models can be considered mature enough for quality analysis purposes, and they expose characteristics that can be considered representative for industry-scale models. Thus, we are confident that a subset of the models is suitable for empirical research. More generally, using a publicly available model corpus or a dedicated subset enables researchers to replicate findings, publish subsequent studies, and use them for validation purposes. We publish our dataset for the sake of replicating our results and fostering future empirical research.
Alexander Boll, Florian Brokhausen, Tiago Amorim 0001, Timo Kehrer, Andreas Vogelsang
Softw. Syst. Model.5
2020 The Role and Potentials of Field User Interaction Data in the Automotive UX Development Lifecycle: An Industry Perspective
abstract
We are interested in the role of field user interaction data in the development of In-Vehicle Information System (IVIS), the potentials practitioners see in analyzing this data, the concerns they share, and how this compares to companies with digital products. We conducted interviews with 14 UX professionals, 8 from automotive and 6 from digital companies, and analyzed the results by emergent thematic coding. Our key findings indicate that implicit feedback through field user interaction data is currently not evident in the automotive UX development process. Most decisions regarding the design of IVIS are made based on personal preferences and the intuitions of stakeholders. However, the interviewees also indicated that user interaction data has the potential to lower the influence of guesswork and assumptions in the UX design process and can help to make the UX development lifecycle more evidence-based and user-centered.
Patrick Ebel 0001, Florian Brokhausen, Andreas Vogelsang
AutomotiveUI3
2020 What Makes Agile Test Artifacts Useful?: An Activity-Based Quality Model from a Practitioners' Perspective
abstract
Background: The artifacts used in Agile software testing and the reasons why these artifacts are used are fairly well-understood. However, empirical research on how Agile test artifacts are eventually designed in practice and which quality factors make them useful for software testing remains sparse. Aims: Our objective is two-fold. First, we identify current challenges in using test artifacts to understand why certain quality factors are considered good or bad. Second, we build an Activity-Based Artifact Quality Model that describes what Agile test artifacts should look like. Method: We conduct an industrial survey with 18 practitioners from 12 companies operating in seven different domains. Results: Our analysis reveals nine challenges and 16 factors describing the quality of six test artifacts from the perspective of Agile testers. Interestingly, we observed mostly challenges regarding language and traceability, which are well-known to occur in non-Agile projects. Conclusions: Although Agile software testing is becoming the norm, we still have little confidence about general do's and don'ts going beyond conventional wisdom. This study is the first to distill a list of quality factors deemed important to what can be considered as useful test artifacts.
Jannik Fischbach, Henning Femmer, Daniel Méndez 0001, Davide Fucci, Andreas Vogelsang
ESEM5
2020 SPECMATE: Automated Creation of Test Cases from Acceptance Criteria
abstract
In the agile domain, test cases are derived from acceptance criteria to verify the expected system behavior. However, the design of test cases is laborious and has to be done manually due to missing tool support. Existing approaches for automatically deriving tests require semi-formal or even formal notations of acceptance criteria, though informal descriptions are mostly employed in practice. In this paper, we make three contributions: (1) a case study of 961 user stories providing an insight into how user stories are formulated and used in practice, (2) an approach for the automatic extraction of test cases from informal acceptance criteria and (3) a study demonstrating the feasibility of our approach in cooperation with our industry partner. In our study, out of 604 manually created test cases, 56 % can be generated automatically and missing negative test cases are added.
Jannik Fischbach, Andreas Vogelsang, Dominik Spies, Andreas Wehrle, Maximilian Junker, Dietmar Freudenstein
ICST2
2020 Destination Prediction Based on Partial Trajectory Data
abstract
Two-thirds of the people who buy a new car prefer to use a substitute instead of the built-in navigation system. However, for many applications, knowledge about a user's intended destination and route is crucial. For example, suggestions for available parking spots close to the destination can be made or ride-sharing opportunities along the route are facilitated. Our approach predicts probable destinations and routes of a vehicle, based on the most recent partial trajectory and additional contextual data. The approach follows a three-step procedure: First, a k-d tree-based space discretization is performed, mapping GPS locations to discrete regions. Secondly, a recurrent neural network is trained to predict the destination based on partial sequences of trajectories. The neural network produces destination scores, signifying the probability of each region being the destination. Finally, the routes to the most probable destinations are calculated. To evaluate the method, we compare multiple neural architectures and present the experimental results of the destination prediction. The experiments are based on two public datasets of non-personalized, timestamped GPS locations of taxi trips. The best performing models were able to predict the destination of a vehicle with a mean error of 1.3 km and 1.43 km respectively.
Patrick Ebel 0001, Ibrahim Emre Göl, Christoph Lingenfelder, Andreas Vogelsang
IV4
2020 Towards Causality Extraction from Requirements
abstract
System behavior is often based on causal relations between certain events (e.g. If event1, then event2). Consequently, those causal relations are also textually embedded in requirements. We want to extract this causal knowledge and utilize it to derive test cases automatically and to reason about dependencies between requirements. Existing NLP approaches fail to extract causality from natural language (NL) with reasonable performance. In this paper, we describe first steps towards building a new approach for causality extraction and contribute: (1) an NLP architecture based on Tree Recursive Neural Networks (TRNN) that we will train to identify causal relations in NL requirements and (2) an annotation scheme and a dataset that is suitable for training TRNNs. Our dataset contains 212,186 sentences from 463 publicly available requirement documents and is a first step towards a gold standard corpus for causality extraction. We encourage fellow researchers to contribute to our dataset and help us in finalizing the causality annotation process. Additionally, the dataset can also be annotated further to serve as a benchmark for other RE-relevant NLP tasks such as requirements classification.
Jannik Fischbach, Benedikt Hauptmann, Lukas Konwitschny, Dominik Spies, Andreas Vogelsang
RE5
2020 Trace Link Recovery using Semantic Relation Graphs and Spreading Activation
abstract
Trace Link Recovery tries to identify and link related existing requirements with each other to support further engineering tasks. Existing approaches are mainly based on algebraic Information Retrieval or machine-learning. Machinelearning approaches usually demand reasonably large and labeled datasets to train. Algebraic Information Retrieval approaches like distance between tf-idf scores also work on smaller datasets without training but are limited in providing explanations for trace links. In this work, we present a Trace Link Recovery approach that is based on an explicit representation of the content of requirements as a semantic relation graph and uses Spreading Activation to answer trace queries over this graph. Our approach is fully automated including an NLP pipeline to transform unrestricted natural language requirements into a graph. We evaluate our approach on five common datasets. Depending on the selected configuration, the predictive power strongly varies. With the best tested configuration, the approach achieves a mean average precision of 40% and a Lag of 50%. Even though the predictive power of our approach does not outperform state-of-the-art approaches, we think that an explicit knowledge representation is an interesting artifact to explore in Trace Link Recovery approaches to generate explanations and refine results.
Aaron Schlutter, Andreas Vogelsang
RE2
2020 Data-driven Risk Management for Requirements Engineering: An Automated Approach based on Bayesian Networks
abstract
Requirements Engineering (RE) is a means to reduce the risk of delivering a product that does not fulfill the stakeholders' needs. Therefore, a major challenge in RE is to decide how much RE is needed and what RE methods to apply. The quality of such decisions is strongly based on the RE expert's experience and expertise in carefully analyzing the context and current state of a project. Recent work, however, shows that lack of experience and qualification are common causes for problems in RE. We trained a series of Bayesian Networks on data from the NaPiRE survey to model relationships between RE problems, their causes, and effects in projects with different contextual characteristics. These models were used to conduct (1) a post-mortem (diagnostic) analysis, deriving probable causes of suboptimal RE performance, and (2) to conduct a preventive analysis, predicting probable issues a young project might encounter. The method was subject to a rigorous cross-validation procedure for both use cases before assessing its applicability to real-world scenarios with a case study.
Florian Wiesweg, Andreas Vogelsang, Daniel Méndez 0001
RE2
2020 How Do Quantifiers Affect the Quality of Requirements?
Katharina Winter, Henning Femmer, Andreas Vogelsang
REFSQ3
2020 What am I testing and where? Comparing testing procedures based on lightweight requirements annotations
abstract
Abstract Context The testing of software-intensive systems is performed in different test stages each having a large number of test cases. These test cases are commonly derived from requirements. Each test stages exhibits specific demands and constraints with respect to their degree of detail and what can be tested. Therefore, specific test suites are defined for each test stage. In this paper, the focus is on the domain of embedded systems, where, among others, typical test stages are Software- and Hardware-in-the-loop. Objective Monitoring and controlling which requirements are verified in which detail and in which test stage is a challenge for engineers. However, this information is necessary to assure a certain test coverage, to minimize redundant testing procedures, and to avoid inconsistencies between test stages. In addition, engineers are reluctant to state their requirements in terms of structured languages or models that would facilitate the relation of requirements to test executions. Method With our approach, we close the gap between requirements specifications and test executions. Previously, we have proposed a lightweight markup language for requirements which provides a set of annotations that can be applied to natural language requirements. The annotations are mapped to events and signals in test executions. As a result, meaningful insights from a set of test executions can be directly related to artifacts in the requirements specification. In this paper, we use the markup language to compare different test stages with one another. Results We annotate 443 natural language requirements of a driver assistance system with the means of our lightweight markup language. The annotations are then linked to 1300 test executions from a simulation environment and 53 test executions from test drives with human drivers. Based on the annotations, we are able to analyze how similar the test stages are and how well test stages and test cases are aligned with the requirements. Further, we highlight the general applicability of our approach through this extensive experimental evaluation. Conclusion With our approach, the results of several test levels are linked to the requirements and enable the evaluation of complex test executions. By this means, practitioners can easily evaluate how well a systems performs with regards to its specification and, additionally, can reason about the expressiveness of the applied test stage.
Florian Pudlitz, Florian Brokhausen, Andreas Vogelsang
Empir. Softw. Eng.3
2020 Views on quality requirements in academia and practice: commonalities, differences, and context-dependent grey areas
Andreas Vogelsang, Jonas Eckhardt, Daniel Méndez 0001, Moritz Berger
Inf. Softw. Technol.1
2020 Feature dependencies in automotive software systems: Extent, awareness, and refactoring
Andreas Vogelsang
J. Syst. Softw.1
2019 Extraction of System States from Natural Language Requirements
abstract
In recent years, simulations have proven to be an important means to verify the behavior of complex software systems. The different states of a system are monitored in the simulations and are compared against the requirements specification. So far, system states in natural language requirements cannot be automatically linked to signals from the simulation. However, the manual mapping between requirements and simulation is a time-consuming task. Named-entity Recognition is a sub-task from the field of automated information retrieval and is used to classify parts of natural language texts into categories. In this paper, we use a self-trained Named-entity Recognition model with Bidirectional LSTMs and CNNs to extract states from requirements specifications. We present an almost entirely automated approach and an iterative semi-automated approach to train our model. The automated and iterative approach are compared and discussed with respect to the usual manual extraction. We show that the manual extraction of states in 2,000 requirements takes nine hours. Our automated approach achieves an F1-score of 0.51 with 15 minutes of manual work and the iterative approach achieves an F1-score of 0.62 with 100 minutes of work.
Florian Pudlitz, Florian Brokhausen, Andreas Vogelsang
RE3
2019 Optimizing for Recall in Automatic Requirements Classification: An Empirical Study
abstract
Using Machine Learning to solve requirements engineering problems can be a tricky task. Even though certain algorithms have exceptional performance, their recall is usually below 100%. One key aspect in the implementation of machine learning tools is the balance between recall and precision. Tools that do not find all correct answers may be considered useless. However, some tasks are very complicated and even requirements engineers struggle to solve them perfectly. If a tool achieves performance comparable to a trained engineer while reducing her workload considerably, it is considered to be useful. One such task is the classification of specification content elements into requirements and non-requirements. In this paper, we analyze this specific requirements classification problem and assess the importance of recall by performing an empirical study. We compared two groups of students who performed this task with and without tool support, respectively. We use the results to compute an estimate of β for the Fβscore, allowing us to choose the optimal balance between precision and recall. Furthermore, we use the results to assess the practical time savings realized by the approach. By using the tool, users may not be able to find all defects in a document, however, they will be able to find close to all of them in a fraction of the time necessary. This demonstrates the practical usefulness of our approach and machine learning tools in general.
Jonas Winkler, Jannis Grönberg, Andreas Vogelsang
RE3
2019 Predicting How to Test Requirements: An Automated Approach
abstract
An important task in requirements engineering is to identify and determine how to verify a requirement (e.g., by manual review, testing, or simulation; also called potential verification method). This information is required to effectively create test cases and verification plans for requirements. [Objective] In this paper, we propose an automatic approach to classify natural language requirements with respect to their potential verification methods (PVM). [Method] Our approach uses a convolutional neural network architecture to implement a multiclass and multilabel classifier that assigns probabilities to a predefined set of six possible verification methods, which we derived from an industrial guideline. Additionally, we implemented a backtracing approach to analyze and visualize the reasons for the network's decisions. [Results] In a 10-fold cross validation on a set of about 27,000 industrial requirements, our approach achieved a macro averaged F1score of 0.79 across all labels. For the classification into test or non-test, the approach achieves an even higher F1score of 0.94. [Conclusions] The results show that our approach might help to increase the quality of requirements specifications with respect to the PVM attribute and guide engineers in effectively deriving test cases and verification plans.
Jonas Winkler, Jannis Grönberg, Andreas Vogelsang
RE3
2019 A Lightweight Multilevel Markup Language for Connecting Software Requirements and Simulations
Florian Pudlitz, Andreas Vogelsang, Florian Brokhausen
REFSQ2
2019 Guest editorial: special section on artificial intelligence for requirements engineering
Eduard C. Groen, Rachel Harrison, Pradeep K. Murukannaiah, Andreas Vogelsang
Autom. Softw. Eng.4
2019 Artefacts in software engineering: a fundamental positioning
Daniel Méndez 0001, Wolfgang Böhm 0002, Andreas Vogelsang, Jakob Mund, Manfred Broy, Marco Kuhrmann, Thorsten Weyer
Softw. Syst. Model.3
2018 Information Extraction from High-level Activity Diagrams to Support Development Tasks
Martin Beckmann 0002, Thomas Karbe, Andreas Vogelsang
MODELSWARD3
2018 Automatic Glossary Term Extraction from Large-Scale Requirements Specifications
abstract
Creating glossaries for large corpora of requirments is an important but expensive task. Glossary term extraction methods often focus on achieving a high recall rate and, therefore, favor linguistic proecssing for extracting glossary term candidates and neglect the benefits from reducing the number of candidates by statistical filter methods. However, especially for large datasets a reduction of the likewise large number of candidates may be crucial. This paper demonstrates how to automatically extract relevant domain-specific glossary term candidates from a large body of requirements, the CrowdRE dataset. Our hybrid approach combines linguistic processing and statistical filtering for extracting and reducing glossary term candidates. In a twofold evaluation, we examine the impact of our approach on the quality and quantity of extracted terms. We provide a ground truth for a subset of the requirements and show that a substantial degree of recall can be achieved. Furthermore, we advocate requirements coverage as an additional quality metric to assess the term reduction that results from our statistical filters. Results indicate that with a careful combination of linguistic and statistical extraction methods, a fair balance between later manual efforts and a high recall rate can be achieved.
Tim Gemkow, Miro Conzelmann, Kerstin Hartig, Andreas Vogelsang
RE4
2018 Coexisting Graphical and Structured Textual Representations of Requirements: Insights and Suggestions
Martin Beckmann 0002, Christian Reuter 0002, Andreas Vogelsang
REFSQ3
2018 Using Tools to Assist Identification of Non-requirements in Requirements Specifications - A Controlled Experiment
Jonas Winkler, Andreas Vogelsang
REFSQ2
2017 Removal of Redundant Elements within UML Activity Diagrams
abstract
As the complexity of systems continues to rise, the use of model-driven development approaches becomes more widely applied. Still, many created models are mainly used for documentation. As such, they are not designed to be used in following stages of development, but merely as a means of improved overview and communication. In an effort to use existing UML2 activity diagrams of an industry partner (Daimler AG) as a source for automatic generation of software artifacts, we discovered, that the diagrams often contain multiple instances of the same element. These redundant instances might improve the readability of a diagram. However, they complicate further approaches such as automated model analysis or traceability to other artifacts because mostly redundant instances must be handled as one distinctive element. In this paper, we present an approach to automatically remove redundant ExecutableNodes within activity diagrams as they are used by our industry partner. The removal is implemented by merging the redundant instances to a single element and adding additional elements to maintain the original behavior of the activity. We use reachability graphs to argue that our approach preserves the behavior of the activity. Additionally, we applied the approach to a real system described by 36 activity diagrams. As a result 25 redundant instances were removed from 15 affected diagrams.
Martin Beckmann 0002, Vanessa N. Michalke, Aaron Schlutter, Andreas Vogelsang
MoDELS4
2017 "What Does My Classifier Learn?" A Visual Approach to Understanding Natural Language Text Classifiers
Jonas Winkler, Andreas Vogelsang
NLDB2
2017 Should I Stay or Should I Go? - On Forces that Drive and Prevent MBSE Adoption in the Embedded Systems Industry
Andreas Vogelsang, Tiago Amorim 0001, Florian Pudlitz, Peter Gersing, Jan Philipps
PROFES1
2017 A Case Study on a Specification Approach Using Activity Diagrams in Requirements Documents
abstract
Rising complexity of systems has long been a major challenge in requirements engineering. This manifests in more extensive and harder to understand requirements documents. At the Daimler AG, an approach is applied that combines the use of activity diagrams with natural language specifications to specify system functions. The approach starts with an activity diagram that is created to get an early overview. The contained information is then transferred to a textual requirements document, where details are added and the behavior is refined. While the approach aims to reduce efforts needed to understand a system's behavior, the application of the approach itself causes new challenges on its own. By examining existing specifications at Daimler, we identified nine categories of inconsistencies and deviations between activity diagrams and their textual representations. In a case study, we examined one system in detail to assess how often these occur. In a follow-up survey, we presented instances of the categories to different stakeholders of the system and let them asses the categories regarding their severity. Our analysis indicates that a coexistence of textual and graphical representations of models without proper tool support results in inconsistencies and deviations that may cause severe maintenance costs or even provoke faults in subsequent development steps.
Martin Beckmann 0002, Andreas Vogelsang, Christian Reuter 0002
RE2
2017 How do Practitioners Perceive the Relevance of Requirements Engineering Research? An Ongoing Study
abstract
The relevance of Requirements Engineering (RE) research to practitioners is a prerequisite for problem-driven research in the area and key for a long-term dissemination of research results to everyday practice. To understand better how industry practitioners perceive the practical relevance of RE research, we have initiated the RE-Pract project, an international collaboration conducting an empirical study. This project opts for a replication of previous work done in two different domains and relies on survey research. To this end, we have designed a survey to be sent to several hundred industry practitioners at various companies around the world and ask them to rate their perceived practical relevance of the research described in a sample of 418 RE papers published between 2010 and 2015 at the RE, ICSE, FSE, ESEC/FSE, ESEM and REFSQ conferences. In this paper, we summarize our research protocol and present the current status of our study and the planned future steps.
Xavier Franch, Daniel Méndez 0001, Marc Oriol, Andreas Vogelsang, Rogardt Heldal, Eric Knauss, Guilherme Horta Travassos, Jeffrey C. Carver, Óscar Dieste Tubío, Thomas Zimmermann 0001
RE4
2016 Are "non-functional" requirements really non-functional?: an investigation of non-functional requirements in practice
abstract
Non-functional requirements (NFRs) are commonly distinguished from functional requirements by differentiating how the system shall do something in contrast to what the system shall do. This distinction is not only prevalent in research, but also influences how requirements are handled in practice. NFRs are usually documented separately from functional requirements, without quantitative measures, and with relatively vague descriptions. As a result, they remain difficult to analyze and test. Several authors argue, however, that many so-called NFRs actually describe behavioral properties and may be treated the same way as functional requirements. In this paper, we empirically investigate this point of view and aim to increase our understanding on the nature of NFRs addressing system properties. We report on the classification of 530 NFRs extracted from 11 industrial requirements specifications and analyze to which extent these NFRs describe system behavior. Our results suggest that most "non-functional" requirements are not non-functional as they describe behavior of a system. Consequently, we argue that many so-called NFRs can be handled similarly to functional requirements.
Jonas Eckhardt, Andreas Vogelsang, Daniel Méndez 0001
ICSE2
2016 On the Distinction of Functional and Quality Requirements in Practice
Jonas Eckhardt, Andreas Vogelsang, Daniel Méndez 0001
PROFES2
2016 Challenging Incompleteness of Performance Requirements by Sentence Patterns
abstract
Performance requirements play an important role in software development. They describe system behavior that directly impacts the user experience. Specifying performance requirements in a way that all necessary content is contained, i.e., the completeness of the individual requirements, is challenging, yet project critical. Furthermore, it is still an open question, what content is necessary to make a performance requirement complete. To address this problem, we introduce a framework for specifying performance requirements. This framework (i) consists of a unified model derived from existing performance classifications, (ii) denotes completeness through a content model, and (iii) is operationalized through sentence patterns. We evaluate both the applicability of the framework as well as its ability uncover incompleteness with performance requirements taken from 11 industrial specifications. In our study, we were able to specify 86% of the examined performance requirements by means of our framework. Furthermore, we show that 68% of the specified performance requirements are incomplete with respect to our notion of completeness. We argue that our framework provides an actionable definition of completeness for performance requirements.
Jonas Eckhardt, Andreas Vogelsang, Henning Femmer, Philipp Mager
RE2
2016 Take Care of Your Modes! An Investigation of Defects in Automotive Requirements
Andreas Vogelsang, Henning Femmer
REFSQ1
2016 Characterizing Implicit Communal Components as Technical Debt in Automotive Software Systems
abstract
Automotive software systems are often characterized by a set of features that are implemented through a network of communicating components. It is common practice to implement or adapt features by an ad hoc (re) use of signals that originate from components of another feature. Thereby, over time some components become so-called implicit communal components. These components increase the necessary efforts for several development activities because they introduce feature dependencies. Refactoring implicit communal components reduces these efforts but also costs refactoring effort. In this paper, we provide empirical evidence that implicit communal components exist in industrial automotive systems. For two cases, we show that less than 10% of the components are responsible for more than 90% of the feature dependencies. Secondly, we propose a refactoring approach for implicit communal components, which makes them explicit by moving them to a dedicated platform component layer. Finally, we characterize implicit communal components as technical debt, which is a metaphor for suboptimal solutions having short-term benefits but causing a long-term negative impact. With this metaphor, we describe the trade-off between accepting the negative effects of implicit communal components and spending the necessary refactoring costs.
Andreas Vogelsang, Henning Femmer, Maximilian Junker
WICSA1
2015 How to Specify Non-Functional Requirements to Support Seamless Modeling? A Study Design and Preliminary Results
abstract
Context: Seamless model-based development provides integrated chains of models, covering all software engineering phases. Non-functional requirements (NFRs), like reusability, further play a vital role in software and systems engineering, but are often neglected in research and practice. It is still unclear how to integrate NFRs in a seamless model-based development. Goal: Our long-term goal is to develop a theory on the specification of NFRs such that they can be integrated in seamless model-based development. Method: Our overall study design includes a multi-staged procedure to infer an empirically founded theory on specifying NFRs to support seamless modeling. In this short paper, we present the study design and provide a discussion of (i) preliminary results obtained from a sample, and (ii) current issues related to the design. Results: Our study already shows significant fields of improvement, e.g., the low agreement during the classification. However, the results indicate to interesting points; for example, many of commonly used NFR classes concern system modeling concepts in a way that shows how blurry the borders between functional and NFRs are. Conclusions: We conclude so far that our overall study design seems suitable to obtain the envisioned theory in the long run, but we could also show current issues that are worth discussing within the empirical software engineering community. The main goal of this contribution is not to present and discuss current results only, but to foster discussions on the issues related to the integration of NFRs in seamless modeling in general and, in particular, discussions on open methodological issues.
Jonas Eckhardt, Daniel Méndez 0001, Andreas Vogelsang
ESEM3
2015 Systematic elicitation of mode models for multifunctional systems
abstract
Many requirements engineering approaches structure and specify requirements based on the notion of modes or system states. The set of all modes is usually considered as the mode model of a system or problem domain.
Andreas Vogelsang, Henning Femmer
RE1
2014 Supporting Concurrent Development of Requirements and Architecture - A Model-based Approach
abstract
A system’s requirements and its architecture are usually developed at least partly in parallel. This demands a continuous and automated assessment to confirm that the architecture conforms to its requirements. To enable such an assessment, the stepwise formalization of informal requirements has been proposed. However, there is no canonical set of artifacts and analysis techniques that has been evaluated for this task in practice yet. In this paper we propose an artifact model and a process that enables the continuous conformance assessment between requirements and architecture in a model-based context. We evaluate both in a development project with a group of students.
Andreas Vogelsang, Sebastian Eder, Georg Hackenberg, Maximilian Junker, Sabine Teufl
MODELSWARD1
2013 Why feature dependencies challenge the requirements engineering of automotive systems: An empirical study
abstract
Functional dependencies and feature interactions in automotive software systems are a major source of erroneous and deficient behavior. To overcome these problems, many approaches exist that focus on modeling these functional dependencies in early stages of system design. However, there are only few empirical studies that report on the extent of such dependencies in industrial software systems and how they are considered in an industrial development context. In this paper, we analyze the functional architecture of a real automotive software system with the aim to assess the extent, awareness and importance of interactions between features of a future vehicle. Our results show that within the functional architecture at least 85% of the analyzed vehicle features depend on each other. They furthermore show that the developers are not aware of a large number of these dependencies when they are modeled solely on an architectural level. Therefore, the developers mention the need for a more precise specification of feature interactions, e.g., for the execution of comprehensive impact analyses. These results challenge the current development methods and emphasize the need for an extensive modeling of features and their dependencies in requirements engineering.
Andreas Vogelsang, Steffen Fuhrmann
RE1
2012 Extent and characteristics of dependencies between vehicle functions in automotive software systems
abstract
Functional dependencies and feature interactions are a major source of erroneous and unwanted behavior in software-intensive systems. To overcome these problems, many approaches exist that focus on modeling these functional dependencies in advance, i.e., in the specification or the design of a system. However, there is little empirical data on the amount of such interactions between system functions in realistic systems. In this paper, we analyze structural models of a modern realistic automotive vehicle system with the aim to assess the extent and characteristics of interactions between system functions. Our results show that at least 69% of the analyzed system functions depend on each other or influence each other. These dependencies stretch all over the system whereby single system functions have dependencies to up to 40% of all system functions. These results challenge the current development methods and processes that treat system functions more or less as independent units of functionality.
Andreas Vogelsang, Stefan Teuchert, Jean-Francois Girard
MiSE1
2010 Software Metrics in Static Program Analysis
Andreas Vogelsang, Ansgar Fehnker, Ralf Huuck, Wolfgang Reif
ICFEM1