VLDB 2026 Research / reviewers in the wild / expert
Rudolf Ramler
dblp:95/4255
· DBLP profile ↗
64ranked-venue papers
20as first author
23since 2021 · last 2026
0000-0001-9903-6107ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 55 · 20 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 5 since 2021Systems, architecture and hardware · 5 · 1 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Drift Adaptation as Supervision Routing Under Heterogeneous Costs
Jorge Martinez-Gil, Florian Bachinger, Rudolf Ramler, Francois Picard, Leïla Belmerhnia, Georgios P. Spathoulas |
DEXA (2) | 3 |
| 2025 | Poster: Unit Testing Past vs. Present: Examining LLMs' Impact on Defect Detection and EfficiencyabstractThe integration of Large Language Models (LLMs), such as ChatGPT and GitHub Copilot, into software engineering workflows has shown potential to enhance productivity, particularly in software testing. This paper investigates whether LLM support improves defect detection effectiveness during unit testing. Building on prior studies comparing manual and tool-supported testing, we replicated and extended an experiment where participants wrote unit tests for a Java-based system with seeded defects within a time-boxed session, supported by LLMs. Comparing LLM supported and manual testing, results show that LLM support significantly increases the number of unit tests generated, defect detection rates, and overall testing efficiency. These findings highlight the potential of LLMs to improve testing and defect detection outcomes, providing empirical insights into their practical application in software testing. Rudolf Ramler, Philipp Straubinger, Reinhold Plösch, Dietmar Winkler 0001 |
ICST | 1 |
| 2025 | Assessing the strength of Metamorphic Testing applied to optimisation software - Experience from industry
Alejandra Duque-Torres, Claus Klammer, Stefan Fischer 0006, Dietmar Pfahl, Rudolf Ramler |
Inf. Softw. Technol. | 5 |
| 2025 | Metamorphic testing for optimisation: A case study on PID controller tuning
Alejandra Duque-Torres, Claus Klammer, Stefan Fischer 0006, Rudolf Ramler, Dietmar Pfahl |
Inf. Softw. Technol. | 4 |
| 2025 | Contemporary Software Modernization: Strategies, Driving Forces, and Research OpportunitiesabstractSoftware modernization is a common activity in software engineering, since technologies advance, requirements change, and business models evolve. Differently from conventional software evolution (e.g., adding new features, enhancing performance, or adapting to new requirements), software modernization involves re-engineering entire legacy systems (e.g., changing the technology stack, migrating to a new architecture style, or programming paradigms). Given the pervasive nature of software today, modernizing legacy systems is paramount to provide customers with competitive and innovative products and services, while keeping companies profitable. Despite the prevalent discussion of software modernization in gray literature, and the many papers in the literature, there is no work presenting a “big picture” of contemporary software modernization, describing challenges, and providing a well-defined research agenda. The goal of this work is to describe the state of the art in software modernization in the past 10 years. We collect the state of the art by performing a rapid review (searching five digital libraries), identifying potential 3,460 studies, leading to a final set of 126. We analyzed these studies to understand which strategies are employed, the driving forces that lead organizations to modernize their systems, and the challenges that need to be addressed. The results show that studies in the last 10 years have explored eight strategies for modernizing legacy systems, namely cloudification, architecture redesign, moving to a new programming language, targeting reuse optimization, software modernization for new hardware integration, practices to leverage automation, database modernization, and digital transformation. Modernization is triggered by 14 driving forces, with the most common ones being reducing operational costs, improving performance and scalability, and reducing complexity. In addition, based on the analysis of existing literature, we present a detailed discussion of research opportunities in this field. The main challenges are providing tooling support, followed by defining a modernization process and considering better evaluation metrics. The main contribution of our work is to equip practitioners and researchers with knowledge of the current state of contemporary software modernization so that they are aware of practices and challenges to be addressed when deciding to modernize legacy systems. Wesley K. G. Assunção, Luciano Marchezan, Lawrence Arkoh, Alexander Egyed, Rudolf Ramler |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2024 | An Overview of Microservice-Based Systems Used for Evaluation in Testing and Monitoring: A Systematic Mapping StudyabstractMicroservice-based systems have emerged as an effective architecture for countless industry applications. They provide applications as small, independent, and modular services. With the increasing interest in such systems, it is important to tackle challenges related to their quality assurance. However, to advance research in this area, systems are required to evaluate new approaches and tools. In this paper, we perform a systematic literature search for systems used in research for testing and monitoring microservice-based systems to aid future research. We provide an overview of the found studies and the systems used in their evaluation. We compose a list of publicly available systems and their characteristics, like size, available tests, and technologies used. Finally, we investigated the context in which these systems were used to provide insights in their usage and additional data that is available for them. Stefan Fischer 0006, Pirmin Urbanke, Rudolf Ramler, Monika Steidl, Michael Felderer |
AST | 3 |
| 2024 | The Metamorphic Lighthouse: Understanding the Input Data Space of Metamorphic RelationsabstractMetamorphic Testing (MT) addresses the test oracle problem by defining how program outputs should change in response to specific input changes. The relations between input changes and their corresponding output changes are called Metamorphic Relations (MRs). Generating suitable MRs is complex and often requires deep domain knowledge. Our previous work introduced MetaTrimmer, a test-data-driven approach for selecting and constraining MRs, involving three steps: Test Data (TD) Generation, MT Process, and MR Analysis. MR Analysis is done to decide whether the violation of an MR for a specific input data pair (original and changed) indicates a failure or simply means that the MR does not apply for the chosen inputs. In this paper, we present an association-rule-based approach that semi-automatically extracts constraints dividing the input space into valid/invalid data during the MR Analysis step of MetaTrimmer. We validate our approach using 44 methods to which six predefined MRs are applied. Our results indicate that the proposed method efficiently identifies correct input data space constraints. More studies are needed to provide additional evidence that MetaTrimmer with the enhanced MR Analysis step is scalable and generalisable. Alejandra Duque-Torres, Dietmar Pfahl, Claus Klammer, Stefan Fischer 0006, Rudolf Ramler |
SEAA | 5 |
| 2024 | How Industry Tackles Anomalies during Runtime: Approaches and Key Monitoring ParametersabstractDeviations from expected behavior during runtime, known as anomalies, have become more common due to the systems' complexity, especially for microservices. Consequently, analyzing runtime monitoring data, such as logs, traces for microservices, and metrics, is challenging due to the large volume of data collected. Developing effective rules or AI algorithms requires a deep understanding of this data to reliably detect unfore-seen anomalies. This paper seeks to comprehend anomalies and current anomaly detection approaches across diverse industrial sectors. Additionally, it aims to pinpoint the parameters necessary for identifying anomalies via runtime monitoring data. Therefore, we conducted semi-structured interviews with fifteen industry participants who rely on anomaly detection during runtime. Additionally, to supplement information from the interviews, we performed a literature review focusing on anomaly detection approaches applied to industrial real-life datasets. Our paper (1) demonstrates the diversity of interpretations and examples of software anomalies during runtime and (2) explores the reasons behind choosing rule-based approaches in the industry over self-developed AI approaches. AI-based approaches have become prominent in published industry-related papers in the last three years. Furthermore, we (3) identified key monitoring parameters collected during runtime (logs, traces, and metrics) that assist practitioners in detecting anomalies during runtime without introducing bias in their anomaly detection approach due to inconclusive parameters. Monika Steidl, Benedikt Dornauer, Michael Felderer, Rudolf Ramler, Mircea-Cristian Racasan, Marko Gattringer |
SEAA | 4 |
| 2024 | The Past, Present, and Future of Research on the Continuous Development of AIabstractSince 2020, 33 literature reviews have systematically synthesized research on the continuous development of AI, also known as Machine Learning Operations (MLOps), reflecting the increasing prevalence of AI models across various fields and the multifaceted challenges in their development, integration, and deployment. Yet, the lack of comprehensive analysis of these literature reviews and their covered topics complicates selecting relevant ones and anticipating future trends and research. In addition, these literature reviews gathered related 1397 primary sources to describe aspects of AI's continuous development, integration, and deployment, posing a hidden gem to gain insights into the past and present work and derive insights into the future of AI's continuous development. With this work, we 1) systematically collected and summarised 33 literature reviews via a Multivocal Literature Review (MLR) that focus on the continuous development, deployment, and integration of AI models. 2) Due to minimal overlap between the literature reviews' primary sources, we offer holistic insights into and interrelations of frequently addressed topics. These topics encompass the AI development pipeline, respective Software Engineering (SE) practices, and associated challenges. 3) We discuss future research directions for AI's continuous development, integration, and deployment. Therefore, we base our arguments on identified clusters in the primary sources of literature reviews. This discussion focuses on AI model reliability and resource consumption, emphasizing the interrelation of proposed future work and the effects on the whole pipeline. Monika Steidl, Rudolf Ramler, Michael Felderer |
SEAA | 2 |
| 2024 | Investigating the readability of test codeabstractAbstract Context The readability of source code is key for understanding and maintaining software systems and tests. Although several studies investigate the readability of source code, there is limited research specifically on the readability of test code and related influence factors. Objective In this paper, we aim at investigating the factors that influence the readability of test code from an academic perspective based on scientific literature sources and complemented by practical views, as discussed in grey literature. Methods First, we perform a Systematic Mapping Study (SMS) with a focus on scientific literature. Second, we extend this study by reviewing grey literature sources for practical aspects on test code readability and understandability. Finally, we conduct a controlled experiment on the readability of a selected set of test cases to collect additional knowledge on influence factors discussed in practice. Results The result set of the SMS includes 19 primary studies from the scientific literature for further analysis. The grey literature search reveals 62 sources for information on test code readability. Based on an analysis of these sources, we identified a combined set of 14 factors that influence the readability of test code. 7 of these factors were found in scientific and grey literature, while some factors were mainly discussed in academia (2) or industry (5) with only limited overlap. The controlled experiment on practically relevant influence factors showed that the investigated factors have a significant impact on readability for half of the selected test cases. Conclusion Our review of scientific and grey literature showed that test code readability is of interest for academia and industry with a consensus on key influence factors. However, we also found factors only discussed by practitioners. For some of these factors we were able to confirm an impact on readability in a first experiment. Therefore, we see the need to bring together academic and industry viewpoints to achieve a common view on the readability of software test code. Dietmar Winkler 0001, Pirmin Urbanke, Rudolf Ramler |
Empir. Softw. Eng. | 3 |
| 2024 | Data pipeline quality: Influencing factors, root causes of data-related issues, and processing problem areas for developersabstractData pipelines are an integral part of various modern data-driven systems. However, despite their importance, they are often unreliable and deliver poor-quality data. A critical step toward improving this situation is a solid understanding of the aspects contributing to the quality of data pipelines. Therefore, this article first introduces a taxonomy of 41 factors that influence the ability of data pipelines to provide quality data. The taxonomy is based on a multivocal literature review and validated by eight interviews with experts from the data engineering domain. Data, infrastructure, life cycle management, development & deployment, and processing were found to be the main influencing themes. Second, we investigate the root causes of data-related issues, their location in data pipelines, and the main topics of data pipeline processing issues for developers by mining GitHub projects and Stack Overflow posts. We found data-related issues to be primarily caused by incorrect data types (33%), mainly occurring in the data cleaning stage of pipelines (35%). Data integration and ingestion tasks were found to be the most asked topics of developers, accounting for nearly half (47%) of all questions. Compatibility issues were found to be a separate problem area in addition to issues corresponding to the usual data pipeline processing areas (i.e., data loading, ingestion, integration, cleaning, and transformation). These findings suggest that future research efforts should focus on analyzing compatibility and data type issues in more depth and assisting developers in data integration and ingestion tasks. The proposed taxonomy is valuable to practitioners in the context of quality assurance activities and fosters future research into data pipeline quality. Harald Foidl, Valentina Golendukhina, Rudolf Ramler, Michael Felderer |
J. Syst. Softw. | 3 |
| 2023 | Towards Automatic Generation of Amplified Regression Test OraclesabstractRegression testing is crucial in ensuring that pure code refactoring does not adversely affect existing software functionality, but it can be expensive, accounting for half the cost of software maintenance. Automated test case generation reduces effort but may generate weak test suites. Test amplification is a promising solution that enhances tests by generating additional or improving existing ones, increasing test coverage, but it faces the test oracle problem. To address this, we propose a test oracle derivation approach that uses object state data produced during System Under Test (SUT) test execution to amplify regression test oracles. The approach monitors the object state during test execution and compares it to the previous version to detect any changes in relation to the SUT’s intended behaviour. Our preliminary evaluation shows that the proposed approach can enhance the detection of behaviour changes substantially, providing initial evidence of its effectiveness. Alejandra Duque-Torres, Claus Klammer, Dietmar Pfahl, Stefan Fischer 0006, Rudolf Ramler |
SEAA | 5 |
| 2023 | Automation and Development Effort in Continuous AI Development: A Practitioners' SurveyabstractThe widespread adoption of AI-enabled systems and their required continuous development and deployment (MLOps) sparks research interest due to the added intricacy of automatically handling data, code, and the model itself. A better understanding of the stages for the continuous development of AI, namely Data Handling, Model Learning, Software Development, and System Operations, and the respective tasks can help to optimize and improve their effectiveness.Thus, this paper explores the degree of automation, development effort, importance, utilization of computing resources, and factors contributing to automation throughout these stages and tasks. We conducted a questionnaire-based global survey to explore these topics by analyzing 150 responses from experienced AI, data, and MLOps engineers.The results determined that the stage System Operations is mainly automated. Whereas several tasks from the other three stages (e.g., data cleaning, data quality assurance, model design, model improvement, and system level quality assurance) are more often partially automated than automated, and documentation-related tasks are mostly not automated or developed. Participants required the highest development effort for the stage Data Handling. Furthermore, the study reveals a negative correlation between automation and the perceived development effort, whereas the importance of the tasks does not seem to affect automation. 93% of participants consider the availability of computing resources, with model training, data transformation, and data cleaning ranked as the most resource-intensive tasks. Monika Steidl, Valentina Golendukhina, Michael Felderer, Rudolf Ramler |
SEAA | 4 |
| 2023 | Testing of highly configurable cyber-physical systems - Results from a two-phase multiple case study
Stefan Fischer 0006, Claus Klammer, Antonio Manuel Gutiérrez, Rick Rabiser, Rudolf Ramler |
J. Syst. Softw. | 5 |
| 2023 | The pipeline for the continuous development of artificial intelligence models - Current state of research and practiceabstractCompanies struggle to continuously develop and deploy Artificial Intelligence (AI) models to complex production systems due to AI characteristics while assuring quality. To ease the development process, continuous pipelines for AI have become an active research area where consolidated and in-depth analysis regarding the terminology, triggers, tasks, and challenges is required. This paper includes a Multivocal Literature Review (MLR) where we consolidated 151 relevant formal and informal sources. In addition, nine-semi structured interviews with participants from academia and industry verified and extended the obtained information. Based on these sources, this paper provides and compares terminologies for Development and Operations (DevOps) and Continuous Integration (CI)/Continuous Delivery (CD) for AI, Machine Learning Operations (MLOps), (end-to-end) lifecycle management, and Continuous Delivery for Machine Learning (CD4ML). Furthermore, the paper provides an aggregated list of potential triggers for reiterating the pipeline, such as alert systems or schedules. In addition, this work uses a taxonomy creation strategy to present a consolidated pipeline comprising tasks regarding the continuous development of AI. This pipeline consists of four stages: Data Handling, Model Learning, Software Development and System Operations. Moreover, we map challenges regarding pipeline implementation, adaption, and usage for the continuous development of AI to these four stages. Monika Steidl, Michael Felderer, Rudolf Ramler |
J. Syst. Softw. | 3 |
| 2023 | An automated evaluation of broker compatibility for the Message Queuing Telemetry Transport protocolabstractAbstract Message Queuing Telemetry Transport (MQTT) is the most widely used protocol within the communication layer of the Internet of Things (IoT). Message brokers are a key component of the MQTT protocol and a single point of failure. Incompatibilities between different MQTT brokers or broker versions with their clients can cause critical failures and become a source of security risks. Thus, every MQTT broker change or update needs to be accompanied by an evaluation of the compatibility between the new and the previous broker. In this work, we develop an automated framework for compatibility evaluation of MQTT brokers, which can be easily generalized to other similar IoT components. We apply this framework to perform a comprehensive experiment conducted with 16 different versions of 6 popular MQTT brokers. We report inconsistencies in the behavior of different MQTT brokers and broker versions. Based on the experiment results, we calculate and provide a visualization of compatibility among the evaluated brokers in terms of their distance, which indicates the risk of incompatibilities when replacing a broker with another one. The calculation of distance measures can be adjusted by giving higher weights to important features. We use this method to show security‐related differences between the brokers. Hannes Sochor, Flavio Ferrarotti, Rudolf Ramler |
J. Softw. Evol. Process. | 3 |
| 2022 | Data smells: categories, causes and consequences, and detection of suspicious data in AI-based systemsabstractHigh data quality is fundamental for today's AI-based systems. However, although data quality has been an object of research for decades, there is a clear lack of research on potential data quality issues (e.g., ambiguous, extraneous values). These kinds of issues are latent in nature and thus often not obvious. Nevertheless, they can be associated with an increased risk of future problems in AI-based systems (e.g., technical debt, data-induced faults). As a counterpart to code smells in software engineering, we refer to such issues as Data Smells. This article conceptualizes data smells and elaborates on their causes, consequences, detection, and use in the context of AI-based systems. In addition, a catalogue of 36 data smells divided into three categories (i.e., Believability Smells, Understandability Smells, Consistency Smells) is presented. Moreover, the article outlines tool support for detecting data smells and presents the result of an initial smell detection on more than 240 real-world datasets. Harald Foidl, Michael Felderer, Rudolf Ramler |
CAIN | 3 |
| 2022 | Requirements for Anomaly Detection Techniques for Microservices
Monika Steidl, Marko Gattringer, Michael Felderer, Rudolf Ramler, Mostafa Shahriari |
PROFES | 4 |
| 2022 | A Replication Study on Predicting Metamorphic Relations at Unit Testing LevelabstractMetamorphic Testing (MT) addresses the test oracle problem by examining the relations between inputs and outputs of test executions. Such relations are known as Metamorphic Relations (MRs). In current practice, identifying and selecting suitable MRs is usually a challenging manual task, requiring a thorough grasp of the SUT and its application domain. Thus, Kanewala et al. proposed the Predicting Metamorphic Relations (PMR) approach to automatically suggest MRs from a list of six pre-defined MRs for testing newly developed methods. PMR is based on a classification model trained on features extracted from the control-flow graph (CFG) of 100 Java methods. In our replication study, we explore the generalizability of PMR. First, since not all details necessary for a replication are provided, we rebuild the entire preprocessing and training pipeline and repeat the original study in a close replication to verify the reported results and establish the basis for further experiments. Second, we perform a conceptual replication to explore the reusability of the PMR model trained on CFGs from Java methods in the first step for functionally identical methods implemented in Python and C++. Finally, we retrain the model on the CFGs from the Python and C++ methods to investigate the dependence on programming language and implementation details. We were able to successfully replicate the original study achieving comparable results for the Java methods set. However, the prediction performance of the Java-based classifiers significantly decreases when applied to functionally equivalent Python and C++ methods despite using only CFG features to abstract from language details. Since the performance improved again when the classifiers were retrained on the CFGs of the methods written in Python and C++, we conclude that the PMR approach can be generalized, but only when classifiers are developed starting from code artefacts in the used programming language. Alejandra Duque-Torres, Dietmar Pfahl, Rudolf Ramler, Claus Klammer |
SANER | 3 |
| 2022 | What Do We Know About Readability of Test Code? - A Systematic Mapping StudyabstractThe readability of software code is a key success criterion for understanding and maintaining software systems and tests. In industry practice, a limited number of guidelines aim for improving and assessing the readability of software (test) code. Although several studies focus on investigating the readability of software code, we observed limited research work that focuses on the readability of software test code. In this paper we focus on systematically investigating the characteristics, factors, and assessment criteria that have an impact on the readability of test code. We build on a Systematic Mapping Study (SMS) to identify key characteristics, factors, and assessment criteria that have an impact on test code readability, legibility, and understandability to support and improve maintenance tasks. The result set includes 16 studies for further analysis. The majority of publications focuses on readability investigations of automatically generated test code (88%), often evaluated with surveys to access the readability of test code (44 %). Although several approaches aim at assessing the readability with focus on isolated factors, a combination of different readability aspects within an assessment framework can help to better assess and justify the readability of test code with focus on improving software and system test maintenance. Dietmar Winkler 0001, Pirmin Urbanke, Rudolf Ramler |
SANER | 3 |
| 2021 | Comparing Automated Reuse of Scripted Tests and Model-Based Tests for Configurable SoftwareabstractHighly configurable software gives developers more flexibility to meet different customer requirements and enables users to better tailor software to their needs. However, variability causes higher complexity in software and complicates many development processes, such as testing. One major challenge for testing of configurable software is adjusting tests to fit different configurations, which often has to be done manually. In our previous work, we evaluated the use of an automated reuse technique to support the reuse of existing tests for new configurations. Research on automated reuse of model variants and on applying model-based testing to configurable software encouraged us to also evaluate the automated reuse of model-based test variants. The goal is to investigate differences in applying automated reuse to the different testing paradigms. Our evaluation provides evidence for the usefulness of automated reuse for both testing paradigms. Nonetheless we found some differences in the robustness of tests to small inaccuracies of the reuse approach. Stefan Fischer 0006, Rudolf Ramler, Lukas Linsbauer |
APSEC | 2 |
| 2021 | Adopting Microservices for Industrial Control Systems: A Five Step Migration PathabstractMicroservices are widely used by large internet companies as they support scalable systems with high resilience and fault-tolerance, flexible and agile development, and continuous delivery enabling fast time to market. Hence, there is an increasing interest in adopting microservices in the field of industrial automation. This raises the question, if and to what degree this architectural style can also be applied for the development of industrial control systems (ICS). In this paper, we have systematically analyzed the applicability of microservices for ICS development. Together with domain experts from industry, we have developed a migration path from a monolithic ICS towards cloud-ready systems based on microservices. By studying the central principles for microservice development and operation, we found that microservices can be applied in the context of ICS and the use of microservices leads to increased flexibility with regard to frequent software releases and the development of new deployment variants. However, communication between real-time services is still an open research challenge that poses a potential technical risk in the migration towards adopting full microservice-based system architectures. Georg Buchgeher, Rudolf Ramler, Heinz Stummer, Hannes Kaufmann |
ETFA | 2 |
| 2021 | Test Benchmarks: Which One Now and in Future?abstractTo evaluate software testing and program analysis tools, the research community relies on collections of sample programs (benchmarks) containing realistic code examples with defects. We investigated 23 benchmark projects for test generation in common programming languages and looked at how they can be categorized according to attributes such as programming language, number of programs and defects, license, and size. From our studies, it is evident that the development and especially maintenance of benchmarks are a big challenge. Out of the 23 benchmark projects we investigated, only four are still active as of today, and only nine have been updated after their initial release. With the underlying programming languages and platforms constantly evolving (often without full backward compatibility), this creates a challenge when comparing new tools to older ones. To exacerbate the situation, many benchmarks do not fully track the provenance and license of the code they include. Sustainable benchmark collections share these key factors: Open hosting of complete (actual) data allowing community involvement, systematic maintenance of license and authorship data, and a unified machine-readable format for such data. Cyrille Artho, Adam Benali, Rudolf Ramler |
QRS | 3 |
| 2020 | Automated security test generation for MQTT using attack patternsabstractThe dramatic increase of attacks and malicious activities has made security a major concern in the development of interconnected cyber-physical systems and raised the need to address this concern also in testing. The goal of security testing is to discover vulnerabilities in the system under test so that they can be fixed before an attacker finds and abuses them. However, testing for security issues faces the challenge of systematically exploring a potentially non-tractable number of interaction scenarios that have to include also invalid inputs and possible harmful interaction attempts. In this paper, we describe an approach for automated generation of test cases for security testing, which are based on attack patterns. These patterns are blueprints that can be used for exploiting common vulnerabilities. The approach combines random test case generation with attack patterns implemented for the Message Queuing Telemetry Transport (MQTT) protocol. We have applied the proposed testing approach to five popular and widely available MQTT brokers, generating 1,804 interaction sequences in form of executable test cases which resulted in numerous test failures, unhandled exceptions and crashes. A detailed manual analysis of these cases have revealed 28 security-relevant issues and critical shortcomings in the tested MQTT broker implementations. Hannes Sochor, Flavio Ferrarotti, Rudolf Ramler |
ARES | 3 |
| 2020 | Applying AI in Practice: Key Challenges and Lessons Learned
Lukas Fischer 0001, Lisa Ehrlinger, Verena Geist, Rudolf Ramler, Florian Sobieczky, Werner Zellinger, Bernhard Moser 0001 |
CD-MAKE | 4 |
| 2020 | An Expert Review on the Applicability of Blockly for Industrial Robot ProgrammingabstractThe paradigm shift triggered by Industry 4.0 leads to a fast rising number of industrial machinery and collaborative robots that increases the need for flexible customization of production processes and automation workflows. End-user programming of industrial robots has become an essential capability for all areas in industry. In this paper, we investigate the applicability of block-based programming languages for large and complex robot programs. Therefore, we implemented a real-world program for industrial robots using Blockly, a block-based visual language, and assessed the results regarding portability, readability, understandability, and maintainability. Our findings are, that (1) large and complex real-world robot programs can be expressed in Blockly, (2) that the readability and understandability of such programs is equal to conventional flow-chart based programming languages, and (3) that the maintainability of Blockly programs can even be considered better than in flowchart based languages. Our results and findings serve as basis for usability studies with end-users in robot programming. Mario Winterer, Christian Salomon, Jessica Köberle, Rudolf Ramler, Markus Schittengruber |
ETFA | 4 |
| 2020 | Live Replay of Screen Videos: Automatically Executing Real Applications as Shown in RecordingsabstractScreencasts and videos with screen recordings are becoming an increasingly popular source of information for users to understand and learn about software applications. However, searching for answers to specific questions in screen videos is notoriously difficult due to the effort for locating specific events of interest and reproducing the application's state up to this event. To increase the efficiency when working with screen videos, we propose a solution for replaying recorded sequences shown in videos directly on live applications. In this paper, we describe the analysis of screen videos to automatically identify and extract user interactions and the construction of visual scripts, which are used to run the application in sync with replaying the video. Currently, a first prototype has been developed to demonstrate the technical feasibility of the approach. The paper provides an overview of the implemented solution concept and discusses technical challenges, open issues, as well as future application scenarios. Rudolf Ramler, Marko Gattringer, Josef Pichler |
SANER | 1 |
| 2020 | Automated test reuse for highly configurable software
Stefan Fischer 0006, Gabriela Karoline Michelon, Rudolf Ramler, Lukas Linsbauer, Alexander Egyed |
Empir. Softw. Eng. | 3 |
| 2018 | Adapting automated test generation to GUI testing of industry applications
Rudolf Ramler, Georg Buchgeher, Claus Klammer |
Inf. Softw. Technol. | 1 |
| 2017 | Exploring code clones in programmable logic controller softwareabstractThe reuse of code fragments by copying and pasting is widely practiced in software development and results in code clones. Cloning is considered an anti-pattern as it negatively affects program correctness and increases maintenance efforts. Programmable Logic Controller (PLC) software is no exception in the code clone discussion as reuse in development and maintenance is frequently achieved through copy, paste, and modification. Even though the presence of code clones may not necessary be a problem per se, it is important to detect, track and manage clones as the software system evolves. Unfortunately, tool support for clone detection and management is not commonly available for PLC software systems or limited to generic tools with a reduced set of features. In this paper, we investigate code clones in a real-world PLC software system based on IEC 61131-3 Structured Text and C/C++. We extended a widely used tool for clone detection with normalization support. Furthermore, we evaluated the different types and natures of code clones in the studied system and their relevance for refactoring. Results shed light on the applicability and usefulness of clone detection in the context of industrial automation systems and it demonstrates the benefit of adapting detection and management tools for IEC 611313-3 languages. Hannes Thaller, Rudolf Ramler, Josef Pichler, Alexander Egyed |
ETFA | 2 |
| 2017 | How to Test in Sixteen Languages? Automation Support for Localization TestingabstractDeveloping for a global market requires the internationalization of software products and their localization to different countries, regions, and cultures. Localization testing verifies that the localized software variants work, look and feel as expected. Localization testing is a perfect candidate for automation. It has a high potential to reduce the manual effort in testing of multiple language variants and to speed-up release cycles. However, localization testing is rarely investigated in scientific work. There are only a few reports on automation approaches for localization testing providing very little empirical results or practical advice. In this paper we describe the approach we applied for automated testing of the different localized variants of a large industrial software system, we report on the various bugs found, and we discuss our experiences and lessons learned. Rudolf Ramler, Robert Hoschek |
ICST | 1 |
| 2017 | Process and Tool Support for Internationalization and Localization Testing in Software Product Development
Rudolf Ramler, Robert Hoschek |
PROFES | 1 |
| 2017 | Special issue on collaboration in software testing between industry and academia
Michael Felderer, Rudolf Ramler |
Softw. Qual. J. | 2 |
| 2017 | Static Code Analysis of IEC 61131-3 Programs: Comprehensive Tool Support and Experiences from Large-Scale Industrial ApplicationabstractStatic code analysis techniques examine programs without actually executing them. The main benefits lie in improving software quality by detecting problematic code constructs and potential defects in early development stages. Today, static code analysis is a widely used quality assurance technique and numerous tools are available for established programming languages like C/C++, Java, or C#. However, in the domain of programmable logic controller (PLC) programming, static code analysis tools are still rare, although many properties of PLC programming languages are beneficial for static analysis techniques. Therefore, an approach and tool for static code analysis of IEC 61131-3 programs has been developed which is capable of detecting a range of issues commonly occurring in PLC programming. The approach employs different analysis methods, like pattern-matching on program structures, control flow and data flow analyses, and, especially, call graph and pointer analysis techniques. Based on results from an initial analysis project, where common issues for static analysis of PLC programs have been investigated, this paper illustrates adoption and extensions of analysis techniques for PLC programs and presents results from large-scale industrial application. Herbert Prähofer, Florian Angerer, Rudolf Ramler, Friedrich Grillenberger |
IEEE Trans. Ind. Informatics | 3 |
| 2016 | Harnessing Automated Test Case Generators for GUI Testing in IndustryabstractModern graphical user interfaces (GUIs) are highly dynamic and support multi-touch interactions and screen gestures besides conventional inputs via mouse and keyboard. Hence, the flexibility of modern GUIs enables countless usage scenarios and combinations including all kind of interactions. From the viewpoint of testing, this flexibility results in a combinatorial explosion of possible interaction sequences. It dramatically raises the required time and effort involved in GUI testing, which brings manual exploration as well as conventional regression testing approaches to its limits. Automated test generation (ATG) has been proposed as a solution to reduce the effort for manually designing test cases and to speed-up test execution cycles. In this paper we describe how we successfully harnessed a state-of-the-art ATG tool (Randoop) developed for code-based API testing to generate GUI test cases. The key is an adapter that transforms API calls to GUI events. The approach is the result of a research transfer project with the goal to apply ATG for testing of human machine interfaces used to control industrial machinery. In this project the ATG tool was used to generate unit test cases for custom GUI controls and system tests for exploring navigation scenarios. It helped to increase the test coverage and was able reveal new defects in the implementation of the GUI controls as well as in the GUI application. Claus Klammer, Rudolf Ramler, Heinz Stummer |
SEAA | 2 |
| 2016 | Analyzing Performance Issues of Industrial User Interfaces: Experiences and ResultsabstractUser interfaces of modern industrial automation systems are required to provide similar user experience as known from mass market consumer devices like smartphones. They need to support an increasing amount of dynamic visualizations as well as multi-touch interactions and screen gestures. However, the rich visualization frameworks that are applied and the elaborated user interfaces built on top have demanding hardware requirements, which can result in critical performance issues when run on resource-limited target platforms. In this paper we share our experiences and findings from analyzing the performance issues of a state-of-the-art user interface based on JavaFX. We describe the performance analysis approach, the tools we used for measurement and analysis, as well as insights gained into the JavaFX platform. We started with measuring and optimizing the performance of a concrete implementation of a user interface. Our results triggered immediate improvement actions, but they also led to the plan to develop of a performance model and corresponding tool support for predicting potential performance issues at design time, before the user interface is completed and deployed to the target hardware. Claus Klammer, Rudolf Ramler, Mario Winterer, Heinz Stummer |
SEAA | 2 |
| 2016 | Requirements for Integrating Defect Prediction and Risk-Based TestingabstractDefect prediction is a powerful method that provides information about the likely defective parts in software system and is applicable to improve effectiveness and efficiency of software quality assurance. This makes defect prediction a perfect candidate to be combined with risk-based testing to optimally guide testing activities towards risky parts of software. As a first step towards a successful combination, this paper presents requirements that have to be fulfilled for enabling the synergies between defect prediction and risk-based testing. Rudolf Ramler, Michael Felderer |
SEAA | 1 |
| 2016 | A Framework for Monkey GUI TestingabstractTesting via graphical user interfaces (GUI) is a complex and labor-intensive task. Numerous techniques, tools and frameworks have been proposed for automating GUI testing. In many projects, however, the introduction of automated tests did not reduce the overall effort of testing but shifted it from manual test execution to test script development and maintenance. As a pragmatic solution, random testing approaches (aka "monkey testing") have been suggested for automated random exploration of the system under test via the GUI. This paper presents a versatile framework for monkey GUI testing. The framework provides reusable components and a predefined, generic workflow with extension points for developing custom-built test monkeys. It supports tailoring the monkey for a particular application scenario and the technical requirements imposed by the system under test. The paper describes the customization of test monkeys for an open source project and in an industry application, where the framework has been used for successfully transferring the idea of monkey testing into an industry solution. Thomas Wetzlmaier, Rudolf Ramler, Werner Putschögl |
ICST | 2 |
| 2016 | Exploring Expectations About Risk-Based Testing: Towards Increasing Effectiveness and Efficiency
Michael Felderer, Rudolf Ramler |
PROFES | 2 |
| 2016 | Risk orientation in software testing processes of small and medium enterprises: an exploratory and comparative study
Michael Felderer, Rudolf Ramler |
Softw. Qual. J. | 2 |
| 2015 | Model-Based Testing of Stateful APIs with ModbatabstractModbat makes testing easier by providing a user-friendly modeling language to describe the behavior of systems, from such a model, test cases are generated and executed. Modbat's domain-specific language is based on Scala, its features include probabilistic and non-deterministic transitions, component models with inheritance, and exceptions. We demonstrate the versatility of Modbat by finding a confirmed defect in the currently latest version of Java, and by testing SAT solvers. Cyrille Artho, Martina Seidl, Quentin Gros, Eun-Hye Choi, Takashi Kitamura 0001, Akira Mori, Rudolf Ramler, Yoriyuki Yamagata |
ASE | 7 |
| 2015 | GRT: Program-Analysis-Guided Random Testing (T)abstractWe propose Guided Random Testing (GRT), which uses static and dynamic analysis to include information on program types, data, and dependencies in various stages of automated test generation. Static analysis extracts knowledge from the system under test. Test coverage is further improved through state fuzzing and continuous coverage analysis. We evaluated GRT on 32 real-world projects and found that GRT outperforms major peer techniques in terms of code coverage (by 13 %) and mutation score (by 9 %). On the four studied benchmarks of Defects4J, which contain 224 real faults, GRT also shows better fault detection capability than peer techniques, finding 147 faults (66 %). Furthermore, in an in-depth evaluation on the latest versions of ten popular real-world projects, GRT successfully detects over 20 unknown defects that were confirmed by developers. Lei Ma 0003, Cyrille Artho, Hiroyuki Sato 0002, Johannes Gmeiner, Rudolf Ramler |
ASE | 6 |
| 2015 | GRT: An Automated Test Generator Using Orchestrated Program AnalysisabstractWhile being highly automated and easy to use, existing techniques of random testing suffer from low code coverage and defect detection ability for practical software applications. Most tools use a pure black-box approach, which does not use knowledge specific to the software under test. Mining and leveraging the information of the software under test can be promising to guide random testing to overcome such limitations. Guided Random Testing (GRT) implements this idea. GRT performs static analysis on software under test to extract relevant knowledge and further combines the information extracted at run-time to guide the whole test generation procedure. GRT is highly configurable, with each of its six program analysis components implemented as a pluggable module whose parameters can be adjusted. Besides generating test cases, GRT also automatically creates a test coverage report. We show our experience in GRT tool development and demonstrate its practical usage using two concrete application scenarios. Lei Ma 0003, Cyrille Artho, Hiroyuki Sato 0002, Johannes Gmeiner, Rudolf Ramler |
ASE | 6 |
| 2015 | A Process for Risk-Based Test Strategy Development and Its Industrial Evaluation
Rudolf Ramler, Michael Felderer |
PROFES | 1 |
| 2014 | Extracting Dependencies from Software Changes: An Industry Experience ReportabstractRetrieving and analyzing information from software repositories and detecting dependencies are important tasks supporting software evolution. Dependency information is used for change impact analysis, defect prediction as well as cohesion and coupling measurement. In this paper we report our experience from extracting dependency information from the change history of a commercial software system. We analyzed the software system's evolution of about six years, from the start of development to the transition to product releases and maintenance. Analyzing the co-evolution of software artifacts allows detecting logical dependencies between system parts implemented with heterogeneous technologies as well as between different types of development artifacts such as source code, data models or documentation. However, the quality of the extracted dependencies relies on established development practices and conformance to a defined change process. In this paper we indicate resulting limitations and recommend further processing and filtering steps to prepare the dependency data for subsequent analysis and measurement activities. Thomas Wetzlmaier, Claus Klammer, Rudolf Ramler |
IWSM/Mensura | 3 |
| 2014 | Integrating risk-based testing in industrial test processes
Michael Felderer, Rudolf Ramler |
Softw. Qual. J. | 2 |
| 2014 | A multiple case study on risk-based testing in industry
Michael Felderer, Rudolf Ramler |
Int. J. Softw. Tools Technol. Transf. | 2 |
| 2013 | A Retrospection on Building a Custom Tool for Automated System TestingabstractThe numerous commercial and open source test tools available today cover almost any of the features one may ever require for automating tests. However, companies still develop in-house solutions or extend existing tools with custom functionality. In our case, all started with the need to automate tests for a machinery system based on non-standard technologies. In this paper we review the experiences and results from building our own test tool. We discuss the unique advantages of this endeavor and contrast them to the actual effort and costs. We list the involved challenges, the solutions we found, and the issues that remained open. In the end, building our own tool was a success. But would we do it again? Rudolf Ramler, Werner Putschögl |
COMPSAC | 1 |
| 2013 | A Replicated Study on Random Test Case Generation and Manual Unit Testing: How Many Bugs Do Professional Developers Find?abstractThis paper describes the replication of an empirical study comparing tool-supported test case generation and manual development of unit tests. As variation to the original study, which was based on test results from students performing manual unit testing for 60 minutes, the replication involves professional software developers with several years of industry experience and extends the initial time restriction. As part of the replication the paper explores the differences in unit testing by students and professionals and investigates the impact of the extended time limit. The main findings are: There are no significant differences in the results produced by students and by professional developers when performing manual unit testing for 60 minutes. Furthermore, there is a non-linear increase in the number of defects found when the time limit is extended from one to two hours, which indicates the transition from the initial ramp-up phase to productive testing. The replication also confirms the conclusions of the original study: Automated test case generation can be equally efficient as manual unit testing under severe time restrictions and it may be used to complement a manual testing approach. Rudolf Ramler, Klaus Wolfmaier, Theodorich Kopetzky |
COMPSAC | 1 |
| 2013 | Points-to analysis of IEC 61131-3 programs: Implementation and applicationabstractA call graph of a program represents the information which executable program element calls which other executable program elements. Based on the call graph, points-to sets can be computed, which represent the memory locations a reference variable can possibly point to. Call graph and points-to sets provide important information for static program analysis. This is especially true for PLC programs which heavily use pointer variables. However, due to the complexity of the algorithms, call graph and points-to analysis methods are not widely available in static analysis. In this paper, we present an approach for call graph and points-to analysis of IEC 61131-3 programs. We present the algorithm for computing call graph and points-to sets and its implementation in a tool environment, show several different application scenarios, and present first results from industrial application. Florian Angerer, Herbert Prähofer, Rudolf Ramler, Friedrich Grillenberger |
ETFA | 3 |
| 2013 | Experiences from an Initial Study on Risk Probability Estimation Based on Expert OpinionabstractBackground: Determining the factor probability in risk estimation requires detailed knowledge about the software product and the development process. Basing estimates on expert opinion may be a viable approach if no other data is available. Objective: In this paper we analyze initial results from estimating the risk probability based on expert opinion to answer the questions (1) Are expert opinions consistent? (2) Do expert opinions reflect the actual situation? (3) How can the results be improved? Approach: An industry project serves as case for our study. In this project six members provided initial risk estimates for the components of a software system. The resulting estimates are compared to each other to reveal the agreement between experts and they are compared to the actual risk probabilities derived in an ex-post analysis from the released version. Results: We found a moderate agreement between the rations of the individual experts. We found a significant accuracy when compared to the risk probabilities computed from the actual defects. We identified a number of lessons learned useful for improving the simple initial estimation approach applied in the studied project. Conclusions: Risk estimates have successfully been derived from subjective expert opinions. However, additional measures should be applied to triangulate and improve expert estimates. Rudolf Ramler, Michael Felderer |
IWSM/Mensura | 1 |
| 2013 | Noise in Bug Report Data and the Impact on Defect Prediction ResultsabstractThe potential benefits of defect prediction have created widespread interest in research and generated a considerable number of empirical studies. Applications with real-world data revealed a central problem: Real-world data is "dirty" and often of poor quality. Noise in bug report data is a particular problem for defect prediction since it effects the correct classification of software modules. Is the module actually defective or not? In this paper we examine different causes of noise encountered when predicting defects in an industrial software system and we provide an overview of commonly reported causes in related work. Furthermore we conduct an experiment to explore the impact of class noise on the predictions performance. The experiment shows that the prediction results for the studied system remain reliable even at a noise level of 20% probability of incorrect links between bug reports and modules. Rudolf Ramler, Johannes Himmelbauer |
IWSM/Mensura | 1 |
| 2012 | Opportunities and challenges of static code analysis of IEC 61131-3 programsabstractStatic code analysis techniques analyze programs by examining the source code without actually executing them. The main benefits lie in improving software quality by detecting potential defects and problematic code constructs in early development stages. Today, static code analysis is widely used and numerous tools are available for established programming languages like C/C++, Java, C# and others. However, in the domain of PLC programming, static code analysis tools are still rare. In this paper we present an approach and tool support for static code analysis of PLC programs. The paper discusses opportunities static code analysis can offer for PLC programming, it reviews techniques for static analysis, and it describes our tool that implements a rule-based analysis approach for IEC 61131-3 programs. Herbert Prähofer, Florian Angerer, Rudolf Ramler, Hermann Lacheiner, Friedrich Grillenberger |
ETFA | 3 |
| 2012 | Rule-Based Detection of Process Conformance Violations in Application Lifecycle Management
Rudolf Ramler, Hermann Lacheiner, Albin Kern |
EuroSPI | 1 |
| 2012 | Combinatorial Test Design in the TOSCA Testsuite: Lessons Learned and Practical ImplicationsabstractThe advantage of combinatorial techniques over less structured approaches is supported by the experience from numerous real-world projects where a significant reduction of the number of test cases has been achieved without compromising functional coverage. However, to fully benefit from combinatorial testing, the applied techniques and tools have to satisfy the requirements and needs of testers and practitioners. In this paper we explore such requirements distilled from testing software systems for over 15 years across a wide range of projects in business and industry. Their practical implications span from mastering the combinatorial explosion over support for fault localization to understandability, changeability and maintainability. Finally, the paper illustrates how the different combinatorial techniques are able to meet these requirements. The combinatorial techniques discussed in this paper are part of the TOSCA Test suite developed by TRICENTIS. Rudolf Ramler, Theodorich Kopetzky, Wolfgang Platz |
ICST | 1 |
| 2012 | Improving Unfamiliar Code with Unit Tests: An Empirical Investigation on Tool-Supported and Human-Based Testing
Dietmar Winkler 0001, Martina Schmidt, Rudolf Ramler, Stefan Biffl |
PROFES | 3 |
| 2010 | The usual suspects: a case study on delivered defects per developerabstractIndividual differences of developers in performance and introduced defects have been reported by many research studies and are frequently observed in software development practice. Thus, when the source of defects in the final product is discussed, developers are usually the first under suspicion. However, defects residing in a released software product are the result of defects introduced throughout the sequence of development activities (e.g., specification, design, implementation, testing and stabilization) less the defects detected and removed in these activities. This case study explores and describes the difference between developers in terms of associated post-release (i.e., delivered) defects. The results are put in relation to the intensity with which a developer's changes and enhancements have been tested to identify a latent influence by pre-release quality assurance measures. Rudolf Ramler, Claus Klammer, Thomas Natschläger |
ESEM | 1 |
| 2009 | Key Questions in Building Defect Prediction Models in Practice
Rudolf Ramler, Klaus Wolfmaier, Erwin Stauder, Felix Kossak, Thomas Natschläger |
PROFES | 1 |
| 2008 | Issues and effort in integrating data from heterogeneous software repositories and corporate databasesabstractSoftware repositories and corporate databases capture different fragments of a project's history. Software cockpits integrate the data from these repositories and databases to provide a holistic view of the project and the capability to drill-down and analyze details. By incorporating existing data, the cockpit can be used effectively from the first day it is introduced. In this paper we describe our findings from integrating several repositories and databases for a large, distributed project. We highlight common issues in data integration, report on the resulting effort for the development of software cockpits, and share our lessons learned from this data integration project. Rudolf Ramler, Klaus Wolfmaier |
ESEM | 1 |
| 2008 | How to Test the Intangible Properties of Graphical User Interfaces?abstractIn this paper we describe our experience from developing and testing a visual graphical user interface (GUI) editor for mobile and multimedia devices. Testing of the editor's highly interactive user interface is critical for its success, yet remains a challenge due to the specification of often intangible quality characteristics of the GUI and its proneness to change. The approach we provide is supporting exploratory testing of the GUI with tools integrated with the tested object. Thus a step-by-step guide for manual exploratory testing can be enhanced with automated elements that directly manipulate the status of the editor, access internal properties of the GUI, and record interactions for bug reporting. Josef Pichler, Rudolf Ramler |
ICST | 2 |
| 2007 | Observing Distributions in Size Metrics: Experience from Analyzing Large Software SystemsabstractIn this paper we observe and compare distributions of popular size metric values from the analysis of different software systems as well as from different consecutive versions of one software system. The typically heavy-tailed distributions are visualized and discussed with the help of Pareto diagrams. We found that the distributions remain remarkable stable over time, support the identification of problem areas by statistical and relative threshold-based filtering, and show the ability to reveal the fundamental characteristics of a software system. Rudolf Ramler, Klaus Wolfmaier, Thomas Natschläger |
COMPSAC (2) | 1 |
| 2004 | From Maintenance to Evolutionary Development of Web Applications: A Pragmatic Approach
Rudolf Ramler, Klaus Wolfmaier, Edgar R. Weippl |
ICWE | 1 |
| 2004 | Decision Support for Test Management in Iterative and Evolutionary Development
Rudolf Ramler |
ASE | 1 |
| 2003 | Unit Testing beyond a Bar in Green and Red
Rudolf Ramler, Gerald Czech, Dietmar Schlosser |
XP | 1 |