EDBT 2026 Demo / reviewers in the wild / expert
Alexander Pretschner
dblp:98/5319 · also Walter Alexander Pretschner
· DBLP profile ↗
121ranked-venue papers
14as first author
40since 2021 · last 2026
0000-0002-5573-1201ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 60 · 9 first-author · 24 since 2021Security and privacy · 42 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information RetrievalabstractJiahui Geng, Fengyu Cai, Shaobo Cui, Qing Li, Liangwei Chen, Chenyang Lyu, Haonan Li, Derui Zhu, Alexander Pretschner, Heinz Koeppl, Fakhri Karray. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiahui Geng, Fengyu Cai, Shaobo Cui 0006, Qing Li 0038, Liangwei Chen, Chenyang Lyu, Derui Zhu, Alexander Pretschner, Heinz Koeppl, Fakhri Karray |
ACL (1) | 9 |
| 2025 | Analyzing the Usage of Donation Platforms for PyPI Libraries
Alexandros Tsakpinis, Alexander Pretschner |
EASE | 2 |
| 2025 | A Taxonomy of Integration-Relevant Faults for Microservice TestingabstractMicroservices have emerged as a popular architectural paradigm, offering a flexible and scalable approach to software development. However, their distributed nature and diverse technology stacks introduce inherent complexities, surpassing those of monolithic systems. The integration of microservices presents numerous challenges, from communication failures to compatibility issues, compromising system reliability. Understanding faults in these distributed components is crucial for preventing defects, devising test strategies, and implementing robustness testing. Despite the significance of these software systems, existing taxonomies are limited, as they primarily focus on non-functional attributes or lack empirical validation. To address these gaps, this paper proposes an extensive taxonomy of the most common integration-relevant faults observed in large-scale microservice systems in industry. Leveraging insights from a systematic literature review and ten semi-structured interviews with industry experts, we identify common integration-related faults encountered in real-world microservice projects. Our final taxonomy was validated through a survey with an additional set of 16 practitioners, confirming that almost all fault categories (21/23) were experienced by at least 50% of the survey participants. Lena Gregor, Anja Hentschel, Leon Kastner, Alexander Pretschner |
ICST | 4 |
| 2025 | RustyRTS: Regression Test Selection for RustabstractRegression testing is a testing activity that aims to ensure that existing functionality is preserved when introducing changes. The goal of regression test selection (RTS) is to reduce the cost of regression testing by only re-executing tests that are affected by changes. Lately, research on RTS has focused on the languages Java and C++. Despite Rust being an increasingly relevant systems programming language, there are no RTS tools available for this language so far. In this paper, we present and evaluate RUSTYRTS, the first RTS technique and tool for Rust. It provides both module-level and function-level RTS. Its function-level variants can rely on either static or dynamic code analysis to select affected tests. We evaluate RUSTYRTS in an empirical study in terms of safety, precision, and effectiveness. When applied to changes resulting from mutation testing on 9 open-source projects, RUSTYRTS selected 99.99%, 97.87% and 99.97% of all tests that failed in consequence of a modification in code using a module-level, static and dynamic RTS approach respectively. For cases of unsafe behavior, i.e., tests that have not been selected although failing due to such changes, we find plausible explanations. By applying RUSTYRTS to changes from the Git history of 13 repositories on GitHub, we effectively reduce the average end-to-end (e2e) testing time on the majority of projects. Our results show that the two function-level approaches outperform the coarser module-level one, especially on longer-running test suites. On average, the$e$2$e$testing time has been reduced to 67.80%, 62.52% and 52.79% of retest-all by module-level, static and dynamic RUSTYRTS respectively. Lastly, we provide a novel solution for dynamic dispatch and compile-time function evaluation, two contexts that impose a special challenge on approaches to RTS. Simon Hundsdorfer, Roland Würsching, Alexander Pretschner |
ICST | 3 |
| 2025 | Should We Evaluate LLM Based Security Analysis Approaches on Open Source Systems?abstractExisting research has demonstrated promising results when applying large language models (LLMs) to detect security vulnerabilities in source code. However, these studies have been exclusively evaluated on benchmarks from open-source systems, using publicly known vulnerabilities that are likely part of the LLMs’ training data. This raises concerns that reported performance metrics may be inflated due to data contamination, providing a misleading view of the models’ actual capabilities.In this paper, we quantify this effect with a case study that evaluates five frontier LLMs on two carefully curated datasets: CWE-Bench-Java (an open-source dataset) and TS-Vuls (based on a closed-source commercial codebase). To provide a second angle, we also split CWE-Bench-Java by CVE record date to explore temporal contamination based on LLM knowledge cutoff dates.Our results reveal that the average F1 score dropped by approximately 20 percentage points when comparing the open-source to the closed-source dataset. Additionally, the precision drops from 56% to 34% on average, which is statistically significant (p < 0.05) for four of five models. This declining trend is consistent across all tested LLMs and metrics. In contrast, the results for the temporal split on the open-source data are inconclusive, suggesting that using a knowledge cutoff may reduce but does not ensure the elimination of contamination effects.Although our study is based on a single closed-source system and thus not generalizable, these findings provide the first empirical evidence that evaluating LLM-based vulnerability detection on open-source benchmarks may lead to overly optimistic results. This motivates the inclusion of closed-source datasets in future LLM evaluations. Kohei Dozono, Jonas Engesser, Benjamin Hummel, Tobias Roehm, Alexander Pretschner |
ASE | 5 |
| 2025 | More Than Just Functional: LLM-as-a-Critique for Efficient Code GenerationabstractLarge language models (LLMs) have demonstrated remarkable progress in generating functional code, leading to numerous AI-based coding program tools. However, their reliance on the perplexity objective during both training and inference primarily emphasizes functionality, often at the expense of efficiency—an essential consideration for real-world coding tasks. Perhaps interestingly, we observed that well-trained LLMs inherently possess knowledge about code efficiency, but this potential remains underutilized with standard decoding approaches. To address this, we design strategic prompts to activate the model’s embedded efficiency understanding, effectively using LLMs as \textit{efficiency critiques} to guide code generation toward higher efficiency without sacrificing—and sometimes even improving—functionality, all without the need for costly real code execution. Extensive experiments on benchmark datasets (EffiBench, HumanEval+) across multiple representative code models demonstrate up to a 70.6\% reduction in average execution time and a 13.6\% decrease in maximum memory usage, highlighting the computational efficiency and practicality of our approach compared to existing alternatives. Derui Zhu, Dingfan Chen, Jinfu Chen 0002, Jens Grossklags, Alexander Pretschner, Weiyi Shang |
NeurIPS | 5 |
| 2024 | Analyzing the Accessibility of GitHub Repositories for PyPI and NPM LibrariesabstractIndustrial applications heavily rely on open-source software (OSS) libraries, which provide various benefits. But, they can also present a substantial risk if a vulnerability or attack arises and the community fails to promptly address the issue and release a fix due to inactivity. To be able to monitor the activities of such communities, a comprehensive list of repositories for the libraries of an ecosystem must be accessible. Based on these repositories, integrated libraries of an application can be monitored to observe whether they are adequately maintained. In this descriptive study, we analyze the accessibility of GitHub repositories for PyPI and NPM libraries. For all available libraries, we extract assigned repository URLs, direct dependencies and use the page rank algorithm to comprehensively analyze the ecosystems from a library and dependency chain perspective. For invalid repository URLs, we derive potential reasons. Both ecosystems show varying accessibility to GitHub repository URLs, depending on the page rank score of the analyzed libraries. For individual libraries, up to 73.8% of PyPI and up to 69.4% of NPM libraries have repository URLs. Within dependency chains, up to 80.1% of PyPI libraries have URLs, while up to 81.1% for NPM. That means, most libraries, especially the ones of increasing importance, can be monitored on GitHub. Among the most common reasons for invalid repository URLs is no URLs being assigned at all, which amounts up to 17.9% for PyPI and up to 39.6% for NPM. Package maintainers should address this issue and update the repository information to enable monitoring of their libraries. Alexandros Tsakpinis, Alexander Pretschner |
EASE | 2 |
| 2024 | Cost of Flaky Tests in Continuous Integration: An Industrial Case StudyabstractResearchers and practitioners alike increasingly often perceive flaky tests as a major challenge in software engineering. They spend a lot of effort trying to detect, repair, and mitigate the negative effects of flaky tests. However, it is yet unclear where and to what extent the costs of flaky tests manifest in industrial Continuous Integration (CI) development processes. In this study, we compile cost factors introduced by flaky tests in CI development from research and practice and derive a cost model that allows gaining insight into the costs incurred. We then instantiate this model in a case study of a large, commercial software project with ~30 developers and ~1M SLoC. We analyze five years of development history, including CI test logs, commits from the Version Control System (VCS), issue tickets, and tracked work time to quantify the cost factors implied by flaky tests. We find that the time spent dealing with flaky tests in the studied project represents at least 2.5% of the productive developer time. This effort is divided into investigating potentially flaky test failures, which accounts for 1.1% of the total time spent, repairing flaky tests adds another 1.3 %, and developing tools to monitor flaky tests adds 0.1 %. Contrary to most other studies, we find the cost for rerunning tests to be negligible and inexpensive. Automatically rerunning a test costs 0.02 cents, while not rerunning and thus letting the pipeline fail results in a manual investigation costing $5.67 in our context. The insights gained from our case study have led to the decision to shift effort from investigation and repair to automatically rerunning tests. Our cost model can help practitioners analyze the cost of flaky tests in their context and make informed decisions. Furthermore, our case study provides a first step to better understand the costs of flaky tests, which can lead researchers to industry-relevant problems. Fabian Leinen, Daniel Elsner, Alexander Pretschner, Andreas Stahlbauer, Michael Sailer, Elmar Jürgens |
ICST | 3 |
| 2024 | System-Level Test Case Generation and Execution for Distributed Cooperative Unmanned Aerial SystemsabstractSystems of cooperating unmanned aerial vehicles (UAVs) can achieve requirements that individual UAVs cannot achieve. Particularly of interest to us are groups composed of UAVs that are cooperating in a distributed way, which would include some types of drone swarms. We refer to these groups as distributed cooperative unmanned aerial systems (DC-UAS). Cyber-physical systems must be tested before operational deployment, and while there has been substantial research into designing DC-UAS capabilities, there is a lack of systematic approaches for generating and executing test cases for system-level verification. Scenario-based testing (SBT) combined with search methods (SBT+search) is an approach that has been used to test autonomous cars and UAVs, but not in the context of distributed cooperative operations. In this paper, we show how the existing SBT+search approach to test an individual UAV can be adapted to conduct system-level testing of a DC-UAS, despite the additional properties inherent to a DC-UAS. Specifically, the adaptation involves incorporating the global task of the DC-UAS and information about the system’s individual UAVs into the approach’s search process. We then utilize a model of a decentralized drone swarm to concretely demonstrate the adapted approach, resulting in the generation of scenarios that challenge the example system’s ability to behave safely and perform global tasks. We also propose additional considerations when formulating the test case generation process. The result is a testing approach for DC-UAS that builds on existing approaches, producing a solution for test case generation and execution through the perspective of system-level verification rather than design. David K. Marson, Diana Deldar, Alexander Pretschner |
IV | 3 |
| 2024 | A Failure Model Library for Simulation-Based Validation of Functional Safety
Tiziano Munaro, Irina Muntean, Alexander Pretschner |
SAFECOMP | 3 |
| 2024 | A References Architecture for Human Cyber Physical Systems, Part II: Fundamental Design Principles for Human-CPS InteractionabstractAs automation increases qualitatively and quantitatively in safety-critical human cyber-physical systems, it is becoming more and more challenging to increase the probability or ensure that human operators still perceive key artifacts and comprehend their roles in the system. In the companion paper, we proposed an abstract reference architecture capable of expressing all classes of system-level interactions in human cyber-physical systems. Here we demonstrate how this reference architecture supports the analysis of levels of communication between agents and helps to identify the potential for misunderstandings and misconceptions. We then develop a metamodel for safe human machine interaction. Therefore, we ask what type of information exchange must be supported on what level so that humans and systems can cooperate as a team, what is the criticality of exchanged information, what are timing requirements for such interactions, and how can we communicate highly critical information in a limited time frame in spite of the many sources of a distorted perception. We highlight shared stumbling blocks and illustrate shared design principles, which rest on established ontologies specific to particular application classes. In order to overcome the partial opacity of internal states of agents, we anticipate a key role of virtual twins of both human and technical cooperation partners for designing a suitable communication. Klaus Bengler, Werner Damm, Andreas Lüdtke, Jochem W. Rieger, Benedikt Austel, Bianca Biebl, Martin Fränzle, Willem Hagemann, Moritz Held, David Hess, Klas Ihme, Severin Kacianka, Alyssa J. Kerscher, Forrest Laine, Sebastian Lehnhoff, Alexander Pretschner, Astrid Rakow, Daniel Sonntag, Janos Sztipanovits, Maike Schwammberger, Mark Schweda, Anirudh Unni, Eric M. S. P. Veith |
ACM Trans. Cyber Phys. Syst. | 16 |
| 2024 | A Reference Architecture of Human Cyber-Physical Systems - Part III: Semantic FoundationsabstractThe design and analysis of multi-agent human cyber-physical systems in safety-critical or industry-critical domains calls for an adequate semantic foundation capable of exhaustively and rigorously describing all emergent effects in the joint dynamic behavior of the agents that are relevant to their safety and well-behavior. We present such a semantic foundation. This framework extends beyond previous approaches by extending the agent-local dynamic state beyond state components under direct control of the agent and belief about other agents (as previously suggested for understanding cooperative as well as rational behavior) to agent-local evidence and belief about the overall cooperative, competitive, or coopetitive game structure. We argue that this extension is necessary for rigorously analyzing systems of human cyber-physical systems because humans are known to employ cognitive replacement models of system dynamics that are both non-stationary and potentially incongruent. These replacement models induce visible and potentially harmful effects on their joint emergent behavior and the interaction with cyber-physical system components. Werner Damm, Martin Fränzle, Alyssa J. Kerscher, Forrest Laine, Klaus Bengler, Bianca Biebl, Willem Hagemann, Moritz Held, David Hess, Klas Ihme, Severin Kacianka, Sebastian Lehnhoff, Andreas Lüdtke, Alexander Pretschner, Astrid Rakow, Jochem W. Rieger, Daniel Sonntag, Janos Sztipanovits, Maike Schwammberger, Mark Schweda, Alexander Trende, Anirudh Unni, Eric M. S. P. Veith |
ACM Trans. Cyber Phys. Syst. | 14 |
| 2024 | A Reference Architecture of Human Cyber-Physical Systems - Part I: Fundamental ConceptsabstractWe propose a reference architecture of safety-critical or industry-critical human cyber-physical systems (CPSs) capable of expressing essential classes of system-level interactions between CPS and humans relevant for the societal acceptance of such systems. To reach this quality gate, the expressivity of the model must go beyond classical viewpoints such as operational, functional, and architectural views and views used for safety and security analysis. The model does so by incorporating elements of such systems for mutual introspections in situational awareness, capabilities, and intentions to enable a synergetic, trusted relation in the interaction of humans and CPSs, which we see as a prerequisite for their societal acceptance. The reference architecture is represented as a metamodel incorporating conceptual and behavioral semantic aspects. We illustrate the key concepts of the metamodel with examples from cooperative autonomous driving, the operating room of the future, cockpit-tower interaction, and crisis management. Werner Damm, David Hess, Mark Schweda, Janos Sztipanovits, Klaus Bengler, Bianca Biebl, Martin Fränzle, Willem Hagemann, Moritz Held, Klas Ihme, Severin Kacianka, Alyssa J. Kerscher, Sebastian Lehnhoff, Andreas Lüdtke, Alexander Pretschner, Astrid Rakow, Jochem W. Rieger, Daniel Sonntag, Maike Schwammberger, Benedikt Austel, Anirudh Unni, Eric M. S. P. Veith |
ACM Trans. Cyber Phys. Syst. | 15 |
| 2024 | Generation of Tailored and Confined Datasets for IDS Evaluation in Cyber-Physical SystemsabstractThe state-of-the-art evaluation of an Intrusion Detection System (IDS) relies on benchmark datasets composed of the regular system's and potential attackers' behavior. The datasets are collected once and independently of the IDS under analysis. This paper questions this practice by introducing a methodology to elicit particularly challenging samples to benchmark a given IDS. In detail, we propose (1) six fitness functions quantifying the suitability of individual samples, particularly tailored for safety-critical cyber-physical systems, (2) a scenario-based methodology for attacks on networks to systematically deduce optimal samples in addition to previous datasets, and (3) a respective extension of the standard IDS evaluation methodology. We applied our methodology to two network-based IDSs defending an advanced driver assistance system. Our results indicate that different IDSs show strongly differing characteristics in their edge case classifications and that the original datasets used for evaluation do not include such challenging behavior. In the worst case, this causes a critical undetected attack, as we document for one IDS. Our findings highlight the need to tailor benchmark datasets to the individual IDS in a final evaluation step. Especially the manual investigation of selected samples from edge case classifications by domain experts is vital for assessing the IDSs. Thomas Hutzelmann, Dominik Mauksch, Ana Petrovska, Alexander Pretschner |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2023 | Better Safe Than Sorry! Automated Identification of Functionality-Breaking Security-Configuration RulesabstractInsecure default values in software settings can be exploited by attackers to compromise the system that runs the software. As a countermeasure, there exist security-configuration guides specifying in detail which values are secure. However, most administrators still refrain from hardening existing systems because the system functionality is feared to deteriorate if secure settings are applied. To foster the application of security-configuration guides, it is necessary to identify those rules that would restrict the functionality.This article presents our approach to use combinatorial testing to find problematic combinations of rules and machine learning techniques to identify the problematic rules within these combinations. The administrators can then apply only the unproblematic rules and, therefore, increase the system’s security without the risk of disrupting its functionality. To demonstrate the usefulness of our approach, we applied it to real-world problems drawn from discussions with administrators at Siemens and found the problematic rules in these cases. We hope that this approach and its open-source implementation motivate more administrators to harden their systems and, thus, increase their systems’ general security. Patrick Stöckle, Michael Sammereier, Bernd Grobauer, Alexander Pretschner |
AST | 4 |
| 2023 | BinaryRTS: Cross-language Regression Test Selection for C++ Binaries in CIabstractContinuous integration (CI) pipelines are commonly used to execute regression tests before pull requests are merged. Regression test selection (RTS) aims to reduce the required testing effort and feedback time for developers. However, existing RTS techniques are imprecise for tests with cross-language links to compiled C++ binaries or unsafe if tests use external files. This is problematic because modern software in fact involves several programming languages and (non-)code artifacts such as configuration files. In this paper, we present BinaryRTS, a novel RTS technique that leverages dynamic binary instrumentation to collect the covered functions and accessed external files for each test. BinaryRTS then selects tests depending on changes issued to C++ binaries or external (non-)code artifacts. When evaluating BinaryRTS in our large-scale industrial context, we are able to exclude on average up to 74% of tests without missing real failures. We release BinaryRTS as the first publicly available RTS tool for software involving C++ code. Daniel Elsner, Severin Kacianka, Stephan Lipp, Alexander Pretschner, Axel Habermann, Maria Graber, Silke Reimer |
ICST | 4 |
| 2023 | DIRTS: Dependency Injection Aware Regression Test SelectionabstractRegression test selection (RTS) aims to reduce regression testing effort by selecting only those tests that are affected by introduced changes. RTS techniques are considered to be safe if they select all affected test cases. Several supposedly safe RTS tools have been developed over the past decades, lately especially for Java projects. However, recent studies have shown that state-of-the-art RTS tools for Java can become unsafe when confronted with dependency injection (DI) mechanisms: despite the widespread use of DI frameworks in Java projects, no existing technique acknowledges DI-related changes. In this paper, we analyze the reasons behind unsafe RTS behavior for DI-related changes and develop Dirts, a novel DI-aware RTS tool for Java. To counteract effects of DI on RTS, Dirts efficiently analyzes source code annotations and metadata employed by popular DI frameworks, and generates a dependency graph including edges for dynamically injected objects. We evaluate Dirts on 228 commits from 9 open-source Java projects that use DI. Our results indicate that in 33.3% of those commits DI-related changes affect some tests, and in 3.1% (7) Dirts identifies affected tests that are clearly missed by the static RTS tool STARTS. Still, Dirts is comparatively efficient and precise. We publish Dirts1,2as an RTS tool that can either be used as a safety extension for existing RTS tools or as a standalone RTS solution. Simon Hundsdorfer, Daniel Elsner, Alexander Pretschner |
ICST | 3 |
| 2023 | Severity-Aware Prioritization of System-Level Regression Tests in Automotive SoftwareabstractIn automotive software engineering, system-level regression testing is crucial to ensure proper integration of often- times safety-critical components. Due to the inherent complexity of such systems and components, testing is commonly performed manually and in a black-box manner, which is particularly costly and leads to slow feedback cycles between testers and developers. Regression Test Prioritization (RTP) aims to reduce feedback time by ordering tests to reveal faults earlier during the testing process. However, most prior RTP research does not incorporate varying fault severity, which must be taken into account when evaluating and designing appropriate RTP approaches for safety-critical automotive software systems. In this work, we present a case study at our industry partner MAN, a leading international provider of commercial vehicles. We design and instantiate a domain-specific, severity-aware RTP assessment model and comparatively assess state-of-the-art RTP approaches. Our results indicate that simple and partly well- known heuristics based on test history and test costs have the best cost-effectiveness, achieving between 85% and 90% of the maximum possible feedback time reduction. On the other hand, search-based and machine-learning-based RTP approaches do not perform better, especially if available test history is sparse. Roland Würsching, Daniel Elsner, Fabian Leinen, Alexander Pretschner, Georg Grueneissl, Thomas Neumeyr, Tobias Vosseler |
ICST | 4 |
| 2023 | Green Fuzzing: A Saturation-Based Stopping Criterion using Vulnerability PredictionabstractFuzzing is a widely used automated testing technique that uses random inputs to provoke program crashes indicating security breaches. A difficult but important question is when to stop a fuzzing campaign. Usually, a campaign is terminated when the number of crashes and/or covered code elements has not increased over a certain period of time. To avoid premature termination when a ramp-up time is needed before vulnerabilities are reached, code coverage is often preferred over crash count to decide when to terminate a campaign. However, a campaign might only increase the coverage on non-security-critical code or repeatedly trigger the same crashes. For these reasons, both code coverage and crash count tend to overestimate the fuzzing effectiveness, unnecessarily increasing the duration and thus the cost of the testing process. Stephan Lipp, Daniel Elsner, Severin Kacianka, Alexander Pretschner, Marcel Böhme, Sebastian Banescu |
ISSTA | 4 |
| 2023 | Revisiting Inter-Class Maintainability IndicatorsabstractOver the last few decades, a variety of static code metrics have been published and promoted to measure the maintainability of software systems.This study evaluates 12 common static code metrics for their correlation with observed maintenance efforts. Leveraging modern repository mining techniques, we examine the historical data of three large open-source software systems with a combined size of over 1M LOC and over 10k classes. We automatically identify maintenance activities and measure the effort needed to perform them through revised lines of code. Then, we investigate if the state of the system as captured by these metrics is an indicator for the required maintenance effort.In contrast to earlier research, our results could not validate a general correlation between any of the examined metrics and maintainability. Instead, all evaluated metrics showed positive and negative correlations with maintenance effort depending on the considered time interval. Strong correlations only hold for specific projects, and within these projects, only for limited time spans. Across the project history, however, all metrics showed moderate correlations at most.As no metric was found to be a good indicator for high maintenance efforts in all contexts, we advocate against using any of the evaluated metrics without project-specific validation. If metrics are to be used to monitor the maintainability of a system, either directly or through models based on these metrics, engineers have to validate their applicability not just for the project at hand, but also for the current time span. Lena Gregor, Markus Schnappinger, Alexander Pretschner |
SANER | 3 |
| 2023 | Decentralized Inverse Transparency with BlockchainabstractEmployee data can be used to facilitate work, but their misusage may pose risks for individuals.Inverse transparencytherefore aims to track all usages of personal data, allowing individuals to monitor them to ensure accountability for potential misusage. This necessitates a trusted log to establish an agreed-upon and non-repudiable timeline of events. The unique properties of blockchain facilitate this by providing immutability and availability. For power asymmetric environments such as the workplace, permissionless blockchain is especially beneficial as no trusted third party is required. Yet, two issues remain: (1) In a decentralized environment, no arbiter can facilitate and attest to data exchanges. Simple peer-to-peer sharing of data, conversely, lacks the required non-repudiation. (2) With data governed by privacy legislation such as the GDPR, the core advantage of immutability becomes a liability. After a rightful request, an individual’s personal data need to be rectified or deleted, which is impossible in an immutable blockchain. To solve these issues, we presentKovacs, a decentralized data exchange and usage logging system for inverse transparency built on blockchain. Its new-usage protocol ensures non-repudiation, and therefore accountability, for inverse transparency. Its one-time pseudonym generation algorithm guarantees unlinkability and enables proof of ownership, which allows data subjects to exercise their legal rights regarding their personal data. With our implementation, we show the viability of our solution. The decentralized communication impacts performance and scalability, but exchange duration and storage size are still reasonable. More importantly, the provided information security meets high requirements. We conclude thatKovacsrealizes decentralized inverse transparency through secure and GDPR-compliant use of permissionless blockchain. Valentin Zieglmeier, Gabriel Loyola Daiqui, Alexander Pretschner |
Distributed Ledger Technol. Res. Pract. | 3 |
| 2023 | Rethinking People Analytics With Inverse Transparency by DesignabstractEmployees work in increasingly digital environments that enable advanced analytics. Yet, they lack oversight over the systems that process their data. That means that potential analysis errors or hidden biases are hard to uncover. Recent data protection legislation tries to tackle these issues, but it is inadequate. It does not prevent data misusage while at the same time stifling sensible use cases for data. We think the conflict between data protection and increasingly data-driven systems should be solved differently. When access to an employees' data is given, all usages should be made transparent to them, according to the concept of inverse transparency. This allows individuals to benefit from sensible data usage while addressing the potentially harmful consequences of data misusage. To accomplish this, we propose a new design approach for workforce analytics software we refer to as inverse transparency by design. To understand the developer and user perspectives on the proposal, we conduct two exploratory studies with students. First, we let small teams of developers implement analytics tools with inverse transparency by design to uncover how they judge the approach and how it materializes in their developed tools. We find that architectural changes are made without inhibiting core functionality. The developers consider our approach valuable and technically feasible. Second, we conduct a user study over three months to let participants experience the provided inverse transparency and reflect on their experience. The study models a software development workplace where most work processes are already digital. Participants perceive the transparency as beneficial and feel empowered by it. They unanimously agree that it would be an improvement for the workplace. We conclude that inverse transparency by design is a promising approach to realize accepted and responsible people analytics. Valentin Zieglmeier, Alexander Pretschner |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2022 | Challenges in Regression Test Selection for End-to-End Testing of Microservice-based Software SystemsabstractDynamic regression test selection (RTS) techniques aim to minimize testing efforts by selecting tests using per-test execution traces. However, most existing RTS techniques are not applicable to microservice-based, or, more generally, distributed systems, as the dynamic program analysis is typically limited to a single system. In this paper, we describe our distributed RTS approach, microRTS, which targets automated and manual end-to-end testing in microservice-based software systems. We employ microRTS in a case study on a set of 20 manual end-to-end test cases across 12 versions of the German COVID-19 contact tracing application, a modern microservice-based software system. The results indicate that initially microRTS selects all manual test cases for each version. Yet, through semi-automated filtering of test traces, we are able to effectively reduce the testing effort by 10--50%. In contrast with prior results on automated unit tests, we find method-level granularity of per-test execution traces to be more suitable than class-level for manual end-to-end testing. Daniel Elsner, Daniel Bertagnolli, Alexander Pretschner, Rudi Klaus |
AST | 3 |
| 2022 | Probe-based Syscall Tracing for Efficient and Practical File-level Test TracesabstractEfficiently collecting per-test execution traces is a common prerequisite of dynamic regression test optimization techniques. However, as these test traces are typically recorded through language-specific code instrumentation, non-code artifacts and multi-language source code are usually not included. In contrast, more complete test traces can be obtained by instrumenting operating system calls and thereby tracing all accessed files during a test's execution. Yet, existing test optimization techniques that use syscall tracing are impractical as they either modify the Linux kernel or operate in user space, thus raising transferability, performance, and security concerns. Recent advances in operating system development provide versatile, lightweight, and safe kernel instrumentation frameworks: They allow to trace syscalls by instrumenting probes in the operating system kernel. Probe-based Syscall Tracing (ProST), our novel technique, harnesses this potential to collect file-level test traces that go beyond language boundaries and consider non-code artifacts. To evaluate ProST's efficiency and the completeness of obtained test traces, we perform an empirical study on 25 multi-language open-source software projects and compare our approach to existing language-specific instrumentation techniques. Our results show that most studied projects use source files from multiple languages (22/25) or non-code artifacts during testing (22/25) that are missed by language-specific techniques. With the low execution time overhead of 4.6% compared to non-instrumented test execution, ProST is more efficient than language-specific instrumentation. Furthermore, it collects on average 89% more files on top of those collected by language-specific techniques. Consequently, ProST paves the way for efficiently extracting valuable information through dynamic analysis to better understand and optimize testing in multi-language software systems. Daniel Elsner, Roland Würsching, Markus Schnappinger, Alexander Pretschner |
AST | 4 |
| 2022 | Hardening with Scapolite: A DevOps-based Approach for Improved Authoring and Testing of Security-Configuration Guides in Large-Scale OrganizationsabstractSecurity Hardening is the process of configuring IT systems to ensure the security of the systems' components and data they process or store. In many cases, so-called security-configuration guides are used as a basis for security hardening. These guides describe secure configuration settings for components such as operating systems and standard applications. Rigorous testing of security-configuration guides and automated mechanisms for their implementation and validation are necessary since erroneous implementations or checks of hardening guides may severely impact systems' security and functionality. At Siemens, centrally maintained security-configuration guides carry machine-readable information specifying both the implementation and validation of each required configuration step. The guides are maintained within git repositories; automated pipelines generate the artifacts for implementation and checking, e.g., PowerShell scripts for Windows, and carry out testing of these artifacts on AWS images. This paper describes our experiences with our DevOps-inspired approach for authoring, maintaining, and testing security-configuration guides. We want to share these experiences to help other organizations with their security hardening and increase their systems' security. Patrick Stöckle, Ionut Pruteanu, Bernd Grobauer, Alexander Pretschner |
CODASPY | 4 |
| 2022 | Why Did the Test Execution Fail? Failure Classification Using Association Rules (Practical Experience Report)abstractTesting automotive electronic control units and their software, the system-under-test (SUT) in our context, requires complex test infrastructure setups. As those setups are developed in parallel to the SUT, often a problem in the test infrastructure instead of the SUT causes a failed test case execution (TCE). We call such unexpectedly failing TCEs invalid. As there are several reasons that lead to such invalid test failure, failed TCEs are manually analyzed and categorized. This failure classification is a time-consuming task. Thus, automatic classification could significantly reduce overall development time and cost. A pre-vious study suggests using association rule learning (ARL) to classify failed TCEs as valid or invalid based solely on test step information. In this work, we extend this ARL-based approach to our multi-class setting and evaluate its application on data from five running verification & validation projects in the automotive industry. In total, we predict the defect classes of more than 75k TCEs and achieve an overall precision up to 86.7% with an overall recall up to 57.4%. With this work, we offer evidence that the application of said approach, originally presented in the context of information systems, can be fruitful in automotive integration- and system-level testing contexts as well. Claudius Jordan, Philipp Foth, Matthias Fruth, Alexander Pretschner |
ISSRE | 4 |
| 2022 | StellaUAV: A Tool for Testing the Safe Behavior of UAVs with Scenario-Based Testing (Tools and Artifact Track)abstractWhen we allow Unmanned Aerial Vehicles (UAVs) to perform their missions autonomously in the near future, we need to ensure their safe behavior. To generate relevant test cases that can reveal potential faults in the tested UAVs, we propose to leverage scenario-based testing from the automotive domain. For a systematic application of this methodology, we present StellaUAV, a tool for testing the safe behavior of UAVs with scenario-based testing. With our proposed tool, we can describe relevant test situations, generate test cases for these situations that can reveal potential faults in the tested UAV, and evaluate the performance of different optimization algorithms and their combinations. To demonstrate its applicability, we apply StellaUAV to generate test cases for various situations and discover several safety distance violations of the tested exemplary UAV in the presence of dynamic obstacles. These experimental results indicate that the given system under test can handle situations with only static obstacles rather well, while it encounters problems when facing dynamic ones. Further, we detect that a combination of optimization algorithms can find safety distance violations for a logical scenario that the widely used algorithm NSGAII deemed safe for the tested system. Overall, our results show that StellaUAV can effectively detect potential faults in the tested UAV. Tabea Schmidt, Alexander Pretschner |
ISSRE | 2 |
| 2022 | An empirical study on the effectiveness of static C code analyzers for vulnerability detectionabstractStatic code analysis is often used to scan source code for security vulnerabilities. Given the wide range of existing solutions implementing different analysis techniques, it is very challenging to perform an objective comparison between static analysis tools to determine which ones are most effective at detecting vulnerabilities. Existing studies are thereby limited in that (1) they use synthetic datasets, whose vulnerabilities do not reflect the complexity of security bugs that can be found in practice and/or (2) they do not provide differentiated analyses w.r.t. the types of vulnerabilities output by the static analyzers. Hence, their conclusions about an analyzer's capability to detect vulnerabilities may not generalize to real-world programs. In this paper, we propose a methodology for automatically evaluating the effectiveness of static code analyzers based on CVE reports. We evaluate five free and open-source and one commercial static C code analyzer(s) against 27 software projects containing a total of 1.15 million lines of code and 192 vulnerabilities (ground truth). While static C analyzers have been shown to perform well in benchmarks with synthetic bugs, our results indicate that state-of-the-art tools miss in-between 47% and 80% of the vulnerabilities in a benchmark set of real-world programs. Moreover, our study finds that this false negative rate can be reduced to 30% to 69% when combining the results of static analyzers, at the cost of 15 percentage points more functions flagged. Many vulnerabilities hence remain undetected, especially those beyond the classical memory-related security bugs. Stephan Lipp, Sebastian Banescu, Alexander Pretschner |
ISSTA | 3 |
| 2022 | Automated Identification of Security-Relevant Configuration Settings Using NLPabstractTo secure computer infrastructure, we need to configure all security-relevant settings. We need security experts to identify security-relevant settings, but this process is time-consuming and expensive. Our proposed solution uses state-of-the-art natural language processing to classify settings as security-relevant based on their description. Our evaluation shows that our trained classifiers do not perform well enough to replace the human security experts but can help them classify the settings. By publishing our labeled data sets and the code of our trained model, we want to help security experts analyze configuration settings and enable further research in this area. Patrick Stöckle, Theresa Wasserer, Bernd Grobauer, Alexander Pretschner |
ASE | 4 |
| 2022 | A Model-based System Engineering Plugin for Safety Architecture Pattern Synthesis
Yuri Gil Dantas, Tiziano Munaro, Carmen Cârlan, Vivek Nigam, Simon Barner, Shiqing Fan, Alexander Pretschner, Ulrich Schöpp, Sergey Tverdyshev |
MODELSWARD | 7 |
| 2022 | Data-Driven Assessment of Parameterized Scenarios for Autonomous Vehicles
Nicola Kolb, Florian Hauer 0002, Mojdeh Golagha, Alexander Pretschner |
SAFECOMP | 4 |
| 2022 | Exploring a Maximal Number of Relevant Obstacles for Testing UAVs
Tabea Schmidt, Florian Hauer 0002, Alexander Pretschner |
SAFECOMP | 3 |
| 2022 | PR-SZZ: How pull requests can support the tracing of defects in software repositoriesabstractThe SZZ algorithm represents a standard way to identify bug fixing commits as well as inducing counterparts. It forms the basis for data sets used in numerous empirical studies. Since its creation, multiple extensions have been proposed to enhance its performance. For historical reasons, related work relies on commit messages to map bug tickets to possibly related code with no additional data used to trace inducing commits from these fixes. Therefore, we present an updated version of SZZ utilizing pull requests, which are widely adopted today. We evaluate our approach in comparison to existing SZZ variants by conducting experiments and analyzing the usage of pull requests, inner commits, and merge strategies. We base our results on 6 open-source projects with more than 50k commits and 35k pull requests. With respect to bug fixing commits, on average 18% of bug tickets can be additionally mapped to a fixing commit, resulting in an overall F-score of 0.75, an improvement of 40 percentage points. By selecting an inducing commit, we manage to reduce the false-positives and increase precision by on average 16 percentage points in comparison to existing approaches. Peter Bludau, Alexander Pretschner |
SANER | 2 |
| 2022 | Defining adaptivity and logical architecture for engineering (smart) self-adaptive cyber-physical systems
Ana Petrovska, Stefan Kugele, Thomas Hutzelmann, Theo Beffart, Sebastian Bergemann, Alexander Pretschner |
Inf. Softw. Technol. | 6 |
| 2021 | Dynamic Taint Analysis versus Obfuscated Self-CheckingabstractSoftware protection in practice addresses the yearly loss of tens of billion USD for software manufacturers, a result of malicious end-users tampering with the software (”software cracking”). Software protection is prevalent in the gaming and license checking industries, and also relevant in the embedded and other industries. State of the art research in the area of software tamper protection against man-at-the-end (MATE) attackers focuses on the localization of integrity checks. The goal of this paper is a general assessment of the resilience of software self-checking, protected themselves by obfuscations against (1) (automated) detection and (2) (automated) bypass, without deobfuscating the code. Using dynamic taint analysis on a benchmark set of programs, we study how easy it is to detect and bypass combinations of self-checking and various obfuscation transformations. We aim at generalizing these findings across different programs rather than focusing on one particular program instance. To this end, we perform a set of controlled experiments using a data set of real-world programs, the MiBench suite and open-source games, and show that all of these can be broken by dynamic taint analysis attacks. To counter such attacks, we propose and implement improvements to an existing obfuscation implementation. We evaluate the implemented improvement and discuss the security-performance trade-offs. Sebastian Banescu, Samuel Valenzuela, Marius Guggenmos, Mohsen Ahmadvand, Alexander Pretschner |
ACSAC | 5 |
| 2021 | Human-level Ordinal Maintainability Prediction Based on Static Code MetricsabstractOne of the greatest challenges in software quality control is the efficient and effective measurement of maintainability. Thorough expert assessments are precise yet slow and expensive, whereas automated static analysis yields imprecise yet rapid feedback. Several machine learning approaches aim to integrate the advantages of both concepts. Markus Schnappinger, Arnaud Fietzke, Alexander Pretschner |
EASE | 3 |
| 2021 | Empirically evaluating readily available information for regression test optimization in continuous integrationabstractRegression test selection (RTS) and prioritization (RTP) techniques aim to reduce testing efforts and developer feedback time after a change to the code base. Using various information sources, including test traces, build dependencies, version control data, and test histories, they have been shown to be effective. However, not all of these sources are guaranteed to be available and accessible for arbitrary continuous integration (CI) environments. In contrast, metadata from version control systems (VCSs) and CI systems are readily available and inexpensive. Yet, corresponding RTP and RTS techniques are scattered across research and often only evaluated on synthetic faults or in a specific industrial context. It is cumbersome for practitioners to identify insights that apply to their context, let alone to calibrate associated parameters for maximum cost-effectiveness. This paper consolidates existing work on RTP and unsafe RTS into an actionable methodology to build and evaluate such approaches that exclusively rely on CI and VCS metadata. To investigate how these approaches from prior research compare in heterogeneous settings, we apply the methodology in a large-scale empirical study on a set of 23 projects covering 37,000 CI logs and 76,000 VCS commits. We find that these approaches significantly outperform established RTP baselines and, while still triggering 90% of the failures, we show that practitioners can expect to save on average 84% of test execution time for unsafe RTS. We also find that it can be beneficial to limit training data, features from test history work better than change-based features, and, somewhat surprisingly, simple and well-known heuristics often outperform complex machine-learned models. Daniel Elsner, Florian Hauer 0002, Alexander Pretschner, Silke Reimer |
ISSTA | 3 |
| 2021 | Understanding Safety for Unmanned Aerial Vehicles in Urban EnvironmentsabstractWhen Unmanned Aerial Vehicles (UAVs) autonomously operate in urban environments, it is especially important for these systems to behave safely and not harm anybody or anything. However, it is challenging to ensure that these systems behave safely in all possible situations and to clearly define this “safe” behavior for each situation. In this work, we provide a methodology for testing the safe behavior of UAV s while considering their environment with the help of scenario-based testing and search-based techniques. Additionally, we explore two cases throughout the paper: (i) A safety distance is specified, and we can use it for testing. (ii) No safety distance is defined, but we still aim to test the safe behavior of UAV s. In our experiments, we show the effectiveness and applicability of the proposed methods by discovering several safety distance violations and questionable behaviors of the tested UAV for both cases and four scenarios that represent all alternatives to avoid an obstacle. Tabea Schmidt, Florian Hauer 0002, Alexander Pretschner |
IV | 3 |
| 2021 | How can manual testing processes be optimized? developer survey, optimization guidelines, and case studiesabstractManual software testing is tedious and costly as it involves significant human effort. Yet, it is still widely applied in industry and will be in the foreseeable future. Although there is arguably a great need for optimization of manual testing processes, research focuses mostly on optimization techniques for automated tests. Accordingly, there is no precise understanding of the practices and processes of manual testing in industry nor about pitfalls and optimization potential that is untapped. To shed light on this issue, we conducted a survey among 38 testing professionals from 16 companies, to investigate their manual testing processes and to identify potential for optimization. We synthesize guidelines when optimization techniques from automated testing can be implemented for manual testing. By means of case studies on two industrial software projects, we show that fault detection likelihood, test feedback time and test creation efforts can be improved when following our guidelines. Roman Haas, Daniel Elsner, Elmar Jürgens, Alexander Pretschner, Sven Apel |
ESEC/SIGSOFT FSE | 4 |
| 2021 | Maat: Automatically Analyzing VirusTotal for Accurate Labeling and Effective Malware DetectionabstractThe malware analysis and detection research community relies on the online platform VirusTotal to label Android apps based on the scan results of around 60 antiviral scanners. Unfortunately, there are no standards on how to best interpret the scan results acquired from VirusTotal, which leads to the utilization of different threshold-based labeling strategies (e.g., if 10 or more scanners deem an app malicious, it is considered malicious). While some of the utilized thresholds may be able to accurately approximate the ground truths of apps, the fact that VirusTotal changes the set and versions of the scanners it uses makes such thresholds unsustainable over time. We implemented a method, Maat , that tackles these issues of standardization and sustainability by automatically generating a Machine Learning ( ML )-based labeling scheme, which outperforms threshold-based labeling strategies. Using the VirusTotal scan reports of 53K Android apps that span 1 year, we evaluated the applicability of Maat ’s Machine Learning ( ML )-based labeling strategies by comparing their performance against threshold-based strategies. We found that such ML -based strategies (a) can accurately and consistently label apps based on their VirusTotal scan reports, and (b) contribute to training ML -based detection methods that are more effective at classifying out-of-sample apps than their threshold-based counterparts. Aleieldin Salem, Sebastian Banescu, Alexander Pretschner |
ACM Trans. Priv. Secur. | 3 |
| 2020 | From Checking to Inference: Actual Causality Computations as Optimization Problems
Amjad Ibrahim, Alexander Pretschner |
ATVA | 2 |
| 2020 | Actual Causality Canvas: A General Framework for Explanation-Based Socio-Technical ConstructsabstractThe rapid deployment of digital systems into all aspects of daily life requires embedding social constructs into the digital world. Because of the complexity of these systems, there is a need for technical support to understand their actions. Social concepts, such as explainability, accountability, and responsibility rely on a notion of actual causality. Encapsulated in the Halpern and Pearl's (HP) definition, actual causality conveniently integrates into the socio-technical world if operationalized in concrete applications. To the best of our knowledge, theories of actual causality such as the HP definition are either applied in correspondence with domain-specific concepts (e.g., a lineage of a database query) or demonstrated using straightforward philosophical examples. On the other hand, there is a lack of explicit automated actual causality theories and operationalizations for helping understand the actions of systems. Therefore, this paper proposes a unifying framework and an interactive platform (Actual Causality Canvas) to address the problem of operationalizing actual causality for different domains and purposes. We apply this framework in such areas as aircraft accidents, unmanned aerial vehicles, and artificial intelligence (AI) systems for purposes of forensic investigation, fault diagnosis, and explainable AI. We show that with minimal effort, using our general-purpose interactive platform, actual causality reasoning can be integrated into these domains. Amjad Ibrahim, Tobias Klesel, Ehsan Zibaei, Severin Kacianka, Alexander Pretschner |
ECAI | 5 |
| 2020 | Defining a Software Maintainability Dataset: Collecting, Aggregating and Analysing Expert Evaluations of Software MaintainabilityabstractBefore controlling the quality of software systems, we need to assess it. In the case of maintainability, this often happens with manual expert reviews. Current automatic approaches have received criticism because their results often do not reflect the opinion of experts or are biased towards a small group of experts. We use the judgments of a significantly larger expert group to create a robust maintainability dataset. In a large scale survey, 70 professionals assessed code from 9 open and closed source Java projects with a combined size of 1.4 million source lines of code. The assessment covers an overall judgment as well as an assessment of several subdimensions of maintainability. Among these subdimensions, we present evidence that understandability is valued the most by the experts. Our analysis also reveals that disagreement between evaluators occurs frequently. Significant dissent was detected in 17% of the cases. To overcome these differences, we present a method to determine a consensus, i.e. the most probable true label. The resulting dataset contains the consensus of the experts for more than 500 Java classes. This corpus can be used to learn precise and practical classifiers for software maintainability. Markus Schnappinger, Arnaud Fietzke, Alexander Pretschner |
ICSME | 3 |
| 2020 | Can We Predict the Quality of Spectrum-based Fault Localization?abstractFault localization and repair are time-consuming and tedious. There is a significant and growing need for automated techniques to support such tasks. Despite significant progress in this area, existing fault localization techniques are not widely applied in practice yet and their effectiveness varies greatly from case to case. Existing work suggests new algorithms and ideas as well as adjustments to the test suites to improve the effectiveness of automated fault localization. However, important questions remain open: Why is the effectiveness of these techniques so unpredictable? What are the factors that influence the effectiveness of fault localization? Can we accurately predict fault localization effectiveness? In this paper, we try to answer these questions by collecting 70 static, dynamic, test suite, and fault-related metrics that we hypothesize are related to effectiveness. Our analysis shows that a combination of only a few static, dynamic, and test metrics enables the construction of a prediction model with excellent discrimination power between levels of effectiveness (eight metrics yielding an AUC of .86; fifteen metrics yielding an AUC of.88). The model hence yields a practically useful confidence factor that can be used to assess the potential effectiveness of fault localization. Given that the metrics are the most influential metrics explaining the effectiveness of fault localization, they can also be used as a guide for corrective actions on code and test suites leading to more effective fault localization. Mojdeh Golagha, Alexander Pretschner, Lionel C. Briand |
ICST | 2 |
| 2020 | Clustering Traffic Scenarios Using Mental Models as Little as PossibleabstractTest scenario generation for testing automated and autonomous driving systems requires knowledge about the recurring traffic cases, known as scenario types. The most common approach in industry is to have experts create lists of scenario types. This poses the risk both that certain types are overlooked; and that the mental model that underlies the manual process is inadequate. We propose to extract scenario types from real driving data by clustering recorded scenario instances, which are composed of timeseries. Existing works in the domain of traffic data either cannot cope with multivariate timeseries; are limited to one or two vehicles per scenario instance; or they use handcrafted features that are based on the mental model of the data scientist. The latter suffers from similar shortcomings as manual scenario type derivation. Our approach clusters scenario instances relying as little as possible on a mental model. As such, we consider the approach an important complement to manual scenario type derivation. It may yield scenario types overlooked by the experts, and it may provide a different segmentation of a whole set of scenarios instances into scenario types, thus overall increasing confidence in the handcrafted scenario types. We present the application of the approach to a real driving dataset. Florian Hauer 0002, Ilias Gerostathopoulos, Tabea Schmidt, Alexander Pretschner |
IV | 4 |
| 2020 | Re-Using Concrete Test Scenarios Generally Is a Bad IdeaabstractMany approaches for testing automated and autonomous driving systems in dynamic traffic scenarios rely on the reuse of test cases, e.g., recording test scenarios during real test drives or creating “test catalogs.” Both are widely used in industry and in literature. By counterexample, we show that the quality of test cases is system-dependent and that faulty system behavior may stay unrevealed during testing if test cases are naïvely re-used. We argue that, in general, system-specific “good” test cases need to be generated. Thus, recorded scenarios in general cannot simply be used for testing, and regression testing strategies needs to be rethought for automated and autonomous driving systems. The counterexample involves a system built according to state-of-the-art literature, which is tested in a traffic scenario using a high-fidelity physical simulation tool. Test scenarios are generated using standard techniques from the literature and state-of-the-art methodologies. By comparing the quality of test cases, we argue against a naïve re-use of test cases. Florian Hauer 0002, Alexander Pretschner, Bernd Holzmüller |
IV | 2 |
| 2020 | Automated Implementation of Windows-related Security-Configuration GuidesabstractHardening is the process of configuring IT systems to ensure the security of the systems' components and data they process or store. The complexity of contemporary IT infrastructures, however, renders manual security hardening and maintenance a daunting task. Patrick Stöckle, Bernd Grobauer, Alexander Pretschner |
ASE | 3 |
| 2020 | Automated Anomaly Detection in CPS Log Files - A Time Series Clustering Approach
Tabea Schmidt, Florian Hauer 0002, Alexander Pretschner |
SAFECOMP | 3 |
| 2019 | Failure clustering without coverageabstractDeveloping and integrating software in the automotive industry is a complex task and requires extensive testing. An important cost factor in testing and debugging is the time required to analyze failing tests. In the context of regression testing, usually, large numbers of tests fail due to a few underlying faults. Clustering failing tests with respect to their underlying faults can, therefore, help in reducing the required analysis time. In this paper, we propose a clustering technique to group failing hardware-in-the-loop tests based on non-code-based features, retrieved from three different sources. To effectively reduce the analysis effort, the clustering tool selects a representative test for each cluster. Instead of analyzing all failing tests, testers only inspect the representative tests to find the underlying faults. We evaluated the effectiveness and efficiency of our solution in a major automotive company using 86 regression test runs, 8743 failing tests, and 1531 faults. The results show that utilizing our clustering tool, testers can reduce the analysis time more than 60% and find more than 80% of the faults only by inspecting the representative tests. Mojdeh Golagha, Constantin Lehnhoff, Alexander Pretschner, Hermann Ilmberger |
ISSTA | 3 |
| 2019 | Learning a classifier for prediction of maintainability based on static analysis toolsabstractStatic Code Analysis Tools are a popular aid to monitor and control the quality of software systems. Still, these tools only provide a large number of measurements that have to be interpreted by the developers in order to obtain insights about the actual quality of the software. In cooperation with professional quality analysts, we manually inspected source code from three different projects and evaluated its maintainability. We then trained machine learning algorithms to predict the human maintainability evaluation of program classes based on code metrics. The code metrics include structural metrics such as nesting depth, cloning information and abstractions like the number of code smells. We evaluated this approach on a dataset of more than 115,000 Lines of Code. Our model is able to predict up to 81% of the threefold labels correctly and achieves a precision of 80%. Thus, we believe this is a promising contribution towards automated maintainability prediction. In addition, we analyzed the attributes in our created dataset and identified the features with the highest predictive power, i.e. code clones, method length, and the number of alerts raised by the tool Teamscale. This insight provides valuable help for users needing to prioritize tool measurements. Markus Schnappinger, Mohd Hafeez Osman, Alexander Pretschner, Arnaud Fietzke |
ICPC | 3 |
| 2019 | Fitness Functions for Testing Automated and Autonomous Driving Systems
Florian Hauer 0002, Alexander Pretschner, Bernd Holzmüller |
SAFECOMP | 2 |
| 2019 | Guest editorial for the special section on MODELS 2016
Jörg Kienzle, Alexander Pretschner |
Softw. Syst. Model. | 2 |
| 2019 | Leveraging Compression-Based Graph Mining for Behavior-Based Malware DetectionabstractBehavior-based detection approaches commonly address the threat of statically obfuscated malware. Such approaches often use graphs to represent process or system behavior and typically employ frequency-based graph mining techniques to extract characteristic patterns from collections of malware graphs. Recent studies in the molecule mining domain suggest that frequency-based graph mining algorithms often perform sub-optimally in finding highly discriminating patterns. We propose a novel malware detection approach that uses so-called compression-based mining on quantitative data flow graphs to derive highly accurate detection models. Our evaluation on a large and diverse malware set shows that our approach outperforms frequency-based detection models in terms of detection effectiveness by more than 600 percent. Tobias Wüchner, Aleksander Cislak, Martín Ochoa, Alexander Pretschner |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2018 | Practical Integrity Protection with Oblivious HashingabstractOblivious hashing (OH) is an integrity protection technique that checks the (side) effects resulting from the executed code, in contrast to checking the code itself as done by self-checking (SC). SC introduces atypical behavior in the program logic, like reading the code section loaded in memory. Since such atypical behavior can be detected by attackers, OH is more appealing to be employed in practice than SC. However, OH is incapable of protecting a presumable majority of program instructions, those that depend on nondeterministic (input) data or branches, which have to be manually identified and subsequently skipped. In this paper, we extend OH into a practical protection scheme by proposing i) a technique for automatic segregation of deterministic instructions, and ii) a novel extension, Short Range Oblivious Hashing (SROH), for OH to cover control-flow instructions dependent on nondeterministic data. Our SROH technique increases the range of instructions that OH can protect to nondeterministic branches. Moreover, we intertwine OH with SC to cover (nondeterministic) data dependent instructions and enhance the resilience against tampering attacks. We evaluate the performance overhead as well as the security of our scheme using the MiBench dataset and 3 open source games. Our experiments show that the proposed technique yields a 20-fold increase in the median number of protected instructions and, on non-CPU-intensive programs, imposes an overhead of 52%. Mohsen Ahmadvand, Anahit Hayrapetyan, Sebastian Banescu, Alexander Pretschner |
ACSAC | 4 |
| 2018 | Software quality assessment in practice: a hypothesis-driven frameworkabstractSoftware quality models describe decompositions of quality characteristics. However, in practice, there is a gap between quality models, quality measurements, and quality assessment activities. As a first step of bridging the gap, this paper presents a novel and structured framework to perform quality assessments. Together with our industrial partner, we applied this framework in two case studies and present our lessons learned. Among others, we found that results from automated tools can be misleading. Manual inspections still need to be conducted to find hidden quality issues, and concrete evidence of quality violations needs to be collected to convince the stakeholders. Markus Schnappinger, Mohd Hafeez Osman, Alexander Pretschner, Markus Pizka, Arnaud Fietzke |
ESEM | 3 |
| 2018 | Understanding and Formalizing Accountability for Cyber-Physical SystemsabstractAccountability is the property of a system that enables the uncovering of causes for events and helps understand who or what is responsible for these events. Definitions and interpretations of accountability differ; however, they are typically expressed in natural language that obscures design decisions and the impact on the overall system. This paper presents a formal model to express the accountability properties of cyber-physical systems. To illustrate the usefulness of our approach, we demonstrate how three different interpretations of accountability can be expressed using the proposed model and describe the implementation implications through a case study. This formal model can be used to highlight context specific-elements of accountability mechanisms, define their capabilities, and express different notions of accountability. In addition, it makes design decisions explicit and facilitates discussion, analysis and comparison of different approaches. Severin Kacianka, Alexander Pretschner |
SMC | 2 |
| 2018 | Data Usage Control for Distributed SystemsabstractData usage control enables data owners to enforce policies over how their data may be used after they have been released and accessed. We address distributed aspects of this problem, which arise if the protected data reside within multiple systems. We contribute by formalizing, implementing, and evaluating a fully decentralized system that (i) generically and transparently tracks protected data across systems, (ii) propagates data usage policies along, and (iii) efficiently and preventively enforces policies in a decentralized manner. The evaluation shows that (i) dataflow tracking and policy propagation achieve a throughput of 21--54% of native execution and (ii) decentralized policy enforcement outperforms a centralized approach in many situations. Florian Kelbert, Alexander Pretschner |
ACM Trans. Priv. Secur. | 2 |
| 2017 | Detecting Patching of Executables without System CallsabstractPopular software applications (e.g. web browsers) are targeted by malicious organizations which develop potentially unwanted programs (PUPs). If such a PUP executes on benign user devices, it is able to manipulate the process memory of popular applications, their locally stored resources or their environment in a profitable way for the attacker and in detriment to benign end-users. We describe the implementation of a tamper detection mechanism based on code self-checksumming, able to detect static and dynamic patching of executables, performed by PUPs or other attackers. As opposed to other works based on code self-checksumming, our approach can also checksum instructions which contain absolute addresses affected by relocation, without using calls to external libraries. We implemented this solution for the x86 ISA and evaluated the performance impact and effectiveness. The results indicate that the run-time overhead of self-checksumming grows proportionally with the level of protection, which can be specified as input to our implementation. We have applied our implementation on the Chromium web-browser and observed that the overhead is practically unobservable for the end-user. Sebastian Banescu, Mohsen Ahmadvand, Alexander Pretschner, Robert Shield, Chris Hamilton |
CODASPY | 3 |
| 2017 | Predicting the Resilience of Obfuscated Code Against Symbolic Execution Attacks via Machine Learning
Sebastian Banescu, Christian S. Collberg, Alexander Pretschner |
USENIX Security Symposium | 3 |
| 2016 | Code obfuscation against symbolic execution attacks
Sebastian Banescu, Christian S. Collberg, Vijay Ganesh 0001, Zack Newsham, Alexander Pretschner |
ACSAC | 5 |
| 2016 | Profiting from Unit Tests for Integration TestingabstractIn practice, integration testing typically focuses on a small selection of components or subsystems to integrate and test. This reduces the effort required to create test cases and test environments. However, many defects are only detected when performing integration testing on all possible integrations. These defects are typically only detected later in the development process and lead to increased testing and fault localization efforts. By describing and operationalizing knowledge of such defects, we are able to (semi-)automatically detect them in integration testing. Our OUTFIT tool targets superfluous or missing functionality and untested exception/fault handling in Matlab Simulink models and generated code. It re-uses existing, or automatically generates, high coverage test cases to measure coverage in an opportunistically assembled integration of components or subsystems. A manual inspection of the coverage results then reveals missing or potentially superfluous behavior and thus reveals defects of the targeted kind. Used in a bottom-up integration testing strategy, OUTFIT front loads the detection of such defects, and reduces the fault localization effort. We evaluate OUTFIT using three components of a real-world electrical engine control system of a hybrid car. We find that the results are reproducible, effective and efficiently produced. The achieved coverage of the results is reproducible for 10 executions within a small standard deviation. OUTFIT is effective in finding a potential defect, and efficiently analyzes all evaluated components within a worst case execution time of 110 minutes. Dominik Holling, Andreas Hofbauer, Alexander Pretschner, Matthias Gemmar |
ICST | 3 |
| 2016 | Failure Models for Testing Continuous ControllersabstractRanging from temperature control to safety-critical applications, continuous controllers are used in a plethora of applications becoming increasingly complex. In turn, testing continuous control systems also is more complex. Particularly, application-specific manual formal analysis or testing the complete input range becomes infeasible. We present a comprehensive failure-based testing methodology and a respective automated tool for continuous controllers. Our methodology is based on an existing automated approach, testing stability, liveness, smoothness and responsiveness in a single value-response scenario only. We performed a practitioner survey and literature review in the domain revealing the quality criteria steadiness and reliability to be vital for meaningful testing of continuous controllers. In addition, we identified 4 further scenarios including disturbance response for comprehensive testing. We contribute a library of failure models and quality criteria for the automated testing of continuous control systems more complete than in previous approaches. On the grounds of our comprehensive experiments on 9 real-world control systems, our results demonstrate our failure-based testing methodology to provide better worst cases than manual testing (effectiveness) within an adequate time frame (efficiency) for any configuration used in our experiments (reproducibility). Dominik Holling, Alvin Stanescu, Kristian Beckers, Alexander Pretschner, Matthias Gemmar |
ISSRE | 4 |
| 2016 | MACKE: compositional analysis of low-level vulnerabilities with symbolic executionabstractConcolic (concrete+symbolic) execution has recently gained popularity as an effective means to uncover non-trivial vulnerabilities in software, such as subtle buffer overflows. However, symbolic execution tools that are designed to optimize statement coverage often fail to cover potentially vulnerable code because of complex system interactions and scalability issues of constraint solvers. In this paper, we present a tool (MACKE) that is based on the modular interactions inferred by static code analysis, which is combined with symbolic execution and directed inter-procedural path exploration. This provides an advantage in terms of statement coverage and ability to uncover more vulnerabilities. Our tool includes a novel feature in the form of interactive vulnerability report generation that helps developers prioritize bug fixing based on severity scores. A demo of our tool is available at https://youtu.be/icC3jc3mHEU. Saahil Ognawala, Martín Ochoa, Alexander Pretschner, Tobias Limmer |
ASE | 3 |
| 2016 | Generating behavior-based malware detection models with genetic programmingabstractMalware remains a major IT security threat and current detection approaches struggle to cope with a professionalized malware development industry. We propose the use of genetic programming to generate effective and robust malware detection models which we call FrankenMods. These are sets of graph metrics that capture characteristic malware behavior. Evolution of FrankenMods with good detection capabilities yields continuously improved detection effectiveness. FrankenMods are operationalized by evaluating them on quantitative data flow graphs that model malware behavior as data flows between system resources caused by issued system calls. We show that FrankenMods are substantially more robust and effective than a state-of-the-art graph metric-based detection approach. Tobias Wüchner, Martín Ochoa, Enrico Lovat, Alexander Pretschner |
PST | 4 |
| 2016 | Enhancing Operation Security using Secret SharingabstractStoring highly confidential data and carrying out security-related operations are crucial to many systems. Starting
from an industrial use case we propose a generic architecture based on secret sharing which address critical
operation authorization. By comparing and benchmarking different scheme from the literature we analyze the
different trade-offs (security, functionality, performance) which can be achieved. Finally by providing an open
source .NET implementation of several secret sharing schemes, this paper aims to rise awareness regarding
the capabilities of such algorithms to increase security in industrial setting. Mohsen Ahmadvand, Antoine Scemama, Martín Ochoa, Alexander Pretschner |
SECRYPT | 4 |
| 2016 | Model-based security testing: a taxonomy and systematic classificationabstractModel-based security testing relies on models to test whether a software system meets its security requirements. It is an active research field of high relevance for industrial applications, with many approaches and notable results published in recent years. This article provides a taxonomy for model-based security testing approaches. It comprises filter criteria (i.e. model of system security, security model of the environment and explicit test selection criteria) as well as evidence criteria (i.e. maturity of evaluated system, evidence measures and evidence level). The taxonomy is based on a comprehensive analysis of existing classification schemes for model-based testing and security testing. To demonstrate its adequacy, 119 publications on model-based security testing are systematically extracted from the five most relevant digital libraries by three researchers and classified according to the defined filter and evidence criteria. On the basis of the classified publications, the article provides an overview of the state of the art in model-based security testing and discusses promising research directions with regard to security properties, coverage criteria and the feasibility and return on investment of model-based security testing. Copyright © 2015 John Wiley & Sons, Ltd. Michael Felderer, Philipp Zech, Ruth Breu, Matthias Büchler, Alexander Pretschner |
Softw. Test. Verification Reliab. | 5 |
| 2015 | A Fully Decentralized Data Usage Control Enforcement Infrastructure
Florian Kelbert, Alexander Pretschner |
ACNS | 2 |
| 2015 | Software-Based Protection against ChangewareabstractWe call changeware software that surreptitiously modifies resources of software applications, e.g., configuration files. Changeware is developed by malicious entities which gain profit if their changeware is executed by large numbers of end-users of the targeted software. Browser hijacking malware is one popular example that aims at changing web-browser settings such as the default search engine or the home page. Changeware tends to provoke end-user dissatisfaction with the target application, e.g. due to repeated failure of persisting the desired configuration. We describe a solution to counter changeware, to be employed by vendors of software targeted by changeware. It combines several protection mechanisms: white-box cryptography to hide a cryptographic key, software diversity to counter automated key retrieval attacks, and run-time process memory integrity checking to avoid illegitimate calls of the developed API. Sebastian Banescu, Alexander Pretschner, Dominic Battré, Stéfano Cazzulani, Robert Shield, Greg Thompson |
CODASPY | 2 |
| 2015 | Automated Translation of End User Policies for Usage Control Enforcement
Prachi Kumari, Alexander Pretschner |
DBSec | 2 |
| 2015 | Robust and Effective Malware Detection Through Quantitative Data Flow Graph Metrics
Tobias Wüchner, Martín Ochoa, Alexander Pretschner |
DIMVA | 3 |
| 2015 | Field Study on the Elicitation and Classification of Defects for Defect Models
Dominik Holling, Daniel Méndez 0001, Alexander Pretschner |
PROFES | 3 |
| 2015 | ISA2R: Improving Software Attack and Analysis Resilience via Compiler-Level Software Diversity
Rafael Fedler, Sebastian Banescu, Alexander Pretschner |
SAFECOMP | 3 |
| 2015 | SHRIFT System-Wide HybRid Information Flow Tracking
Enrico Lovat, Alexander Fromm, Martin Mohr, Alexander Pretschner |
SEC | 4 |
| 2014 | Decentralized Distributed Data Usage Control
Florian Kelbert, Alexander Pretschner |
CANS | 2 |
| 2014 | Malware detection with quantitative data flow graphsabstractWe propose a novel behavioral malware detection approach based on a generic system-wide quantitative data flow model. We base our data flow analysis on the incremental construction of aggregated quantitative data flow graphs. These graphs represent communication between different system entities such as processes, sockets, files or system registries. We demonstrate the feasibility of our approach through a prototypical instantiation and implementation for the Windows operating system. Our experiments yield encouraging results: in our data set of samples from common malware families and popular non-malicious applications, our approach has a detection rate of 96% and a false positive rate of less than 1.6%. In comparison with closely related data flow based approaches, we achieve similar detection effectiveness with considerably better performance: an average full system analysis takes less than one second. Tobias Wüchner, Martín Ochoa, Alexander Pretschner |
AsiaCCS | 3 |
| 2014 | On quantitative dynamic data flow trackingabstractWe present a non-probabilistic model for dynamic quantitative data flow tracking. Estimations of the amount of data stored in a particular representation at runtime - a file, a window, a network packet - enable the adoption of fine-grained policies which authorize or prohibit partial leaks of data. We prove the correctness of the estimations, provide an implementation that we evaluate w.r.t. precision and performance, and analyze one instantiation at the OS level. Enrico Lovat, Johan Oudinet, Alexander Pretschner |
CODASPY | 3 |
| 2014 | 8Cage: lightweight fault-based test generation for simulinkabstractMatlab Simulink models, mainly used for the specification of continuous embedded systems, employ a data flow-driven notation well understood by engineers. This notation abstracts from the underlying computational model, hiding run time failures such as over-/underflows and divisions by zero. They are often detected late in the development process by the use of static analysis tools on the completely developed system. The responsible underlying faults are sometimes attributable to a single operation in a model. 8Cage is an automated test case generator for the early detection of such single operation related faults. It is configurable to detect these faults and runs automatically in the background. It tries to find potentially failure-causing operations and generates a test case to gather evidence for an actual fault. 8Cage is usable by developing/testing engineers with knowledge of Matlab. It does not require an expert to perform result validation or fault localization. Dominik Holling, Alexander Pretschner, Matthias Gemmar |
ASE | 2 |
| 2014 | DAVAST: data-centric system level activity visualizationabstractHost-based intrusion detection systems need to be complemented by analysis tools that help understand if malware or attackers have indeed intruded, what they have done, and what the consequences are. We present a tool that visualizes system activities as data flow graphs: nodes are operating system entities such as processes, files, and sockets; edges are data flows between the nodes. Pattern matching identifies structures that correspond to (suspected) malicious and (suspected) normal behaviors. Matches are highlighted in slices of the data flow graph. As a proof of concept, we show how email worm attacks, drive-by downloads, and data leakage are detected, visualized, and analyzed. Tobias Wüchner, Alexander Pretschner, Martín Ochoa |
VizSEC | 2 |
| 2013 | Enforcing privacy through usage-controlled video surveillanceabstractIncreasing capabilities of intelligent video surveillance systems require the enforcement of privacy-related requirements. Data usage control technologies offer appropriate solutions in this problem domain. We first present specific requirements for a privacy enforcement infrastructure for modern surveillance systems that we align with a generic architecture and a privacy-aware workflow template for operating such systems. To ensure the compliance of a surveillance system's operation with such a workflow, we then derive respective usage control requirements. We show that the conceptual framework of usage control provides suitable instruments for specifying these requirements and for implementing the corresponding enforcement mechanisms. Our architecture has been implemented prototypically. Pascal Birnstill, Alexander Pretschner |
AVSS | 2 |
| 2013 | Data usage control enforcement in distributed systemsabstractDistributed usage control is concerned with how data may or may not be used in distributed system environments after initial access has been granted. If data flows through a distributed system, there exist multiple copies of the data on different client machines. Usage constraints then have to be enforced for all these clients. We extend a generic model for intra-system data flow tracking---that has been designed and used to track the existence of copies of data on single clients---to the cross-system case. When transferring, i.e., copying, data from one machine to another, our model makes it possible to (1) transfer usage control policies along with the data to the end of local enforcement at the receiving end, and (2) to be aware of the existence of copies of the data in the distributed system. As one example, we concretize "transfer of data" to the Transmission Control Protocol (TCP). Based on this concretized model, we develop a distributed usage control enforcement infrastructure that generically and application-independently extends the scope of usage control enforcement to any system receiving usage-controlled data. We instantiate and implement our work for OpenBSD and evaluate its security and performance. Florian Kelbert, Alexander Pretschner |
CODASPY | 2 |
| 2013 | A Generic Fault Model for Quality Assurance
Alexander Pretschner, Dominik Holling, Robert Eschbach, Matthias Gemmar |
MoDELS | 1 |
| 2012 | Deriving implementation-level policies for usage control enforcementabstractUsage control is concerned with how data is used after access to it has been granted. As such, it is particularly relevant to end users who own the data. System implementations of access and usage control enforcement mechanisms, however, do not always adequately reflect end user requirements. This is due to several reasons, one of which is the problem of mapping concepts in the end user's domain to technical events and artifacts. For instance, semantics of basic operators such as "copy" or "delete", which are fundamental for specifying privacy policies, tend to vary according to context. For this reason they can be mapped to different sets of system events. The behaviour users expect from the system, therefore, may differ from the actual behaviour. In this paper we present a translation of specification-level usage control policies into implementation-level policies which takes into account the precise semantics of domain-specific abstractions. A tool for automating the translation has also been implemented. Prachi Kumari, Alexander Pretschner |
CODASPY | 2 |
| 2012 | SPaCiTE - Web Application Testing EngineabstractWeb applications and web services enjoy an ever-increasing popularity. Such applications have to face a variety of sophisticated and subtle attacks. The difficulty of identifying respective vulnerabilities steadily increases with the complexity of applications. Moreover, the art of penetration testing predominantly depends on the skills of highly trained test experts. The difficulty to test web applications hence represents a daunting challenge to their developers. As a step towards improving security analyses, model checking has, at the model level, been found capable of identifying complex attacks and thus moving security analyses towards a push-button technology. In order to bridge the gap with actual systems, we present Spa Cite. This tool relies on a dedicated model-checker for security analyses that generates potential attacks with regard to common vulnerabilities in web applications. Then, it semi-automatically runs those attacks on the System Under Validation (SUV) and reports which vulnerabilities were successfully exploited. We applied Spa Cite to Role-Based-Access-Control (RBAC) and Cross-Site Scripting (XSS) lessons of Web Goat, an insecure web application maintained by OWASP. The tool successfully reproduced RBAC and XSS attacks. Matthias Büchler, Johan Oudinet, Alexander Pretschner |
ICST | 3 |
| 2012 | Data Loss Prevention Based on Data-Driven Usage ControlabstractInadvertent data disclosure by insiders is considered as one of the biggest threats for corporate information security. Data loss prevention systems typically try to cope with this problem by monitoring access to confidential data and preventing their leakage or improper handling. Current solutions in this area, however, often provide limited means to enforce more complex security policies that for instance specify temporal or cardinal constraints on the execution of events. This paper presents UC4Win, a data loss prevention solution for Microsoft Windows operating systems that is based on the concept of data-driven usage control to allow such a fine-grained policy-based protection. UC4Win is capable of detecting and controlling data-loss related events at the level of individual function calls. This is done with function call interposition techniques to intercept application calls to the Windows API in combination with methods to track the flows of confidential data through the system. Tobias Wüchner, Alexander Pretschner |
ISSRE | 2 |
| 2012 | Towards a policy enforcement infrastructure for distributed usage controlabstractDistributed usage control is concerned with how data may or may not be used after initial access to it has been granted and is therefore particularly important in distributed system environments. We present an application- and application-protocol-independent infrastructure that allows for the enforcement of usage control policies in a distributed environment. We instantiate the infrastructure for transferring files using FTP and for a scenario where smart meters are connected to a Facebook application. Florian Kelbert, Alexander Pretschner |
SACMAT | 2 |
| 2012 | PrefaceabstractThis series of Workshops is organized by ERCIM (European Research Consortium in Informatics and Mathematics) Working Group on Security and Trust Management, that was established in 2005 to foster collaborative work on all theoretical and practical aspects of security, trust and privacy in ICT within the European research community and to increase co-operation of the research community with the European industry.The four articles in this Special Issue were selected from the 17 papers that were accepted for presentation at the Workshop, out of 40 submissions.The articles cover representative topics of STM, such as access control, privacy, security and trust policies, or formal methods.The paper "Scalable automated symbolic analysis of administrative role-based access control policies by SMT solving" by Armando and Ranise considers the problem of automatically analysing administrative role based access control (ARBAC) policies.Such analyses are extremely useful to understand the subtle implications of complex combinations of authorizations, and more generally to ensure that ARBAC systems are scalable and can evolve gracefully.Authors isolate a new class of properties (user-role reachability) for ARBAC policies and leverage state-of-the-art SMT solvers to check automatically their validity.Finally, the strength of their approach is demonstrated through extensive experimental evaluation.In their paper "Stateful authorization logic -proof theory and a case study", Garg and Pfenning embark on the proof-theoretical study of logics for stateful authorization policies; that is, policies that depend on externally controlled conditions.They introduce an expressive logic, called BL, for reasoning about stateful authorization policies.The expressiveness of the logic is achieved through the adoption of a modal connective that allows delegation, and through an explicit modeling of time.Thanks to a careful design that crisply separates between the validity of state predicates and BL judgments, the authors provide an elegant treatment of the meta-theory of BL, and prove that it enjoys admissibility of cut, an essential property that entails consistency.The BL logic is empirically validated through a detailed case study of the US authorization policies for accessing to sensitive intelligence information.The paper "Modeling and preventing inferences from sensitive value distributions in data release" by Bezzi, De Capitani di Vimercati, Foresti, Livraga, Samarati and Sassi considers the following setting: assume a database contains many entries, each of which in its own is considered non-sensitive and can be declassified, but some Gilles Barthe, Jorge Cuéllar, Javier López 0001, Alexander Pretschner |
J. Comput. Secur. | 4 |
| 2012 | A taxonomy of model-based testing approachesabstractSUMMARY Model‐based testing (MBT) relies on models of a system under test and/or its environment to derive test cases for the system. This paper discusses the process of MBT and defines a taxonomy that covers the key aspects of MBT approaches. It is intended to help with understanding the characteristics, similarities and differences of those approaches, and with classifying the approach used in a particular MBT tool. To illustrate the taxonomy, a description of how three different examples of MBT tools fit into the taxonomy is provided. Copyright © 2011 John Wiley & Sons, Ltd. Mark Utting, Alexander Pretschner, Bruno Legeard |
Softw. Test. Verification Reliab. | 2 |
| 2011 | A Hypervisor-Based Bus System for Usage ControlabstractData usage control is concerned with requirements on data after access has been granted. In order to enforce usage control requirements, it is necessary to track the different representations that the data may take (among others, file, window content, network packet). These representations exist at different layers of abstraction. As a consequence, in order to enforce usage control requirements, multiple data flow tracking and usage control enforcement monitors must exist, one at each layer. If a new representation is created at some layer of abstraction, e.g., if a cache file is created for a picture after downloading it with a browser, then the initiating layer (in the example, the browser) must notify the layer at which the new representation is created (in the example, the operating system). We present a bus system for system-wide usage control that, for security and performance reasons, is implemented in a hyper visor. We evaluate its security and performance. Cornelius Moucha, Enrico Lovat, Alexander Pretschner |
ARES | 3 |
| 2011 | A Trustworthy Usage Control Enforcement FrameworkabstractUsage control policies specify restrictions on the handling of data after access has been granted. We present the design and implementation of a framework for enforcing usage control requirements and demonstrate its genericity by instantiating it to two different levels of abstraction, those of the operating system and an enterprise service bus. This framework consists of a policy language, an automatic conversion of policies into enforcement mechanisms, and technology implemented on the grounds of trusted computing technology that makes it possible to detect tampering with the infrastructure. We show how this framework can, among other things, be used to enforce separation-of-duty policies. We provide a performance analysis. Ricardo Neisse, Alexander Pretschner, Valentina Di Giacomo |
ARES | 2 |
| 2011 | Implementing Trust in Cloud InfrastructuresabstractToday's cloud computing infrastructures usually require customers who transfer data into the cloud to trust the providers of the cloud infrastructure. Not every customer is willing to grant this trust without justification. It should be possible to detect that at least the configuration of the cloud infrastructure -- as provided in the form of a hyper visor and administrative domain software -- has not been changed without the customer's consent. We present a system that enables periodical and necessity-driven integrity measurements and remote attestations of vital parts of cloud computing infrastructures. Building on the analysis of several relevant attack scenarios, our system is implemented on top of the Xen Cloud Platform and makes use of trusted computing technology to provide security guarantees. We evaluate both security and performance of this system. We show how our system attests the integrity of a cloud infrastructure and detects all changes performed by system administrators in a typical software configuration, even in the presence of a simulated denial-of-service attack. Ricardo Neisse, Dominik Holling, Alexander Pretschner |
CCGRID | 3 |
| 2011 | Distributed data usage control for web applications: a social network implementationabstractUsage control is concerned with how data is used after access to it has been granted. Respective enforcement mechanisms need to be implemented at different layers of abstraction in order to monitor or control data at and across all these layers. We present a usage control enforcement mechanism at the application layer. It is implemented for a common web browser and, as an example, is used to control data in a social network application. With the help of the mechanism, a data owner can, on the grounds of assigned trust values, prevent data from being printed, saved, copied&pasted, etc., after this data has been downloaded by other users. Prachi Kumari, Alexander Pretschner, Jonas Peschla, Jens-Michael Kuhn |
CODASPY | 2 |
| 2011 | Data-centric multi-layer usage control enforcement: a social network exampleabstractUsage control is concerned with how data is used after access to it has been granted. Data may exist in multiple representations which potentially reside at different layers of abstraction, including operating system, window manager, application level, DBMS, etc. Consequently, enforcement mechanisms need to be implemented at different layers, in order to monitor and control data at and across all of them. Enrico Lovat, Alexander Pretschner |
SACMAT | 2 |
| 2011 | On the number and nature of faults found by random testingabstractAbstract Intuition suggests that random testing should exhibit a considerable difference in the number of faults detected by two different runs of equal duration. As a consequence, random testing would be rather unpredictable. This article first evaluates the variance over time of the number of faults detected by randomly testing object‐oriented software that is equipped with contracts. It presents the results of an empirical study based on 1215 h of randomly testing 27 Eiffel classes, each with 30 seeds of the random number generator. The analysis of over 6 million failures triggered during the experiments shows that therelative numberof faults detected by random testing over time is predictable, but that different runs of the random test case generator detectdifferent faults. The experiment also suggests that the random testing quickly finds faults: the first failure is likely to be triggered within 30 s. The second part of this article evaluates thenatureof the faults found by random testing. To this end, it first explains a fault classification scheme, which is also used to compare the faults found through random testing with those found through manual testing and with those found in field use of the software and recorded in user incident reports. The results of the comparisons show that each technique is good at uncovering different kinds of faults. None of the techniques subsumes any of the others; each brings distinct contributions. This supports a more general conclusion on comparisons between testing strategies: thenumberof detected faults is too coarse a criterion for such comparisons—thenatureof faults must also be considered. Copyright © 2009 John Wiley & Sons, Ltd. Ilinca Ciupa, Alexander Pretschner, Manuel Oriol, Andreas Leitner, Bertrand Meyer 0001 |
Softw. Test. Verification Reliab. | 2 |
| 2009 | Formal Analyses of Usage Control PoliciesabstractUsage control is a generalization of access control that also addresses how data is handled after it is released. Usage control requirements are specified in policies. We present tool support for the following analysis problems. Is a policy consistent, i.e., satisfiable? Is an abstractly specified usage controlmechanism capable of enforcing a given policy? Can we configure such a mechanism by analyzing respective policies? In the context of propagation, where upon re-distribution of data duties may only be increased and rights decreased, can we check if a policy is only strengthened in this sense? — Our solution uses a modelchecker as theorem prover and is based on a translation ofusage control policies into a Linear Time Logic (LTL) dialect. We provide evidence that even complex policies can be analyzed efficiently. Alexander Pretschner, Judith Rüesch, Christian Schaefer, Thomas Walter 0001 |
ARES | 1 |
| 2009 | On the Effectiveness of Test Extraction without OverheadabstractDevelopers write and execute ad-hoc tests as they implement software. While these tests reflect important insights of the developers (e.g., which parts of the software need testing and what inputs should be used), they are usually not persistent and are easily forgotten. They cannot always be re-executed automatically, for example to debug or to test for regressions. Several methods that make such test cases persistent and automatically executable have been proposed. They rely on capturing state and/or events at runtime and thus induce significant overhead or require specialized hardware. In previous work we proposed a method that, in the event of a failure, extracts test cases solely from the state at the time of the failure (and not from before the failure). We call this method "failure-state extraction". Capturing the state only at the moment of failure reduces the run-time overhead to zero, but comes at a cost: state extracted in this way cannot always be used to reproduce the failure. This paper provides an experimental evaluation of failure-state extraction. The results show that the method is highly effective: in the experiment, 90% of all failures were reproducible using failure-state extraction and thus could be extracted without run-time overhead. Andreas Leitner, Alexander Pretschner, Stefan Mori, Bertrand Meyer 0001, Manuel Oriol |
ICST | 2 |
| 2009 | State-Based Usage Control Enforcement with Data Flow Tracking using System Call InterpositionabstractUsage control generalizes access control to what happens to data in the future. We contribute to the enforcement of usage control requirements at the level of system calls by also taking into account data flow: Restrictions on the dissemination of data, for instance, as stipulated by data protection regulations, of course relate not to just one file containing the data, but likely to all copies of that file as well. In order to enforce the dissemination restrictions on all copies of the sensitive data item, we introduce a data flow model that tracks how the content of a file flows through the system (files, network sockets, main memory). By using this model, the existence of potential copies of the data is reflected in the state of the data flow model. This allows us to enforce the dissemination restrictions by relating to the state rather than all sequences of events that possibly yield copies. Generalizing this idea, we describe how usage control policies can be expressed in a related state-based manner. Finally, we present an implementation of the data flow model and state-based policy enforcement as well as first encouraging performance measurements. Matús Harvan, Alexander Pretschner |
NSS | 2 |
| 2008 | Negotiation of Usage Control Policies - Simply the Best?abstractThe term "negotiation" suggests that multi-step bidirectional communication takes place. In this position paper, we play the devil's advocate and argue that (automated) policy negotiation essentially is one of the following, at least in the area of usage control. It can come down to a three-phase protocol that consists of a client request, a set of offers by the server, and the client's choice of an offer or to abort. Policy negotiation can also consist of a client request together with acceptable conditions plus the server's choice of one condition or to abort. In other words, negotiation of policies is a mere choice among alternatives; there is no negotiation in the intuitive sense of the word. - The goal of this position paper is to stimulate the discussion on what (automated) "policy negotiation" really is or can be. Alexander Pretschner, Thomas Walter 0001 |
ARES | 1 |
| 2008 | Mechanisms for usage controlabstractUsage control is a generalization of access control that also addresses how data is used after it is released. We present a formal model for different mechanisms that can enforce usage control policies on the consumer side. Alexander Pretschner, Manuel Hilty, David A. Basin, Christian Schaefer, Thomas Walter 0001 |
AsiaCCS | 1 |
| 2008 | On the Predictability of Random Tests for Object-Oriented SoftwareabstractIntuition suggests that random testing of object-oriented programs should exhibit a significant difference in the number of faults detected by two different runs of equal duration. As a consequence, random testing would be rather unpredictable. We evaluate the variance of the number of faults detected by random testing over time. We present the results of an empirical study that is based on 1215 hours of randomly testing 27 Eiffel classes, each with 30 seeds of the random number generator. Analyzing over 6 million failures triggered during the experiments, the study provides evidence that the relative number of faults detected by random testing over time is predictable but that different runs of the random test case generator detect different faults. The study also shows that random testing quickly finds faults: the first failure is likely to be triggered within 30 seconds. Ilinca Ciupa, Alexander Pretschner, Andreas Leitner, Manuel Oriol, Bertrand Meyer 0001 |
ICST | 2 |
| 2008 | Model-Based Tests for Access Control PoliciesabstractWe present a model-based approach to testing access control requirements. By using combinatorial testing, we first automatically generate test cases from and without access control policies-i.e., the model- and assess the effectiveness of the test suites by means of mutation testing. We also compare them to purely random tests. For some of the investigated strategies, non-random tests kill considerably more mutants than the same number of random tests. Since we rely on policies only, no information on the application is required at this stage. As a consequence, our methodology applies to arbitrary implementations of the policy decision points. Alexander Pretschner, Tejeddine Mouelhi, Yves Le Traon |
ICST | 1 |
| 2008 | Test-Driven Assessment of Access Control in Legacy ApplicationsabstractIf access control policy decision points are not neatly separated from the business logic of a system, the evolution of a security policy likely leads to the necessity of changing the system's code base. This is often the case with legacy systems. We present a test- driven methodology to assess the flexibility of a system, a property that describes the degree of coupling between the access control logic and the business logic of a system. A low flexibility indicates that a modification of the policy will lead to substantial changes of the code. In this paper, we analyze the notion of flexibility which is related to the presence of hidden and implicit security mechanisms in the business logic. We detail how testing can be used for detecting such mechanisms and how it may drive the incremental evolution of a security policy. We use several case studies to illustrate and validate the methodology. Yves Le Traon, Tejeddine Mouelhi, Alexander Pretschner, Benoit Baudry |
ICST | 3 |
| 2008 | Finding Faults: Manual Testing vs. Random+ Testing vs. User ReportsabstractThe usual way to compare testing strategies, whether theoretically or empirically, is to compare the number of faults they detect. To ascertain definitely that a testing strategy is better than another, this is a rather coarse criterion: shouldn't the nature of faults matter as well as their number? The empirical study reported here confirms this conjecture. An analysis of faults detected in Eiffel libraries through three different techniques-random tests, manual tests, and user incident reports-shows that each is good at uncovering significantly different kinds of faults. None of the techniques subsumes any of the others, but each brings distinct contributions. Ilinca Ciupa, Bertrand Meyer 0001, Manuel Oriol, Alexander Pretschner |
ISSRE | 4 |
| 2008 | Doctoral Symposium at MODELS 2008
Alexander Pretschner |
MoDELS | 1 |
| 2007 | A Policy Language for Distributed Usage Control
Manuel Hilty, Alexander Pretschner, David A. Basin, Christian Schaefer, Thomas Walter 0001 |
ESORICS | 2 |
| 2007 | Usage Control in Service-Oriented Architectures
Alexander Pretschner, Fabio Massacci, Manuel Hilty |
TrustBus | 1 |
| 2007 | Engineering Automotive SoftwareabstractThe amount of software in cars grows exponentially. Driving forces of this development are the availability of cheaper and more powerful hardware, as well as the demand for innovation through new functionality. The rapidly growing significance of software and software-based functionality is at the root of various challenges in the automotive industries. These concern their organization, definition of key competencies, processes, methods, tools, models, product structures, division of labor, logistics, maintenance, and long-term strategies. This paper pinpoints the idiosyncrasies of the domain, characterizes the essentials of automotive software, and discusses the challenges of automotive software engineering Manfred Broy, Ingolf Krüger, Alexander Pretschner, Chris Salzmann |
Proc. IEEE | 3 |
| 2007 | Computing refactorings of state machines
Alexander Pretschner, Wolfgang Prenninger |
Softw. Syst. Model. | 1 |
| 2006 | 3rd international workshop on software engineering for automotive systems - SEAS 2006abstractThis workshop summary presents an overview of the one-day International Workshop on Software Engineering for Automotive Systems (SEAS 2006), held in conjunction with the 28th International Conference on Software Engineering (ICSE'06). Details about SEAS 2006 may be found at: http://www.inf.ethz.ch/personal/pretscha/events/seas06/. Martin Rappl, Alexander Pretschner, Chris Salzmann, Thomas Stauner |
ICSE | 2 |
| 2005 | On Obligations
Manuel Hilty, David A. Basin, Alexander Pretschner |
ESORICS | 3 |
| 2005 | Model-Based Testing in Practice
Alexander Pretschner |
FM | 1 |
| 2005 | Model-based testingabstractModel-based testing has become increasingly popular in recent years. Major reasons include: (1) the need for quality assurance for increasingly complex systems, (2) the emerging model-centric development paradigm, e.g., UML and MDA, with its seemingly direct connection to testing, and (3) the advent of test-centered development methodologies. Model-based testing relies on execution traces of behavior models. They are used as test cases for an implementation: input and expected output. This complements the ideas of model-driven testing. The latter uses static models to derive test drivers to automate test execution. This assumes the existence of test cases, and is, like the particular intricacies of OO testing, not in the focus of this tutorial. We cover major methodological and technological issues: the business case of model-based testing within model-based development, the need for abstraction and inverse concretization, test selection, and test case generation. We (1) discuss different scenarios of model-based testing, (2) present common abstractions when building models, and their consequences for testing, (3) explain how to use functional, structural, and stochastic test selection criteria, and (4) describe today's test generation technology. We provide both practical guidance and a discussion of the state-of-the-art. Potentials of model-based testing in practical applications and future research are highlighted. Alexander Pretschner |
ICSE | 1 |
| 2005 | One evaluation of model-based testing and its automationabstractModel-based testing relies on behavior models for the generation of model traces: input and expected output---test cases---for an implementation. We use the case study of an automotive network controller to assess different test suites in terms of error detection, model coverage, and implementation coverage. Some of these suites were generated automatically with and without models, purely at random, and with dedicated functional test selection criteria. Other suites were derived manually, with and without the model at hand. Both automatically and manually derived model-based test suites detected significantly more requirements errors than hand-crafted test suites that were directly derived from the requirements. The number of detected programming errors did not depend on the use of models. Automatically generated model-based test suites detected as many errors as hand-crafted model-based suites with the same number of tests. A sixfold increase in the number of model-based tests led to an 11% increase in detected errors. Alexander Pretschner, Wolfgang Prenninger, Stefan Wagner 0001, Christian Kühnel, Martin Baumgartner, Bernd Sostawa, Rüdiger Zölch, Thomas Stauner |
ICSE | 1 |
| 2005 | 2nd international workshop on software engineering for automotive systemsabstractNo abstract available Chris Salzmann, Thomas Stauner, Alexander Pretschner |
ICSE | 3 |
| 2004 | ICSE Workshop: Software Engineering for Automotive Systems
Chris Salzmann, Thomas Stauner, Alexander Pretschner |
ICSE | 3 |
| 2004 | Coverage Metrics for Continuous Function ChartsabstractContinuous Function Charts are a diagrammatical language for the specification of mixed discrete-continuous embedded systems, similar to the languages of Matlab/Simulink, and often used in the domain of transportation systems. Both control and data flows are explicitly specified when atomic units of computation are composed. The obvious way to assess the quality of integration test suites is to compute known coverage metrics for the generated code. This production code does not exhibit those structures that would make it amenable to "relevant" coverage measurements. We define a translation scheme that results in structures relevant for such measurements, apply coverage criteria for both control and dataflows at the level of composition of atomic computational units, and argue for their usefulness on the grounds of detected errors. Vadim Alyokhin, Benedikte Elbel, Martin Rothfelder, Alexander Pretschner |
ISSRE | 4 |
| 2004 | Model based testing in incremental system development
Alexander Pretschner, Heiko Lötzbeyer, Jan Philipps |
J. Syst. Softw. | 1 |
| 2004 | Model-based testing for real
Alexander Pretschner, Oscar Slotosch, Ernst Aiglstorfer, Stefan Kriebel |
Int. J. Softw. Tools Technol. Transf. | 1 |
| 2003 | Ontology-based personalized search and browsing
Susan Gauch, Jason Chaffee, Alexander Pretschner |
Web Intell. Agent Syst. | 3 |
| 2000 | Specification based test sequence generation with propositional logicabstractIn the domain of concurrent reactive systems, much work has been devoted to (semi-)automatically validating a system's correctness. In this paper a novel approach to the automated generation of test sequences is presented. It may be used for both glass box testing a specification and black box testing an implementation (software/hardware). Finite system models specified within the CASE tool AutoFocus as well as user-friendly test case specifications are automatically translated into propositional logic and fed into the propositional solver SATO. Results are interpreted as input/output traces (test sequences) of the system, and may be displayed as message sequence charts. A small example illustrates the basic ideas as well as the method's advantages and shortcomings. The testing process is integrated into an overall development process. Main contributions include the implementation of a tool for graphical specification of test cases and the description of an efficient method to compute test sequences fully automatically as well as its integration into the same CASE tool. Copyright © 2000 John Wiley & Sons, Ltd. Guido Wimmel, Heiko Lötzbeyer, Alexander Pretschner, Oscar Slotosch |
Softw. Test. Verification Reliab. | 3 |
| 1999 | Ontology-Based Web Site Mapping for Information ExplorationabstractCentralized search process requires that the whole collection reside at a single site. This imposes a burden on both the system storage of the site and the network traffic near the site. It thus comes to require the search process to be distributed. Recently, more and more Web sites provide the ability to search their local collection of Web pages. Query brokering systems are used to direct queries to the promising sites and merge the results from these sites. Creation of meta-information of the sites plays an important role in such systems. In this article, we introduce an ontology-based web site mapping method used to produce conceptual meta-information, the Vector Space approach, and present a serial of experiments comparing it with Naïve-Bayes approach. We found that the Vector Space approach produces better accuracy in ontology-based web site mapping. Xiaolan Zhu, Susan Gauch, Lutz Gerhard, Nicholas Kral, Alexander Pretschner |
CIKM | 5 |
| 1999 | Ontology Based Personalized SearchabstractWith the exponentially growing amount of information available on the Internet, the task of retrieving documents of interest has become increasingly difficult. Search engines usually return more than 1,500 results per query, yet out of the top twenty results, only one half turn out to be relevant to the user. One reason for this is that Web queries are in general very short and give an incomplete specification of individual users' information needs. This paper explores ways of incorporating users' interests into the search process to improve the results. The user profiles are structured as a concept hierarchy of 4,400 nodes. These are populated by 'watching over a user's shoulder' while he is surfing. No explicit feedback is necessary. The profiles are shown to converge and to reflect the actual interests quite well. One possible deployment of the profiles is investigated: re-ranking and filtering search results. Increases in performance are moderate but noticeable and show that fully automatic creation of large hierarchical user profiles is possible. Alexander Pretschner, Susan Gauch |
ICTAI | 1 |