Katja Tuma

dblp:211/8144 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0001-7189-2817ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 13 · 6 first-author · 8 since 2021Security and privacy · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Less is more: usefulness of data flow diagrams and large language models for security threat validation
abstract
The arrival of recent cybersecurity standards has raised the bar for security assessments in organizations, but existing techniques require a high manual effort. Threat analysis and risk assessment are used to identify security threats for new or refactored systems. Still, there is a lack of definition-of-done, so identified threats have to be validated which slows down the analysis. Existing literature has focused on the overall effectiveness of threat analysis, but no previous work has investigated what material must the analysts use to effectively validate the identified security threats. We conduct a controlled experiment with practitioners to investigate whether having some analysis material (either the system's graphical model or LLM-generated advice) is better than none, and whether having both the system's graphical model and LLM-generated advice is better than having only one of them. We run a pilot of the experiment with 41 MSc students, a think-aloud study with three practitioners, and the experiment survey with 68 recruited practitioners. Our main findings suggest that, in terms of additional material needed for threat validation, less is more. We also find that participants perceived the graphical model as equally useful compared to LLMs and that, despite LLMs not always providing conclusive advice, practitioners still perceived it as somewhat useful. The experimental material and data analysis scripts is publicly available in a replication package.
Winnie Mbaka, Katja Tuma
Empir. Softw. Eng.2
2026 Automated Analysis of Security Policy Violations in Helm Charts
abstract
The advent of Infrastructure-as-Code (IaC) and cloud platforms has transformed applications into ephemeral deployments of configuration files, where containers live for only a few minutes. Several industry-level static analyzers are available to check security misconfigurations before deployment, but the experimental evidence that we report in this paper is that they provide different and possibly inconsistent results. We developed an automated pipeline to evaluate and compare static analyzers for Helm charts, a popular package manager to deploy Kubernetes (K8s) applications, in finding a functional configuration adhering to the principle of least privilege. We evaluated seven open-source chart analyzer tools on the 60 most common Artifact Hub Helm charts (returned by the Application Programming Interface — API upon first invocation) and found that overly permissive ClusterRoles are the most common misconfiguration, and using a high user ID is the most commonly needed permission. During the evaluation, we also found several bugs, both false positives and negatives, that we reported to the tool developers. Securing cloud configurations still requires significant manual intervention, and more effort should be spent on standardizing the analysis of misconfiguration.
Francesco Minna, Agathe Blaise, Katja Tuma, Fabio Massacci
IEEE Trans. Dependable Secur. Comput.3
2025 Large Language Models Are Unreliable for Cyber Threat Intelligence
Emanuele Mezzi, Fabio Massacci, Katja Tuma
ARES (2)3
2025 In Specs We Trust? Conformance-Analysis of Implementation to Specifications in Node-RED and Associated Security Risks
Simon Schneider, Komal Kashish, Katja Tuma, Riccardo Scandariato
ARES (1)3
2025 Assessing the usefulness of Data Flow Diagrams for validating security threats
abstract
Threat analysis is a pillar of security-by-design which plays an important role in the elicitation and refinement of security threats. In preparation for the analysis, a model of the system under analysis e.g., the Data Flow Diagram (DFD for short) is often created. Empirical measures of success are important for practitioners that are struggling to meet the current demands for expertise. But no previous work has investigated the role of these diagrams during the validation of identified security threats. This paper presents an experiment conducted with 98 students in two countries. We measured the impact of the DFD on the perceived and actual effectiveness of validating a list of identified security threats including both fabricated and actual threats. In presence of sequence diagrams, the participants perceived DFDs as more useful. However, when exposed to both a DFD and a sequence diagram, DFDs had no significant impact on the participants’ ability to validate security threats.
Winnie Mbaka, Yunduo Wang, Tong Li 0001, Fabio Massacci, Katja Tuma
Comput. Secur.6
2025 Analyzing and mitigating (with LLMs) the security misconfigurations of Helm charts from Artifact Hub
abstract
Helm is a package manager that allows defining, installing, and upgrading applications with Kubernetes (K8s), a popular container orchestration platform. A Helm chart is a collection of files describing all dependencies, resources, and parameters required for deploying an application within a K8s cluster. This study aimed to mine and empirically evaluate the security of Helm charts, comparing the performance of existing tools in terms of misconfigurations reported by policies available by default, and measuring to what extent LLMs could be used for removing misconfigurations. For these reasons, we proposed a pipeline to mine Helm charts from Artifact Hub, a popular centralized repository, and analyze them using state-of-the-art open-source tools like Checkov and KICS. First, the pipeline runs several chart analyzers and identifies the common and unique misconfigurations reported by each tool. Secondly, it uses LLMs to suggest a mitigation for each misconfiguration. Finally, the LLM refactored chart previously generated is analyzed again by the same tools to see whether it satisfies the tool's policies. We also performed a manual analysis on a subset of charts to evaluate whether there are false positive misconfigurations from the tool's reporting and in the LLM refactoring. We found that (i) there is a significant difference between LLMs, (ii) providing a snippet of the YAML template as input might be insufficient compared to all resources, and (iii) even though LLMs can generate correct fixes, they may also delete other irrelevant configurations that break the application.
Francesco Minna, Fabio Massacci, Katja Tuma
Empir. Softw. Eng.3
2025 On the effects of program slicing for vulnerability detection during code inspection
abstract
Abstract Slicing is a fault localization technique that has been proposed to support debugging and program comprehension. Yet, its empirical effectiveness during code inspection by humans has received limited attention. The goal of our study is two-fold. First, we aim to define what it means for a code reviewer to identify the vulnerable lines correctly. Second, we investigate whether reducing the number of to-be-inspected lines by method-level slicing supports code reviewers in detecting security vulnerabilities. We propose a novel approach based on the notion of a $$\delta $$ δ -neighborhood (intuitively based on the idea of the context size of the command ) to define correctly identified lines. Then, we conducted a multi-year controlled experiment (2017-2023) in which MSc students attending security courses ( $$n=236$$ n = 236 ) were tasked with identifying vulnerable lines in original or sliced Java files from Apache Tomcat. We provide perfect seed lines for a slicing algorithm to control for confounding factors. Each treatment differs in the pair (Vulnerability, Original/Sliced) with a balanced design with vulnerabilities from the OWASP Top 10 2017: A1 (Injection), A5 (Broken Access Control), A6 (Security Misconfiguration), and A7 (Cross-Site Scripting). To generate smaller slices for human consumption, we used a variant of intra-procedural thin slicing. We report the results for $$\delta = 0$$ δ = 0 which corresponds to exactly matching the vulnerable ground truth lines, and $$\delta = 3$$ δ = 3 which represents the scenario of identifying the vulnerable area. For both cases, we found that slicing helps in ‘finding something’ (the participant has found at least some vulnerable lines) as opposed to ‘finding nothing’. For the case of $$\delta = 0$$ δ = 0 analyzing a slice and analyzing the original file are statistically equivalent from the perspective of lines found by those who found something. With $$\delta = 3$$ δ = 3 slicing helps to find more vulnerabilities compared to analyzing an original file, as we would normally expect. Given the type of population, additional experiments are necessary to be generalized to experienced developers.
Aurora Papotti, Katja Tuma, Fabio Massacci
Empir. Softw. Eng.2
2024 Does trainer gender make a difference when delivering phishing training? A new experimental design to capture bias
abstract
Phishing is the most common attack vector for initial access. Current defenses such as spam filters and self-reporting phishing are unfortunately insufficient. Past research has found that gender may impact the perception of risk and that background may impact an individual’s susceptibility to phishing threats. However, no previous research has empirically measured the role of the trainer’s gender in identifying and assessing the risk of phishing. To address this gap, we designed a novel experimental setup focused on the trainer and surveyed 145 students at two universities. By adopting a controlled approach with AI-generated trainers we measured (a) the effect of gender and background on the perception of the trainer and (b) the effect of gender and background on identifying and assessing phishing risks. We found that background has a significant impact on the identification and assessment of phishing risks and that no gender bias was present towards the trainer in either a technical or non-technical population.
André Palheiros Da Silva, Winnie Mbaka, Johann Mayer, Jan-Willem Bullee, Katja Tuma
EASE5
2024 The Equality Maturity Model: An Actionable Tool to Advance Gender Balance in Leadership and Participation Roles
abstract
The underrepresentation of women in Computer Science and Engineering is a pervasive issue, impacting the enrolment and graduation rates of female students as well as the presence of women in leadership positions in academia and industry. The European Network For Gender Balance in Informatics (EUGAIN) COST action seeks to share data, experiences, best practices, and lessons from failures, and to provide actionable tools that may contribute to the advancement of gender balance in the field. This paper summarises results from the Ph.D./Postdoc to Professor workgroup that were gathered in two booklets of best practices. Specifically, we introduce the Equality Maturity Model (EMM), a conceptual tool aimed at supporting organisations in measuring how they are doing concerning equality and identifying potential areas of improvement and that was inspired by both booklets.
Paloma Díaz 0001, Paula Alexandra Silva, Katja Tuma
EDUCON3
2024 On the Measures of Success in Replication of Controlled Experiments with STRIDE
abstract
To avoid costly security patching after software deployment, security-by-design techniques (e.g. threat analysis) are adopted in organizations to find and mitigate security issues before the system is ever implemented. Organizations are ramping up such (heavily manual) activities, but there is a global gap in the security workforce. Favorable performance indicators would result in cost savings for organizations with scarce security experts. However, past empirical studies were inconclusive regarding some performance indicators of threat analysis techniques, thus practitioners have little evidence for choosing the technique to adopt. To address this issue, we replicated a controlled experiment with STRIDE. Our study aimed to measure and compare the performance indicators (productivity and precision) of two STRIDE variants (per-element and per-interaction). Since we made some similar observations to the original study, we conclude that the two approaches are not different enough to make a practical impact. To this end, the choice of which variant to adopt should be informed by the needs of the organization performing threat analysis. We conclude by discussing some of the unexplored yet relevant topic domains in the context of STRIDE that will be considered in future work.
Winnie Mbaka, Katja Tuma
Int. J. Softw. Eng. Knowl. Eng.2
2023 A new, evidence-based, theory for knowledge reuse in security risk analysis
abstract
Abstract Security risk analysis (SRA) is a key activity in software engineering but requires heavy manual effort. Community knowledge in the form of security patterns or security catalogs can be used to support the identification of threats and security controls. However, no evidence-based theory exists about the effectiveness of security catalogs when used for security risk analysis. We adopt a grounded theory approach to propose a conceptual, revised and refined theory of SRA knowledge reuse. The theory refinement is backed by evidence gathered from conducting interviews with experts (20) and controlled experiments with both experts (15) and novice analysts (18). We conclude the paper by providing insights into the use of catalogs and managerial implications.
Katsiaryna Labunets, Fabio Massacci, Federica Paci, Katja Tuma
Empir. Softw. Eng.4
2023 Checking security compliance between models and code
abstract
Abstract It is challenging to verify that the planned security mechanisms are actually implemented in the software. In the context of model-based development, the implemented security mechanisms must capture all intended security properties that were considered in the design models. Assuring this compliance manually is labor intensive and can be error-prone. This work introduces the first semi-automatic technique for secure data flow compliance checks between design models and code. We develop heuristic-based automated mappings between a design-level model (SecDFD, provided by humans) and a code-level representation (Program Model, automatically extracted from the implementation) in order to guide users in discovering compliance violations, and hence, potential security flaws in the code. These mappings enable an automated, and project-specific static analysis of the implementation with respect to the desired security properties of the design model. We developed two types of security compliance checks and evaluated the entire approach on open source Java projects.
Katja Tuma, Sven Peldszus, Daniel Strüber 0001, Riccardo Scandariato, Jan Jürjens
Softw. Syst. Model.1
2022 Precise Analysis of Purpose Limitation in Data Flow Diagrams
abstract
Data Flow Diagrams (DFDs) are primarily used for modelling functional properties of a system. In recent work, it was shown that DFDs can be used to also model non-functional properties, such as security and privacy properties, if they are annotated with appropriate security- and privacy-related information. An important privacy principle one may wish to model in this way is purpose limitation. But previous work on privacy-aware DFDs (PA-DFDs) considers purpose limitation only superficially, without explaining how the purpose of DFD activators and flows ought to be specified, checked or inferred. In this paper, we define a rigorous formal framework for (1) annotating DFDs with purpose labels and privacy signatures, (2) checking the consistency of labels and signatures, and (3) inferring labels from signatures. We implement our theoretical framework in a proof-of concept tool consisting of a domain-specific language (DSL) for specifying privacy signatures and algorithms for checking and inferring purpose labels from such signatures. Finally, we evaluate our framework and tool through a case study based on a DFD from the privacy literature.
Hanaa Alshareef, Katja Tuma, Sandro Stucki, Gerardo Schneider, Riccardo Scandariato
ARES2
2022 Towards a Security Stress-Test for Cloud Configurations
abstract
Securing cloud configurations is an elusive task, which is left up to system administrators who have to base their decisions on "trial and error" experimentations or by observing good practices (e.g., CIS Benchmarks). We propose a knowledge, AND/OR, graphs approach to model cloud deployment security objects and vulnerabilities. In this way, we can capture relationships between configurations, permissions (e.g., CAP_SYS_ADMIN), and security profiles (e.g., AppArmor and SecComp). Such an approach allows us to suggest alternative and safer configurations, support administrators in the study of what-if scenarios, and scale the analysis to large scale deployments. We present an initial validation and illustrate the approach with three real vulnerabilities from known sources.
Francesco Minna, Fabio Massacci, Katja Tuma
CLOUD3
2021 Finding security threats that matter: Two industrial case studies
Katja Tuma, Christian Sandberg, Urban Thorsson, Mathias Widman, Thomas Herpel, Riccardo Scandariato
J. Syst. Softw.1
2020 Automating the early detection of security design flaws
abstract
Security by design is a key principle for realizing secure software systems and it is advised to hunt for security flaws from the very early stages of development. At design-time, security analysis is often performed manually by means of either threat modeling or expert-based design inspections. However, when leveraging the wide range of established knowledge bases on security design flaws (e.g., CWE, CAWE), these manual assessments become too time consuming, error-prone, and infeasible in the context of contemporary development practices with frequent iterations. This paper focuses on design inspection and explores the potential for automating the application of inspection rules to speed up the security analysis.
Katja Tuma, Laurens Sion, Riccardo Scandariato, Koen Yskout
MoDELS1
2019 Flaws in Flows: Unveiling Design Flaws via Information Flow Analysis
abstract
This paper presents a practical and formal approach to analyze security-centric information flow policies at the level of the design model. Specifically, we focus on data confidentiality and data integrity objectives. In its guiding principles, the approach is meant to be amenable for designers (e.g., software architects) that have very limited or no background in formal models, logics, and the like. To this aim, we provide an intuitive graphical notation, which is based on the familiar Data Flow Diagrams, and which requires as little effort as possible in terms of extra security-centric information the designer has to provide. The result of the analysis algorithm is the early discovery of design flaws in the form of violations of the intended security properties. The approach is implemented as a publicly available plugin for Eclipse and evaluated with four real-world case studies from publicly available literature.
Katja Tuma, Riccardo Scandariato, Musard Balliu
ICSA1
2019 Secure Data-Flow Compliance Checks between Models and Code Based on Automated Mappings
abstract
During the development of security-critical software, the system implementation must capture the security properties postulated by the architectural design. This paper presents an approach to support secure data-flow compliance checks between design models and code. To iteratively guide the developer in discovering such compliance violations we introduce automated mappings. These mappings are created by searching for correspondences between a design-level model (Security Data Flow Diagram) and an implementation-level model (Program Model). We limit the search space by considering name similarities between model elements and code elements as well as by the use of heuristic rules for matching data-flow structures. The main contributions of this paper are three-fold. First, the automated mappings support the designer in an early discovery of implementation absence, convergence, and divergence with respect to the planned software design. Second, the mappings also support the discovery of secure data-flow compliance violations in terms of illegal asset flows in the software implementation. Third, we present our implementation of the approach as a publicly available Eclipse plugin and its evaluation on five open source Java projects (including Eclipse secure storage).
Sven Peldszus, Katja Tuma, Daniel Strüber 0001, Jan Jürjens, Riccardo Scandariato
MoDELS2
2018 Two Architectural Threat Analysis Techniques Compared
Katja Tuma, Riccardo Scandariato
ECSA1
2018 Back to the Drawing Board - Bringing Security Constraints in an Architecture-centric Software Development Process
abstract
Today, security is still poorly considered in early phases of software engineering. Architects and software engineers still lack knowledge about architectural security design as well as implementing it compliantly. However, a software system that is not designed for security or does not adhere to this design can hardly meet its security requirements. In this paper, we present an approach we are working on. The approach consists of two parts: Firstly, we improve the architecture’s security level through model transformation. Secondly, we derive rules and constraints from the secured architecture in order to check the implementation’s conformance. Through these activities we aim to support architects and software developers in building a secure software system. We plan to evaluate our approach in industrial case studies.
Stefanie Jasser, Katja Tuma, Riccardo Scandariato, Matthias Riebisch
ICISSP2
2018 Threat analysis of software systems: A systematic literature review
Katja Tuma, Gül Çalikli, Riccardo Scandariato
J. Syst. Softw.1