VLDB 2026 Research / reviewers in the wild / expert
Winnie Mbaka
dblp:326/0364 · also Winnie Bahati Mbaka
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0001-6913-1971ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 3 first-author · 4 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Less is more: usefulness of data flow diagrams and large language models for security threat validationabstractThe arrival of recent cybersecurity standards has raised the bar for security assessments in organizations, but existing techniques require a high manual effort. Threat analysis and risk assessment are used to identify security threats for new or refactored systems. Still, there is a lack of definition-of-done, so identified threats have to be validated which slows down the analysis. Existing literature has focused on the overall effectiveness of threat analysis, but no previous work has investigated what material must the analysts use to effectively validate the identified security threats. We conduct a controlled experiment with practitioners to investigate whether having some analysis material (either the system's graphical model or LLM-generated advice) is better than none, and whether having both the system's graphical model and LLM-generated advice is better than having only one of them. We run a pilot of the experiment with 41 MSc students, a think-aloud study with three practitioners, and the experiment survey with 68 recruited practitioners. Our main findings suggest that, in terms of additional material needed for threat validation, less is more. We also find that participants perceived the graphical model as equally useful compared to LLMs and that, despite LLMs not always providing conclusive advice, practitioners still perceived it as somewhat useful. The experimental material and data analysis scripts is publicly available in a replication package. Winnie Mbaka, Katja Tuma |
Empir. Softw. Eng. | 1 |
| 2025 | Assessing the usefulness of Data Flow Diagrams for validating security threatsabstractThreat analysis is a pillar of security-by-design which plays an important role in the elicitation and refinement of security threats. In preparation for the analysis, a model of the system under analysis e.g., the Data Flow Diagram (DFD for short) is often created. Empirical measures of success are important for practitioners that are struggling to meet the current demands for expertise. But no previous work has investigated the role of these diagrams during the validation of identified security threats. This paper presents an experiment conducted with 98 students in two countries. We measured the impact of the DFD on the perceived and actual effectiveness of validating a list of identified security threats including both fabricated and actual threats. In presence of sequence diagrams, the participants perceived DFDs as more useful. However, when exposed to both a DFD and a sequence diagram, DFDs had no significant impact on the participants’ ability to validate security threats. Winnie Mbaka, Yunduo Wang, Tong Li 0001, Fabio Massacci, Katja Tuma |
Comput. Secur. | 1 |
| 2024 | New experimental design to capture bias using LLM to validate security threatsabstractThe usage of Large Language Models is already well understood in software engineering and security and privacy. Yet, little is known about the effectiveness of LLMs in threat validation or the possibility of biased output when assessing security threats for correctness. To mitigate this research gap, we present a pilot study investigating the effectiveness of chatGPT in the validation of security threats. One main observation made from the results was that chatGPT assessed bogus threats as realistic regardless of the assumptions provided which negated the feasibility of certain threats occurring. Winnie Mbaka |
EASE | 1 |
| 2024 | Does trainer gender make a difference when delivering phishing training? A new experimental design to capture biasabstractPhishing is the most common attack vector for initial access. Current defenses such as spam filters and self-reporting phishing are unfortunately insufficient. Past research has found that gender may impact the perception of risk and that background may impact an individual’s susceptibility to phishing threats. However, no previous research has empirically measured the role of the trainer’s gender in identifying and assessing the risk of phishing. To address this gap, we designed a novel experimental setup focused on the trainer and surveyed 145 students at two universities. By adopting a controlled approach with AI-generated trainers we measured (a) the effect of gender and background on the perception of the trainer and (b) the effect of gender and background on identifying and assessing phishing risks. We found that background has a significant impact on the identification and assessment of phishing risks and that no gender bias was present towards the trainer in either a technical or non-technical population. André Palheiros Da Silva, Winnie Mbaka, Johann Mayer, Jan-Willem Bullee, Katja Tuma |
EASE | 2 |
| 2024 | On the Measures of Success in Replication of Controlled Experiments with STRIDEabstractTo avoid costly security patching after software deployment, security-by-design techniques (e.g. threat analysis) are adopted in organizations to find and mitigate security issues before the system is ever implemented. Organizations are ramping up such (heavily manual) activities, but there is a global gap in the security workforce. Favorable performance indicators would result in cost savings for organizations with scarce security experts. However, past empirical studies were inconclusive regarding some performance indicators of threat analysis techniques, thus practitioners have little evidence for choosing the technique to adopt. To address this issue, we replicated a controlled experiment with STRIDE. Our study aimed to measure and compare the performance indicators (productivity and precision) of two STRIDE variants (per-element and per-interaction). Since we made some similar observations to the original study, we conclude that the two approaches are not different enough to make a practical impact. To this end, the choice of which variant to adopt should be informed by the needs of the organization performing threat analysis. We conclude by discussing some of the unexplored yet relevant topic domains in the context of STRIDE that will be considered in future work. Winnie Mbaka, Katja Tuma |
Int. J. Softw. Eng. Knowl. Eng. | 1 |