EDBT 2026 Demo / reviewers in the wild / expert
Florian Skopik
dblp:29/3125
· DBLP profile ↗
13ranked-venue papers in the field
5as first author
4since 2021 · last 2025
0000-0002-1922-7892ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 5Database Systems & Data Management · 4 (1 first)Information Retrieval & Web Search · 2 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)Other / Interdisciplinary · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cybersecurity Text Classification: Challenging the Perceived Superiority of LLMs Over Conventional Machine Learning
Dzenan Hamzic, Markus Wurzenberger, Florian Skopik, Max Landauer, Lukas Linauer, Andreas Rauber |
IEEE Big Data | 3 |
| 2024 | Semi-supervised Configuration and Optimization of Anomaly Detection Algorithms on Log DataabstractCyber threats are evolving rapidly, making anomaly detection (AD) in system log data increasingly important for detection of known and unknown attacks. The configuration of AD algorithms heavily depends on the data at hand. It often involves a complex feature selection process and the determination of parameters such as thresholds or window sizes. In many cases, configuration requires manual intervention by domain experts, which limits accessibility and effectiveness of AD algorithms. This work introduces a Configuration-Engine (CE), which employs a semi-supervised approach to automate the configuration process or optimize existing configurations. The CE utilizes statistical methods to identify log line properties to recognize meaningful tokens for AD methods to monitor. It categorizes variables by their characteristics and behavior over time, then specifies which log parts a detector should observe, and sets appropriate configuration parameters.The CE was evaluated using four different detectors. Evaluations on different Apache Access and audit datasets containing attack traces showed that the CE achieved an average precision of over 0.94 for Apache and over 0.79 for audit datasets, while maintaining high recall, competing with the performance of expert-crafted configurations. The optimization approach was able to strongly improve the precision of both the CE’s and the experts’ configurations for Apache data in 7 out of 16 cases. Furthermore, the CE’s configurations were significantly dissimilar to each other when generated on audit data, highlighting the importance of automated configuration. Viktor Beck, Max Landauer, Markus Wurzenberger, Florian Skopik, Andreas Rauber |
IEEE Big Data | 4 |
| 2024 | Evaluation and Comparison of Open-Source LLMs Using Natural Language Generation Quality MetricsabstractThe rapid advancement of Large Language Models (LLMs) has transformed natural language processing, yet comprehensive evaluation methods are necessary to ensure their reliability, particularly in Retrieval-Augmented Generation (RAG) tasks. This study aims to evaluate and compare the performance of open-source LLMs by introducing a rigorous evaluation framework. We benchmark 20 LLMs using a combination of established metrics such as BLEU, ROUGE, BERTScore, along with and a novel metric, RAGAS. The models were tested across two distinct datasets to assess their text generation quality. Our findings reveal that models like nous-hermes-2-solar-10.7b and mistral-7b-instruct-v0.1 consistently excel in tasks requiring strict instruction adherence and effective use of large contexts, while other models show areas for improvement. This research contributes to the field by offering a comprehensive evaluation framework that aids in selecting the most suitable LLMs for complex RAG applications, with implications for future developments in natural language processing and big data analysis. Dzenan Hamzic, Markus Wurzenberger, Florian Skopik, Max Landauer, Andreas Rauber |
IEEE Big Data | 3 |
| 2022 | A User and Entity Behavior Analytics Log Data Set for Anomaly Detection in Cloud ComputingabstractCyber criminals utilize compromised user accounts to gain access into otherwise protected systems without the need for technical exploits. User and Entity Behavior Analytics (UEBA) leverages anomaly detection techniques to recognize such intrusions by comparing user behavior patterns against profiles derived from historical log data. Unfortunately, hardly any real log data sets suitable for UEBA are publicly available, which prevents objective comparison and reproducibility of approaches. Synthetic data sets are only able to alleviate this problem to some extent, because simulations are unable to adequately induce the dynamic and unstable nature of real user behavior in generated log data. We therefore present a real system log data set from a cloud computing platform involving more than 5000 users and spanning over more than five years. To evaluate our data set for the scenario of account hijacking, we outline a method for attack injection and subsequently disclose the resulting manifestations with an adaptive anomaly detection mechanism. Max Landauer, Florian Skopik, Georg Höld, Markus Wurzenberger |
IEEE Big Data | 2 |
| 2019 | A Framework for Cyber Threat Intelligence Extraction from Raw Log DataabstractIntrusion Detection Systems (IDS) rely on the availability and correctness of Indicators of Compromise (IoC), i.e., artifacts such as IP addresses that are known to correspond to malicious system activities. However, the simple nature and limited validity of these indicators impairs protection against cyber threats. Tactics, Techniques and Procedures (TTP) provide abstract information on attacker behavior, but are only available in human-readable format that prevents automatic detection using IDSs. In this paper we therefore propose an approach that extracts cyber threat intelligence from raw log data and combines the advantages of IoCs and TTPs by producing detectable patterns of complex system behavior. Other than existing approaches, our approach employs log data anomaly detection to disclose suspicious log events, which are used for iterative clustering, pattern recognition, and refinement. Our evaluations show that automatically extracted threat intelligence corresponding to a multi-step attack is suitable for detection of the same attack on another system. Max Landauer, Florian Skopik, Markus Wurzenberger, Wolfgang Hotwagner, Andreas Rauber |
IEEE BigData | 2 |
| 2016 | Complex log file synthesis for rapid sandbox-benchmarking of security- and computer network analysis tools
Markus Wurzenberger, Florian Skopik, Giuseppe Settanni, Wolfgang Scherrer |
Inf. Syst. | 2 |
| 2012 | Discovering and Managing Social Compositions in Collaborative Enterprise Crowdsourcing SystemsabstractCrowdsourcing is an increasingly used model to outsource certain tasks to be carried out by external experts on the Web. Especially when lacking experience or expertise with certain task types, crowdsourcing offers a convenient way to receive instant support. In this paper, we introduce an in-house enterprise crowdsourcing model, which leverages the crowdsourcing concept and transfers it to traditional organizations. Here, a company's staff is considered a crowd that — besides its regularly assigned tasks — can also receive tasks from colleagues from other departments and across hierarchical structures. The aim is to offer instant support and utilize free capacities throughout a large organization more efficiently. In our work, we describe this concept and supporting mechanisms in context of an agile software development use case. However, in contrast to usually crowdsourced microtasks, complex software architectures usually consist of tens and hundreds of connected modules that can be potentially crowdsourced. These technical dependencies between modules require active coordination and interactions between crowd members that process the single artifacts. Hence, technical dependencies of artifacts result in social dependencies of collaborating crowd members that create them. In order to efficiently discover member compositions based on artifact dependencies, we introduce an indexing and discovery approach based on subgraph matching. Typically, assigning tasks to well-rehearsed teams results in more reliable task processing, faster results, and higher quality of work. We evaluate our approach in terms of system scalability and overall applicability by mining and analyzing the popular SourceForge community. We show that our approach of member composition discovery is feasibly in terms of scalability and quality of discovery results. Our findings deliver important input for the design and implementation of supporting information systems for future large-scale collaboration platforms. Florian Skopik, Daniel Schall 0001, Schahram Dustdar |
Int. J. Cooperative Inf. Syst. | 1 |
| 2011 | An Analysis of the Structure and Dynamics of Large-Scale Q/A Communities
Daniel Schall 0001, Florian Skopik |
ADBIS | 2 |
| 2011 | Opportunistic Information Flows through Strategic Social Link EstablishmentabstractSocial networks have emerged from niche existence to a mass phenomenon. Nowadays, their fundamental concepts, such as managing personal contacts and sharing profile information, are increasingly harnessed for businesses in professional environments. Similar to service-oriented networks, they allow flexible discovery on demand and loose coupling of participants. Establishing social links facilitates cooperation and enables selective sharing of information. Intuitively, one shares more information with his connected neighbors and less or even none with unrelated individuals. Today, information is one of the most important and valuable goods in business networks. Being informed about ongoing collaborations and upcoming trends is a key success factor. Thus, in professional networks, participants aim at strategically establishing connections to enable reliable information flows. In this paper, we especially highlight an opportunistic model that let mediators connect actually unrelated actors in order to benefit from information mediation. We further discuss a framework that implements this model for service-oriented professional virtual communities. Florian Skopik, Daniel Schall 0001, Schahram Dustdar |
Web Intelligence | 1 |
| 2011 | Interaction mining and skill-dependent recommendations for multi-objective team compositionabstractWeb-based collaboration and virtual environments supported by various Web 2.0 concepts enable the application of numerous monitoring, mining and analysis tools to study human interactions and team formation processes. The composition of an effective team requires a balance between adequate skill fulfillment and sufficient team connectivity. The underlying interaction structure reflects social behavior and relations of individuals and determines to a large degree how well people can be expected to collaborate. In this paper we address an extended team formation problem that does not only require direct interactions to determine team connectivity but additionally uses implicit recommendations of collaboration partners to support even sparsely connected networks. We provide two heuristics based on Genetic Algorithms and Simulated Annealing for discovering efficient team configurations that yield the best trade-off between skill coverage and team connectivity. Our self-adjusting mechanism aims to discover the best combination of direct interactions and recommendations when deriving connectivity. We evaluate our approach based on multiple configurations of a simulated collaboration network that features close resemblance to real world expert networks. We demonstrate that our algorithm successfully identifies efficient team configurations even when removing up to 40% of experts from various social network configurations. Christoph Mayr-Dorn, Florian Skopik, Daniel Schall 0001, Schahram Dustdar |
Data Knowl. Eng. | 2 |
| 2010 | Modeling and mining of dynamic trust in complex service-oriented systems
Florian Skopik, Daniel Schall 0001, Schahram Dustdar |
Inf. Syst. | 1 |
| 2009 | Trust and Reputation Mining in Professional Virtual Communities
Florian Skopik, Hong Linh Truong 0001, Schahram Dustdar |
ICWE | 1 |
| 2009 | Start Trusting Strangers? Bootstrapping and Prediction of Trust
Florian Skopik, Daniel Schall 0001, Schahram Dustdar |
WISE | 1 |