Yeongwoo Kim

dblp:280/3612 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0009-0002-6228-9332ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Dynamic Alert Prioritization for Real-Time Situational Awareness: A Hidden Markov Model Framework With Active Learning
abstract
Real-time cyber situational awareness (SA) is crucial for effective and timely incident response. However, maintaining SA requires substantial human effort; security analysts must analyze large volumes of alerts, many of which are false positives triggered by anomaly-based intrusion detection systems (IDSs). Efficiently prioritizing these alerts is vital to enable analysts to focus on real threats without delay. In this paper, we present two key contributions designed to improve real-time SA. First, we propose modeling dynamic alert prioritization as an active learning problem in a hidden Markov model (HMM) with the objective to minimize the mean squared error (MSE) of the belief. We propose to use the uncertainty of the belief as a proxy for the MSE of the belief, and we develop two computationally tractable policies for choosing alerts to investigate. Second, we propose and evaluate a state space and an exploit space reduction method to reduce the computational complexity of the belief update. We use simulations on synthetic and real dependency graphs to evaluate the proposed policies. Our results show that the proposed investigation policies reduce the MSE of the belief by up to 50% compared to baseline policies, and they are robust to high false alert rates and to investigation errors. Our results also show that state space reduction can reduce the computation time by 85% without a significant increase in the belief MSE.
Yeongwoo Kim, György Dán
IEEE Trans. Dependable Secur. Comput.1
2024 Anomaly Detection in Security Logs using Sequence Modeling
abstract
As cyberattacks are becoming more sophisticated, automated activity logging and anomaly detection are becoming important tools for defending computer systems. Recent deep learning-based approaches have demonstrated promising results in cybersecurity contexts, typically using supervised learning combined with large amounts of labeled data. Self-supervised learning has seen growing interest as a method of training models because it does not require labeled training data, which can be difficult and expensive to collect. However, existing self-supervised approaches to anomaly detection in user authentication logs either suffer from low precision or rely on large pre-trained natural language models. This makes them slow and expensive both during training and inference. Building on previous works, we therefore propose an end-to-end trained self-supervised transformer-based sequence model for anomaly detection in user authentication events. Thanks in part to an adapted masked-language modeling (MLM) learning task and domain knowledge-based improvements to the anomaly detection method, our proposed model outperforms previous long short-term memory (LSTM)-based approaches at detecting red-team activity in the "Comprehensive, Multi-Source Cyber-Security Events" authentication event dataset, improving the area under the receiver operating characteristic curve (AUC) from 0.9760 to 0.9989 and achieving an average precision of 0.0410. Our work presents the first application of end-to-end trained self-supervised transformer models to user authentication data in a cybersecurity context, and demonstrates the potential of transformer-based approaches for anomaly detection.
Simon G. E. Gökstorp, Jakob Nyberg, Yeongwoo Kim, Pontus Johnson, György Dán
NOMS3
2024 Human-in-the-Loop Cyber Intrusion Detection Using Active Learning
abstract
Timely detection of cyber attacks is essential for minimizing attack impact, but it requires accurate real-time situational awareness (SA). In practice, SA is hampered by frequent false alerts from anomaly-based intrusion detection systems (IDS), causing alarm fatigue. Investigating alerts by humans can enhance SA, but it is resource-intensive and it is often unclear which alerts to prioritize. In this paper, we propose a framework for optimizing human-in-the-loop attack detection, consisting of three key components: 1) dynamic alert prioritization, which ranks alerts based on previous alerts and investigations, 2) human alert investigation, referring to the manual analysis of alerts, and 3) sequential hypothesis testing, a method that confirms a hypothesis based on incoming alerts, with pruned hidden Markov models (HMMs). We formulate the problem as that of active learning in an HMM, and we propose two alert prioritization policies, namely Max Ratio and Max KL. The proposed policies aim to select the most informative alerts based on historical data and prior investigations, thereby minimizing the detection time. Simulation results show that our proposed policies reduce the time to detection by up to 79% compared to a static baseline policy, while maintaining a target mean time between false detections (MTBFD).
Yeongwoo Kim, György Dán, Quanyan Zhu
IEEE Trans. Inf. Forensics Secur.1
2021 Dynamic Clustering in Federated Learning
abstract
In the resource management of wireless networks, Federated Learning has been used to predict handovers. However, non-independent and identically distributed data degrade the accuracy performance of such predictions. To overcome the problem, Federated Learning can leverage data clustering algorithms and build a machine learning model for each cluster. However, traditional data clustering algorithms, when applied to the handover prediction, exhibit three main limitations: the risk of data privacy breach, the fixed shape of clusters, and the non-adaptive number of clusters. To overcome these limitations, in this paper, we propose a three-phased data clustering algorithm, namely: generative adversarial network-based clustering, cluster calibration, and cluster division. We show that the generative adversarial network-based clustering preserves privacy. The cluster calibration deals with dynamic environments by modifying clusters. Moreover, the divisive clustering explores the different number of clusters by repeatedly selecting and dividing a cluster into multiple clusters. A baseline algorithm and our algorithm are tested on a time series forecasting task. We show that our algorithm improves the performance of forecasting models, including cellular network handover, by 43%.
Yeongwoo Kim, Ezeddin Al Hakim, Johan Haraldson, Henrik Eriksson, Jose Mairton B. da Silva Jr., Carlo Fischione
ICC1