MyungKeun Yoon 0001

dblp:36/6231-1 · also Myungkeun Yoon 0001 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
5since 2021 · last 2024
0000-0003-1987-1394ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 5 first-author · 4 since 2021Systems, architecture and hardware · 4 · 2 first-authorSecurity and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2024 Detecting Internet of Things Malware on Evidence Generation
abstract
Malware has been a real threat to Internet of Things (IoT). Although commercial antivirus solutions can detect malware files and provide label information indicating malware types or families, no clear evidence explaining the detection is provided. Therefore, even security experts using the antivirus solutions do not know why some files are reported malicious and they hesitate to take an immediate action. In this article, we study this problem from the viewpoint of antivirus solution users instead of product developers or sellers. We present a new data-driven scheme that can automatically generate a set of readable common strings from the IoT malware files as a detection evidence. These generated string signatures not only provide a clear detection evidence for suspicious files but can also be used as unique high-precision detection criteria. The new data-driven scheme divides any long evasive string embedded in malware files into short n-grams to mitigate the detection evasion, and a limited number of n-grams are selected as representative n-grams on a bipartite graph that improves the efficiency and accuracy of clustering. A set of n-grams per cluster, which plays the role of an unique detection evidence is generated. Through experiments with the real malware data sets, including the public data sets for the experimental reproducibility, we confirm that the new data-driven scheme not only detects malware files as accurately as the current state-of-the-art (SOTA), especially no benign files mistakenly considered as malicious but also provides readable strings as a detection evidence, which has not been achieved by the previous work.
YoonSeok Han, HyungBin Seo, MyungKeun Yoon 0001
IEEE Internet Things J.3
2024 Relative Frequency-Rank Encoding for Unsupervised Network Anomaly Detection
abstract
Network-based anomaly detection plays a pivotal role in cybersecurity. Most detection models are based on unsupervised machine learning to learn such a normal flow pattern of network traffic as the numbers of incoming/outgoing packets, traffic volumes in bytes, duration time, etc., most of which are numerical features. On the contrary, non-numerical features have not been fully utilized yet although they often give a decisive hint to the detection of unseen attacks; for example, rarely observed combinations of IP addresses and port numbers can reveal an uncommon attack attempt. This heuristic has already been used by human experts for decades, but not fully utilized yet by deep learning models. In this paper, we present a new encoding scheme for non-numerical features such as IP addresses and port numbers that might have been mistakenly considered as numerical features. The new encoding scheme first ranks non-numerical features in their frequency order and then evenly places each rank between 0 and 1, which transforms raw data into a form that is easy for machines to understand. The anomaly detection performance is significantly improved when this new encoding scheme is applied to the same deep learning model. For example, a simple autoencoder model with the new encoding scheme achieved the Area Under Receiver Operating Characteristic, AUROC, of 0.99 for the well-known CICIDS2017 dataset while the previous record was 0.91. Experimental results from three different open datasets show that the proposed encoding scheme can significantly enhance the performance of anomaly detection models.
Minsong Kim, Woohyuk Jang, JunNyung Hur, MyungKeun Yoon 0001
IEEE/ACM Trans. Netw.4
2023 Generative Intrusion Detection and Prevention on Data Stream
HyungBin Seo, MyungKeun Yoon 0001
USENIX Security Symposium2
2023 Packet Chunking for File Detection
abstract
Network-based intrusion detection and data leakage prevention systems inspect packets to detect if critical files such as malware or confidential documents are transferred. However, this kind of detection requires heavy computing resources in reassembling packets and only well-known protocols can be interpreted. Besides, finding similar files from a storage requires pairwise comparisons. In this paper, we present a new network-based file identification scheme that inspects packets independently without reassembly and finds similar files through inverted indexing instead of pairwise comparison. We use a content-based chunking algorithm to consistently divide both files and packets into multiple byte sequences, called chunks. If a packet is a part of a file, they would have common chunks. The challenging problem is that packet chunking and inverted-index search should be fast and scalable enough for packet processing. The file identification should be accurate although many chunks are noises. In this paper, we use a small Bloom filter and a two-level threshold strategy to solve the problems. To the best of our knowledge, this is the first scheme that identifies a specific critical file from a packet over unknown protocols. Experimental results show that the proposed scheme can successfully identify a critical file from a packet without packet reassembly.
JunNyung Hur, Hyeon Gy Shon, Young Jae Kim, MyungKeun Yoon 0001
IEEE/ACM Trans. Netw.4
2021 Finding Critical Files from a Packet
abstract
Network-based intrusion detection and data leakage prevention systems inspect packets to detect if critical files such as malware or confidential documents are transferred. However, this kind of detection requires heavy computing resources in reassembling packets and only well-known protocols can be interpreted. Besides, finding similar files from a storage requires pairwise comparisons. In this paper, we present a new network-based file identification scheme that inspects packets independently without reassembly and finds similar files through inverted indexing instead of pairwise comparison. We use a contents-based chunking algorithm to consistently divide both files and packets into multiple byte sequences, called chunks. If a packet is a part of a file, they would have common chunks. The challenging problem is that packet chunking and inverted-index search should be fast and scalable enough for packet processing. The file identification should be accurate although many chunks are noises. In this paper, we use a small Bloom filter and a delayed query strategy to solve the problems. To the best of our knowledge, this is the first scheme that identifies a specific critical file from a packet over unknown protocols. Experimental results show that the proposed scheme can successfully identify a critical file from a packet.
JunNyung Hur, Hahoon Jeon, Hyeon Gy Shon, Young Jae Kim, MyungKeun Yoon 0001
INFOCOM5
2016 When Bloom Filters Are No Longer Compact: Multi-Set Membership Lookup for Network Applications
abstract
Many important network functions require online membership lookup against a large set of addresses, flow labels, signatures, and so on. This paper studies a more difficult, yet less investigated problem, called multi-set membership lookup, which involves multiple (sometimes in hundreds or even thousands) sets. The lookup determines not only whether an element is a member of the sets but also which set it belongs to. To facilitate the implementation of multi-set membership lookup in on-die memory of a network processor for line-speed packet inspection, the existing work uses the variants of Bloom filters to encode set IDs. However, through a thorough analysis of the mechanism and the performance of the prior art, much to our surprise, we find that Bloom filters-which were originally designed for encoding binary membership information-are actually not efficient for encoding set IDs. This paper takes a different solution path by separating membership encoding and set ID storage in two data structures, called index filter and set-id table, respectively. With a new ID placement strategy called uneven candidate-entry distribution and a two-level design of an index filter, we demonstrate through analysis and simulation that when compared with the best existing work, our new approach is able to achieve significant memory saving under the same lookup accuracy requirement, or achieve significantly better lookup accuracy under the same memory constraint.
Shigang Chen, Zhen Mo, MyungKeun Yoon 0001
IEEE/ACM Trans. Netw.4
2014 Bloom tree: A search tree based on Bloom filters for multiple-set membership testing
abstract
A Bloom filter is a compact and randomized data structure popularly used for networking applications. A standard Bloom filter only answers yes/no questions about membership, but recent studies have improved it so that the value of a queried item can be returned, supporting multiple-set membership testing. In this paper, we design a new data structure for multiple-set membership testing, Bloom tree, which not only achieves space compactness, but also operates more efficiently than existing ones. For example, when existing work requires 107 bits per item and 11 memory accesses for a search operation, a Bloom tree requires only 47 bits and 8 memory accesses. The advantages come from a new data structure that consists of multiple Bloom filters in a tree structure. We study a theoretical analysis model to find optimal parameters for Bloom trees, and its effectiveness is verified through experiments.
MyungKeun Yoon 0001, JinWoo Son, Seon-Ho Shin
INFOCOM1
2014 A grand spread estimator using a graphics processing unit
Seon-Ho Shin, Eun-Jin Im, MyungKeun Yoon 0001
J. Parallel Distributed Comput.3
2012 An incrementally deployable path address scheme
MyungKeun Yoon 0001, Shigang Chen
J. Parallel Distributed Comput.1
2012 An efficient incentive scheme with a distributed authority infrastructure in peer-to-peer networks
Shigang Chen, Zhen Mo, MyungKeun Yoon 0001
J. Parallel Distributed Comput.4
2011 Fit a Compact Spread Estimator in Small High-Speed Memory
abstract
The spread of a source host is the number of distinct destinations that it has sent packets to during a measurement period. A spread estimator is a software/hardware module on a router that inspects the arrival packets and estimates the spread of each source. It has important applications in detecting port scans and distributed denial-of-service (DDoS) attacks, measuring the infection rate of a worm, assisting resource allocation in a server farm, determining popular Web contents for caching, to name a few. The main technical challenge is to fit a spread estimator in a fast but small memory (such as SRAM) in order to operate it at the line speed in a high-speed network. In this paper, we design a new spread estimator that delivers good performance in tight memory space where all existing estimators no longer work. The new estimator not only achieves space compactness, but operates more efficiently than the existing ones. Its accuracy and efficiency come from a new method for data storage, called virtual vectors, which allow us to measure and remove the errors in spread estimation. We also propose several ways to enhance the range of spread values that the estimator can measure. We perform extensive experiments on real Internet traces to verify the effectiveness of the new estimator .
MyungKeun Yoon 0001, Tao Li 0013, Shigang Chen, Jih-Kwon Peir
IEEE/ACM Trans. Netw.1
2010 Minimizing the Maximum Firewall Rule Set in a Network with Multiple Firewalls
abstract
A firewall's complexity is known to increase with the size of its rule set. Empirical studies show that as the rule set grows larger, the number of configuration errors on a firewall increases sharply, while the performance of the firewall degrades. When designing a security-sensitive network, it is critical to construct the network topology and its routing structure carefully in order to reduce the firewall rule sets, which helps lower the chance of security loopholes and prevent performance bottleneck. This paper studies the problems of how to place the firewalls in a topology during network design and how to construct the routing tables during operation such that the maximum firewall rule set can be minimized. These problems have not been studied adequately despite their importance. We have two major contributions. First, we prove that the problems are NP-complete. Second, we propose a heuristic solution and demonstrate the effectiveness of the algorithm by simulations. The results show that the proposed algorithm reduces the maximum firewall rule set by 2-5 times when comparing with other algorithms.
MyungKeun Yoon 0001, Shigang Chen
IEEE Trans. Computers1
2010 Aging Bloom Filter with Two Active Buffers for Dynamic Sets
abstract
A Bloom filter is a simple but powerful data structure that can check membership to a static set. As Bloom filters become more popular for network applications, a membership query for a dynamic set is also required. Some network applications require high-speed processing of packets. For this purpose, Bloom filters should reside in a fast and small memory, SRAM. In this case, due to the limited memory size, stale data in the Bloom filter should be deleted to make space for new data. Namely the Bloom filter needs aging like LRU caching. In this paper, we propose a new aging scheme for Bloom filters. The proposed scheme utilizes the memory space more efficiently than double buffering, the current state of the art. We prove theoretically that the proposed scheme outperforms double buffering. We also perform experiments on real Internet traces to verify the effectiveness of the proposed scheme.
MyungKeun Yoon 0001
IEEE Trans. Knowl. Data Eng.1
2009 Fit a Spread Estimator in Small Memory
abstract
The spread of a source host is the number of distinct destinations that it has sent packets to during a measurement period. A spread estimator is a software/hardware module on a router that inspects the arrival packets and estimates the spread of each source. It has important applications in detecting port scans and DDoS attacks, measuring the infection rate of a worm, assisting resource allocation in a server farm, determining popular Web contents for caching, to name a few. The main technical challenge is to fit a spread estimator in a fast but small memory (such as SRAM) in order to operate it at the line speed in a high-speed network. In this paper, we design a new spread estimator that delivers good performance in tight memory space where all existing estimators no longer work. The new estimator not only achieves space compactness but operates more efficiently than the existing ones. Its accuracy and efficiency come from a new method for data storage, called virtual vectors, which allow us to measure and remove the errors in spread estimation. We perform experiments on real Internet traces to verify the effectiveness of the new estimator.
MyungKeun Yoon 0001, Tao Li 0013, Shigang Chen, Jih-Kwon Peir
INFOCOM1
2008 Real-Time Detection of Invisible Spreaders
abstract
Detecting spreaders can help an intrusion detection system identify potential attackers. The existing work can only detect aggressive spreaders that scan a large number of distinct addresses in a short period of time. However, stealthy spreaders may perform scanning deliberately at a low rate. We observe that these spreaders can easily evade the detection because their small traffic footprint will be covered by the large amount of background normal traffic that frequently flushes any spreader information out of the intrusion detection system's memory. We propose a new streaming scheme to detect stealthy spreaders that are invisible to the current systems. The new scheme stores information about normal traffic within a limited portion of the allocated memory, so that it will not interfere with spreaders' information stored elsewhere in the memory. The proposed scheme is light weight; it can detect invisible spreaders in high-speed networks while residing in SRAM. Through experiments using real Internet traffic traces, we demonstrate that our new scheme detects invisible spreaders efficiently while keeping both false-positives (normal sources misclassified as spreaders) and false-negatives (spreaders misclassified as normal sources) to low level.
MyungKeun Yoon 0001, Shigang Chen
GLOBECOM1
2007 Reducing the Size of Rule Set in a Firewall
abstract
A firewall's complexity is known to increase with the size of its rule set. Complex firewalls are more likely to have configuration errors which cause security loopholes. Until now, two rules can be merged into one only when they are exactly same for all the dimensions except one for which each value of two rules should be adjacent to each other. In this paper, we propose a new and aggressive reduction algorithm which finds a group of rules and replace it with a smaller new group so that the total size of rule set can be reduced. This can not be achievable by any previous work because all of them eliminate rules only when these rules are redundant by other rules in the same rule set. The proposed algorithm is also orthogonal to the previous works so that it can be used to supplement them.
MyungKeun Yoon 0001, Shigang Chen
ICC1
2007 MARCH: A Distributed Incentive Scheme for Peer-to-Peer Networks
abstract
As peer-to-peer networks grow larger and include more diverse users, the lack of incentive to encourage cooperative behavior becomes one of the key problems. This challenge cannot be fully met by traditional incentive schemes, which suffer from various attacks based on false reports. Especially, due to the lack of central authorities in typical P2P systems, it is difficult to detect colluding groups. Members in the same colluding group can cooperate to manipulate their history information, and the damaging power increases dramatically with the group size. In this paper, we propose a new distributed incentive scheme, in which the benefit that a node can obtain from the system is proportional to its contribution to the system, and a colluding group cannot gain advantage by cooperation regardless of its size. Consequently, the damaging power of colluding groups is strictly limited. The proposed scheme includes three major components: a distributed authority infrastructure, a key sharing protocol, and a contract verification protocol.
Shigang Chen, MyungKeun Yoon 0001
INFOCOM3
1997 A modeling of security management system for electronic data interchange
abstract
KT-EDI, an EDI system based on X.435, has been developed jointly by Korea Telecom and ETRI (Electronics and Telecommunications Research Institute) in Korea. We describe the design of the security management system for KT-EDI. We specified the requirements and functions of security management for KT-EDI on the basis of X.800 and other standards. By the above specifications, we designed the security management system for KT-EDI and the prototype is being developed.
Taekyoung Kwon 0002, MyungKeun Yoon 0001, JooSeok Song, Chang-Goo Kang
ISCC2