VLDB 2026 Research / reviewers in the wild / expert
Jinoh Kim
dblp:17/4055
· DBLP profile ↗
46ranked-venue papers
14as first author
20since 2021 · last 2026
0000-0002-9835-1866ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 16 · 3 first-author · 11 since 2021Systems, architecture and hardware · 14 · 7 first-author · 1 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021Security and privacy · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Early Attack Identification in the WildabstractWhile characterizing network connections in their early stage is vital for providing timely responses against network threats, existing methods become less attractive due to the requirement of complete connection information (thus unable to make timely identification) or packet payload inspection (there-fore limited to unencrypted packets under no privacy regulation). To this end, this paper takes an approach ofpacket stream analysisreferencing statistical information of packet sequences, requiringneitherpacket inspectionnorcomplete connection information. To enable practical packet stream analysis, there exist several challenges, such asout-of-order packet sequencesintroduced by network dynamics andclass imbalancewith a tiny fraction of attack connections. To overcome these challenges, we design two deep sequence models: (i) abidirectional recurrent structuredesigned for greater resilience to out-of-order packet streams, and (ii) apre-training-enabled sequence-to-sequence structuredesigned for creating consistent representations from unbalanced class distributions using self-supervised learning. We evaluate the presented deep sequence models using real and synthetic network data collections for extensive experimentation. The experimental results support the feasibility of the proposed models outperforming baseline deep learning models, yielding up to 94.8% (F1 score) only with the first five packets (k=5) from the Internet traffic collection containing a substantial fraction of network flows experiencing out-of-order delivery. Dongeun Lee 0001, Kookjin Lee, Doowon Kim, Jinpyo Kim, Sangman Lee, Jinoh Kim |
IEEE Trans. Netw. | 7 |
| 2025 | Conditional Recurrent Neural Networks for Enhancing Throughput Prediction and Slow File Transfers Detection in Large Science WorkflowsabstractEfficient data transfer across scientific computing facilities is critical for enabling timely scientific discoveries. In this work, we explore the options of anticipating extremely slow data transfers to enable preventive actions. However, the dynamic nature of the large distributed scientific workflows driving these data transfers presents significant challenges for predicting network throughput. This study introduces a Conditional Recurrent Neural Network (CondRNN) model, specifically utilizing Conditional Long Short-Term Memory (CondLSTM), to integrate both static and dynamic features for enhanced throughput prediction. By leveraging historical transfers as proxy features, more than 60% of predictions achieved an absolute percentage error (APE) of less than 20%, and slow transfers were detected with a precision of 91.7% and recall of 100%, outperforming traditional RNN models. Implementing CondLSTM in scientific computing environments can optimize network resource utilization, ensuring efficient data transmission, thereby supporting the continuous progression of scientific research. Boyu Fan, Alex Sim, Kesheng Wu, Jinoh Kim |
CCNC | 4 |
| 2025 | WATCH '25: First Workshop on Analytics, Telemetry, and Cybersecurity for HPCCabstractThe Workshop on Analytics, Telemetry, and Cybersecurity for High-Performance Computing and Communications (HPCC) is newly launched and takes place in Taipei, Taiwan, on October 17, 2025, in conjunction with the ACM Conference on Computer and Communications Security (CCS'25). As its title suggests, the workshop centers on strengthening resilience and security in HPCC applications and infrastructures by leveraging leading technologies, including data-driven methodologies and machine intelligence techniques. The primary objective of this workshop is to provide a dedicated platform for researchers, practitioners, and industry experts to engage in discussions on cutting-edge topics in analytics, telemetry, and cybersecurity for HPCC. This year's call for contributions welcomes both full research papers and work-in-progress submissions, resulting in the acceptance of five full-length research papers. In addition, the workshop features a distinguished keynote presentation by Dr. Hsu-Chun Hsiao, Associate Professor in the Department of Computer Science and Information Engineering and the Graduate Institute of Networking and Multimedia at National Taiwan University. The WATCH'25 complete workshop proceedings can be found at: https://dl.acm.org/citation.cfm?id=3733826. Massimo Cafaro, Eric Chan-Tin, Jerry Chou 0001, Jinoh Kim |
CCS | 4 |
| 2025 | Base Station Certificate and Authentication for 5G Radio Control SecurityabstractCurrent cellular networking remains vulnerable to fake base stations due to the lack of base station authentication mechanism or even a key to enable authentication. We design and build a base station certificate (certifying the base station’s public key and location) and a multi-factor authentication (making use of the certificate and the information transmitted in the online radio control communications) to provide authenticity of the source base station and its radio resource control communications. We advance beyond the state-of-the-art research by introducing greater authentication factors and by using blockchain to deliver the base station digital certificate offline, enabling greater key length/security strength and computational/networking efficiency. The multi-factor authentication at the user equipment involves multiple factors verified through the ledger database, the location sensing, and the cryptographic digital signature verification of the cellular radio control communication (SIB1 broadcasting). We analyze our scheme’s security, performance, and the fit to the existing standardized networking protocols. Our work involves the implementation building on X.509 certificate (adapted), smart contract-based blockchain, 5G-standardized radio resource control communications, and software-defined radios. Our analyses show that our scheme effectively defends against more security threats and can enable stronger security, i.e., ECDSA with greater key lengths. Furthermore, our scheme achieves over threefold improvement in computing and energy efficiency on the mobile user equipment compared to previous research. Sourav Purification, Simeon Wuthier, Jinoh Kim, Ikkyun Kim, Sang-Yoon Chang |
MASS | 3 |
| 2024 | Fake Base Station Detection and BlacklistingabstractA fake base station is a well-known security issue in mobile networking. The fake base station exploits the vulnerability in the broadcasting message announcing the base station’s presence, which is called SIB1 in telecommunications protocols such as 4G LTE and 5G NR, to get the user equipment to connect to itself. Once connected, the fake base station can deprive the user of connectivity and access to the Internet/cloud. We discover that a fake base station (which engages the user equipment until parts of the connectivity setup and then discontinues with the protocol) can disable the victim user equipment’s connectivity for an indefinitely long time, which we validate using our threat prototype against current 4G/5G practice. We design and build a detection and blacklisting identification of the fake base station so that the user equipment can avoid the base station and move on to connecting to a legitimate base station for the connectivity availability. Our detection and blacklisting scheme builds on the standardized 5G protocol and requires the implementation only on the user equipment (no further protocol changes), facilitating practicality. Our scheme uses the real-time information of both the time duration and the number of request transmissions, which features are directly impacted by the fake base station’s threat and have not been studied in the previous research. We implement both the base station and the user defense on software-defined radio using open-source 5G software (srsRAN and Open5GS) for validations. We vary the base station implementation to simulate legitimate vs. faulty-but-legitimate vs. fake-and-malicious base stations, where the faulty base station notifies the connectivity disruption and releases the session while the fake base station continues to hold the session. We empirically analyze the detection and identification thresholds, which vary with the fake base station’s power and the channel condition. By strategically selecting the threshold parameters, our scheme provides zero errors, including zero false positives to avoid blacklisting the temporarily faulty base stations which can not provide the connectivity at the time. Sourav Purification, Simeon Wuthier, Jinoh Kim, Jonghyun Kim 0005, Sang-Yoon Chang |
ICCCN | 3 |
| 2024 | Intelligent Trajectory-based Approach to UAV Location Integrity ChecksabstractWhile unmanned aerial vehicles (UAVs) are increasingly utilized in many domains, there is a growing concern about location integrity for securely deploying and managing the vehicles. A body of studies tackled this problem, e.g., using hardware sensors, cryptographic mechanisms, and machine learning (ML) approaches, but they concentrate primarily on GPS-related information (e.g., jamming and noise). In this study, we take a different approach that performs the checks by analyzing actual movement information. This new approach keeps track of location updates across the flight path (‘trajectory’) rather than relying only on point-wise GPS-specific features to test the validity of the location information. To this end, we develop a deep sequence method that takes a sequence of flight data samples with a minimal set of attributes capturing location movement over time. Our extensive experimental results support the feasibility of our approach, showing up to 98.9% accuracy for ensuring location consistency (even without referring to any of the GPS-specific features). Sang-Yoon Chang, Jonghyun Kim 0005, Kyungmin Park, Jinoh Kim |
ICCCN | 5 |
| 2024 | Base station gateway to secure user channel access at the first hop edge
Sang-Yoon Chang, Arijet Sarker, Simeon Wuthier, Jinoh Kim, Jonghyun Kim 0005, Xiaobo Zhou 0002 |
Comput. Networks | 4 |
| 2023 | Version++: Cryptocurrency Blockchain Handshaking With Software AssuranceabstractCryptocurrency software implements the cryptocurrency operations, including the distributed consensus protocol and the peer-to-peer networking. We design a software assurance scheme for cryptocurrency and advance the cryptocurrency handshaking protocol. Since we focus on Bitcoin (the most popular cryptocurrency) for implementation and integration, we call our scheme Version++, built on and advancing the current Bitcoin handshaking protocol based on the Version message. Our Version++ protocol providing software assurance is distinguishable from the previous research because it is permissionless, distributed, and lightweight to fit its cryptocurrency application. Our scheme is permissionless since it does not require a centralized trusted authority (unlike the remote software attestation techniques from trusted computing); it is distributed since the peer checks the software assurances of its own peer connections; and it is designed for efficiency/lightweight due to the dynamic nature of the peer connections and the large-scale broadcasting in cryptocurrency networking. Utilizing Merkle Tree for the efficiency of the proof verification, we implement and test Version++ on Bitcoin software and conduct experiments in an active Bitcoin node prototype connected to the Bitcoin Mainnet. Our prototype-based performance analyses demonstrate the lightweight design of Version++. The peer-specific verification grows logarithmically with the number of software files in processing time and in storage. In addition, the Version++ verification overhead is small compared to the overall handshaking process; our measured overhead of 2.22% with minimal networking latency between the virtual machines provides an upper bound in the real-world networking with greater handshaking duration, i.e., the relative Version++ overhead in the real world with physically separate machines will be smaller. Arijet Sarker, Simeon Wuthier, Jinoh Kim, Jonghyun Kim 0005, Sang-Yoon Chang |
CCNC | 3 |
| 2023 | Version++ Protocol Demonstration for Cryptocurrency Blockchain Handshaking with Software AssuranceabstractCryptocurrency software implements the cryptocurrency operations. We design a software assurance scheme for cryptocurrency and advance the cryptocurrency handshaking protocol. More specifically, we focus on Bitcoin for implementation and integration and advance its Version-message based hand-shaking and thus call our scheme Version++, The Version++ protocol provides software assurance, which is distinguishable from the previous research because it is permissionless, distributed, and lightweight to fit its cryptocurrency application. Utilizing Merkle Tree for the verification efficiency, we implement and test Version++ on Bitcoin software and conduct experiments in an active Bitcoin node prototype connected to the Bitcoin Mainnet. This paper for the conference demonstration supplements our technical paper at CCNC 2023 for synergy but highlights the prototyping and demonstration components of our research. Arijet Sarker, Simeon Wuthier, Jinoh Kim, Jonghyun Kim 0005, Sang-Yoon Chang |
CCNC | 3 |
| 2023 | An Empirical Evaluation of Autoencoding-Based Location Spoofing DetectionabstractLocation integrity is highly crucial in mobile communications. In this regard, the attack attempting to falsify the position of mobile agents (known as location spoofing) is critical, and thus, detecting such spoofing attacks should be a vital function in the mobile communication setting. With its importance, previous studies explored location spoofing attacks. Still, they mainly used classification techniques based on supervised learning, confining the detector's capability to detect known attack patterns. This study evaluates the feasibility of the autoencoder-based scheme that constructs a profile for legitimate data instances to be resilient to intelligent, previously unseen types of attacks (e.g., evading attacks). We examine three types of autoencoder models designed based on different structures and conduct extensive experiments to measure the performance of the autoencoder models with both standard and variation attacks, with a comparison study with conventional supervised learning-based classification techniques. Our experimental results show that the autoencoder models produce comparable or even better performance than supervised learners, which may be limited only to detecting known patterns. Chiho Kim, Sang-Yoon Chang, Jonghyun Kim 0005, Jinoh Kim |
ICMLA | 4 |
| 2023 | Lightweight and Identifier-Oblivious Engine for Cryptocurrency Networking Anomaly DetectionabstractThe distributed cryptocurrency networking is critical because the information delivered through it drives the mining consensus protocol and the rest of the operations. However, the cryptocurrency peer-to-peer (P2P) network remains vulnerable, and the existing security approaches are either ineffective or inefficient because of the permissionless requirement and the broadcasting overhead. We design and build a Lightweight and Identifier-Oblivious eNgine (LION) for the anomaly detection of the cryptocurrency networking. LION is not only effective in permissionless networking but is also lightweight and practical for the computation-intensive miners. We build LION for anomaly detection and use traffic analyses so that it minimally affects the mining rate and is substantially superior in its computational efficiency than the previous approaches based on machine learning. We implement a LION prototype on an active Bitcoin node to show that LION yields less than 1% of mining rate reduction subject to our prototype, in contrast to the state-of-the-art machine-learning approaches costing 12% or more depending on the algorithms subject to our prototype as well, while having detection accuracy of greater than 97% F1-score against the attack prototypes and real-world anomalies. LION therefore can be deployed on the existing miners without the need to introduce new entities in the cryptocurrency ecosystem. Wenjun Fan, Hsiang-Jen Hong, Jinoh Kim, Simeon Wuthier, Makiya Nakashima, Xiaobo Zhou 0002, C. Edward Chow, Sang-Yoon Chang |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2023 | Automated, Reliable Zero-Day Malware Detection Based on Autoencoding ArchitectureabstractWhile a body of studies has been carried out for malware detection with its significance, they are often limited to known malware patterns due to the reliance on signature-based or supervised learning approaches. The semi-supervised learning approach would be an option for identifying previously unseen patterns (i.e., zero-day detection); however, our preliminary study reveals critical limitations from existing methods, including (i) the profiling-based approach using an autoencoder can provide better detection but is sensitive to the threshold setting, and (ii) one-class (OC) classification does not require a manual threshold discovery but may be limited with low detection rates. In this paper, we present a new detection method incorporating the concept of autoencoding and OC classification, designed to benefit from strong abstraction by neural networks (using an autoencoder) and the removal of the complex threshold selection (using an OC classifier). For this combined architecture, a challenge is concurrent training of the autoencoder and the OC classifier, which may cause an ill-suited learner due to no reference to malware instances. To this end, we introduce a new model selection method that discovers well-optimized models from a variety of combinations. The experimental results performed with public malware datasets (Meraz’18 and Drebin) show the effectiveness of our presented methods with up to 97.1% accuracy, comparable to the supervised learning-based detection. We also examine the impact of evading attacks using adversarial attack tools, the result of which shows resilience to malware variants with over 99% detection rates. Chiho Kim, Sang-Yoon Chang, Jonghyun Kim 0005, Dongeun Lee 0001, Jinoh Kim |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2022 | SNTA'22: The 5th Workshop on Systems and Network Telemetry and AnalyticsabstractHPC and distributed systems are the driving force for the advancement of many emerging technologies, such as exascale systems, quantum machines, terabit networking, 5G/6G wireless, and cloud/edge computing. The tasks of systems and network telemetry are a key element for effective operations and management of the advancement of many emerging systems and technologies, and require more scalable telemetry and analysis techniques for comprehensive monitoring and analysis. Various input sources such as end systems, switches, firewalls, intrusion sensors and the emerging network elements speaking with different syntax and semantics make organizing and incorporating the generated data challenging for the quantitative and qualitative analysis. This workshop looks for new approaches and methods at the intersection of HPC systems and data sciences to address these difficult challenges of emerging technologies from the diverse angles of systems/network performance, availability, reliability, and security. Jinoh Kim, Massimo Cafaro, Jerry Chou 0001, Alex Sim |
HPDC | 1 |
| 2022 | Deep Sequence Models for Packet Stream Analysis and Early DecisionsabstractThe packet stream analysis is essential for the early identification of attack connections while in progress, enabling timely responses to protect system resources. However, there are several challenges for implementing effective analysis, including out-of-order packet sequences introduced due to network dynamics and class imbalance with a small fraction of attack connections available to characterize. To overcome these challenges, we present two deep sequence models: (i) a bidirectional recurrent structure designed for resilience to out-of-order packets, and (ii) a pre-training-enabled sequence-to-sequence structure designed for better dealing with unbalanced class distributions using self-supervised learning. We evaluate the presented models using a real network dataset created from month-long real traffic traces collected from backbone links with the associated intrusion log. The experimental results support the feasibility of the presented models with up to 94.8% in F1 score with the first five packets (k=5), outperforming baseline deep learning models. Dongeun Lee 0001, Kookjin Lee, Doowon Kim, Sangman Lee, Jinoh Kim |
LCN | 6 |
| 2022 | Lightweight Code Assurance Proof for Wireless SoftwareabstractSoftware-defined radio (SDR) and the softwarization of the wireless and mobile systems enable intelligent processing and control in wireless networking. We design and build a lightweight code assurance proof scheme for wireless system software implementations. More specifically, our scheme assures that a wireless user/prover holds the correct software codes, e.g., the correct version, for its wireless networking implementations. In contrast to the previous research for code attestation in trusted computing, our scheme forgoes hardware-based security and real-time networking, thus substantially increasing the application feasibility. We further design our scheme to be efficient in computing by using a Merkle tree for the efficiency of the verification of the assurance proof. We implement our scheme for proof-of-concept on srsRAN (a popular open-source software for cellular technology) and conduct preliminary measurements to demonstrate the lightweight design. We envision our scheme to be orthogonal and supplementary to the previous trustworthy code attestation because it provides different properties (assurance vs. attestation) and because the lightweight aspect yields greater applicability and lower overheads in hardware and networking. Our scheme will therefore be appropriate for the wireless/mobile environment which uses broadcasting (where receiving/verifications occur more frequently than transmitting/generations) and whose devices are resource-constrained. Theo Gamboni-Diehl, Simeon Wuthier, Jinoh Kim, Jonghyun Kim 0005, Sang-Yoon Chang |
WISEC | 3 |
| 2022 | Robust P2P networking connectivity estimation engine for permissionless Bitcoin cryptocurrency
Hsiang-Jen Hong, Wenjun Fan, Simeon Wuthier, Jinoh Kim, C. Edward Chow, Xiaobo Zhou 0002, Sang-Yoon Chang |
Comput. Networks | 4 |
| 2022 | A Machine Learning Approach to Anomaly Detection Based on Traffic Monitoring for Secure Blockchain NetworkingabstractWhile blockchain technology provides strong cryptographic protection on the ledger and the system operations, the underlying blockchain networking remains vulnerable due to potential threats such as denial of service (DoS), Eclipse, spoofing, and Sybil attacks. Effectively detecting such malicious events should thus be an essential task for securing blockchain networks and services. Due to its importance, several studies investigated anomaly detection in Bitcoin and blockchain networks, but their analyses mainly focused on the blockchain ledger in the application context (e.g., transactions) and targets specific types of attacks (e.g., double-spending, deanonymization, etc). In this study, we present a security mechanism based on the analysis of blockchain network traffic statistics (rather than ledger data) to detect malicious events, through the functions of data collection and anomaly detection. The data collection engine senses the underlying blockchain traffic and generates multi-dimensional data streams in a periodic, real-time manner. The anomaly detection engine then detects anomalies from the created data instances based on semi-supervised learning, which is capable of detecting previously unseen patterns, and we introduce our profiling-based detection engine implemented on top of AutoEncoder (AE). Our experimental results evaluated with real and simulated traffic data support the effectiveness of our security mechanism and design choices based on the AE structure, with the approximate detection performance to the supervised learning methods only through the profiling of normal instances. The measured time complexity is sufficiently cheap to perform real-time analysis, with less than 1.4 msec for per-instance testing on a single core setting. Jinoh Kim, Makiya Nakashima, Wenjun Fan, Simeon Wuthier, Xiaobo Zhou 0002, Ikkyun Kim, Sang-Yoon Chang |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2021 | Zero-day Malware Detection using Threshold-free Autoencoding ArchitectureabstractThe impact of malware attacks has been getting more significant, targeting critical infrastructures as well as commodity computing devices. A body of studies has been carried out for detecting malware with its devastating impacts, but they are often limited to known malware attacks due to the nature of the signature-based and supervised machine learning approaches. The semi-supervised learning approach would be an option for identifying previously unseen types of malware attacks (i.e., zero-day detection); however, our preliminary studies suggest two limitations in this avenue: (1) one class (OC) classifiers can be limited with relatively low detection rates, and (2) the profiling-based approach (using an autoencoder) may yield better detection performance but under the assumption of the "ideal" threshold setting. In this paper, we tackle these challenges and present a new detection method, which combines the concepts of autoencoding and OC classification, to benefit from strong abstractions by neural networks (using an autoencoder) but to remove the necessity of the complex threshold selection (using an OC classifier). Our extensive experimental results with a recent malware dataset (Meras’18) show the effectiveness of our method with up to 96% accuracy for zero-day malware detection, which is comparable to the supervised learning-based detection (limited to known types of malware). The proposed method also shows the resilience to adversarial attacks, yielding better performance for identifying synthetic samples generated to evade the detection process than supervised learning algorithms. Chiho Kim, Sang-Yoon Chang, Jonghyun Kim 0005, Dongeun Lee 0001, Jinoh Kim |
IEEE BigData | 5 |
| 2021 | Robust P2P Connectivity Estimation for Permissionless Bitcoin NetworkabstractBlockchain relies on the underlying peer-to-peer (p2p) networking to broadcast and get up-to-date on the blocks and transactions. It is therefore imperative to have high p2p connectivity for the quality of the blockchain system operations. High p2p networking connectivity ensures that a peer node is connected to multiple other peers providing a diverse set of observers of the current state of the blockchain and transactions. However, in a permissionless blockchain network, using the peer identifiers—including the current approach of counting the number of distinct IP addresses and port numbers—can be ineffective in measuring the number of peer connections and estimating the networking connectivity. Such current approach is further challenged by the networking threats manipulating the identifiers. We build a robust estimation engine for the p2p networking connectivity by sensing and processing the p2p networking traffic. We implement a working Bitcoin prototype connected to the Bitcoin Mainnet to validate and improve our engine’s performances and evaluate the estimation accuracy and cost efficiency of our estimation engine. Hsiang-Jen Hong, Wenjun Fan, Simeon Wuthier, Jinoh Kim, Xiaobo Zhou 0002, C. Edward Chow, Sang-Yoon Chang |
IWQoS | 4 |
| 2021 | A Machine Learning Approach to Peer Connectivity Estimation for Reliable Blockchain NetworkingabstractPeer connectivity plays a significant role in a blockchain network since any poor connectivity may result in the nodes operating on outdated data (e.g., cryptocurrency transactions). Although connectivity information is maintained by individual nodes, such identifier-based information might be unreliable due to the possibility of bogus identifiers. This paper tackles the problem of peer connectivity estimation through data-driven analytics of blockchain traffic for reliable blockchain networking. We define a set of variables to represent traffic characteristics and estimate peer connectivity from the collected data using a machine learning methodology. We also investigate the feasibility of feature prioritization to minimize estimation complexities. Our experimental results show that the presented estimation mechanism makes accurate predictions, with less than 0.1 difference between the measurement and estimation for over 99.7% of predictions. The time complexity measured on a commodity machine shows a microsecond scale for completing a single prediction task, enabling real-time operations. Jinoh Kim, Makiya Nakashima, Wenjun Fan, Simeon Wuthier, Xiaobo Zhou 0002, Ikkyun Kim, Sang-Yoon Chang |
LCN | 1 |
| 2020 | Botnet Detection Using Recurrent Variational AutoencoderabstractBotnet detection is an active research topic as botnets are a source of many malicious activities, including distributed denial-of-service (DDoS), click-fraud, spamming, and crypto-mining attacks. However, it is getting more complicated to identify botnets due to the continuous evolution of botnet software and families that harness new types of devices and attack vectors. Recent studies employing machine learning (ML) showed improved performance to detect botnets to some extent, but they are still limited and ineffective with the lack of sequential pattern analysis, which is a key to detect various classes of botnets. In this paper, we propose a novel botnet detection method, built upon Recurrent Variational Autoencoder (RVAE), that effectively captures sequential characteristics of botnet anomalies. We validate the feasibility of the proposed method with the CTU-13 dataset that have been widely employed for botnet detection studies, and show that our method is at least comparable to existing techniques in terms of detection accuracy. In addition, our experimental results show that the proposed method can detect previously unseen botnets by utilizing sequential patterns of network traffic. We will also show how our method can detect botnets in the streaming mode, which is the essential requirement to perform real-time, on-line detection. Jeeyung Kim, Alex Sim, Jinoh Kim, Kesheng Wu |
GLOBECOM | 3 |
| 2020 | A Learning-based Data Augmentation for Network Anomaly DetectionabstractWhile machine learning technologies have been remarkably advanced over the past several years, one of the fundamental requirements for the success of learning-based approaches would be the availability of high-quality data that thoroughly represent individual classes in a problem space. Unfortunately, it is not uncommon to observe a significant degree of class imbalance with only a few instances for minority classes in many datasets, including network traffic traces highly skewed toward a large number of normal connections while very small in quantity for attack instances. A well-known approach to addressing the class imbalance problem is data augmentation that generates synthetic instances belonging to minority classes. However, traditional statistical techniques may be limited since the extended data through statistical sampling should have the same density as original data instances with a minor degree of variation. This paper takes a learning-based approach to data augmentation to enable effective network anomaly detection. One of the critical challenges for the learning-based approach is the mode collapse problem resulting in a limited diversity of samples, which was also observed from our preliminary experimental result. To this end, we present a novel "Divide-Augment-Combine" (DAC) strategy, which groups the instances based on their characteristics and augments data on a group basis to represent a subset independently using a generative adversarial model. Our experimental results conducted with two recently collected public network datasets (UNSW-NB15 and IDS-2017) show that the proposed technique enhances performances up to 21.5% for identifying network anomalies. Mohammad Al Olaimat, Dongeun Lee 0001, Youngsoo Kim 0002, Jonghyun Kim 0005, Jinoh Kim |
ICCCN | 5 |
| 2020 | Enhancing IoT Anomaly Detection Performance for Federated LearningabstractWhile federated learning (FL) has gained great attention for mobile and Internet of Things (IoT) computing with the benefits of scalable cooperative learning and privacy protection capabilities, there still exist a great deal of technical challenges to make it practically deployable. For instance, the distribution of the training process to a myriad of devices limits the classification performance of machine learning (ML) algorithms, often showing a significantly degraded accuracy compared to centralized learning. In this paper, we investigate the problem of performance limitation under FL and present the benefit of data augmentation with an application of anomaly detection using an IoT dataset. Our initial study reveals that one of the critical reasons for the performance degradation is that each device sees only a small fraction of data (that it generates), which limits the efficacy of the local ML model (constructed by the device). This becomes more critical if the data holds the class imbalance problem, observed not infrequently in practice (e.g., a small fraction of anomalies). Moreover, device heterogeneity with respect to data quantity is an open challenge in FL. Based on these observations, we examine the impact of data augmentation on detection performance in FL settings (both homogeneous and heterogeneous). Our experimental results show that even a simple random oversampling can improve detection performance with manageable learning complexity. Brett Weinger, Jinoh Kim, Alex Sim, Makiya Nakashima, Nour Moustafa, Kesheng Wu |
MSN | 2 |
| 2020 | Learning-based dynamic cache management in a cloud
Jinhwan Choi, Yu Gu 0001, Jinoh Kim |
J. Parallel Distributed Comput. | 3 |
| 2019 | Federated Wireless Network Intrusion DetectionabstractWi-Fi has become the wireless networking standard that allows short- to medium-range device to connect without wires. For the last 20 year, the Wi-Fi technology has so pervasive that most devices in use today are mobile and connect to the internet through Wi-Fi. Unlike wired network, a wireless network lacks a clear boundary, which leads to significant Wi-Fi network security concerns, especially because the current security measures are prone to several types of intrusion. To address this problem, machine learning and deep learning methods have been successfully developed to identify network attacks. However, collecting data to develop models is expensive and raises privacy concerns. The goal of this paper is to evaluate a federated learning approach that would alleviate such privacy concerns. This initial work on intrusion detection is performed in a simulated environment. Once proven feasible, this process would allow edge devices to collaboratively update global anomaly detection models, without sharing sensitive training data. On a set of tests with the AWID intrusion detection data set, we show that our federated approach is effective in terms of classification accuracy, computation cost, as well as communication cost. Burak Cetin, Alina Lazar, Jinoh Kim, Alex Sim, Kesheng Wu |
IEEE BigData | 3 |
| 2019 | A New Approach to Multivariate Network Traffic Analysis
Jinoh Kim, Alex Sim |
J. Comput. Sci. Technol. | 1 |
| 2018 | An Encoding Technique for CNN-based Network Anomaly DetectionabstractAn important challenge in the cyber-space is the effective identification of network anomalies, often caused by malicious activities. With the remarkable advances, machine learning algorithms have widely been studied for network intrusion and anomaly detection. In particular, deep learning based on neural network structures has recently been given a greater attention to deal with the growing complexity of data with higher dimensions and non-linearity. Convolutional Neural Networks (CNNs) is one of the widely employed deep learning methods. In this work, we introduce a new encoding technique that enhances the performance for the identification of anomalous events using a CNN structure. To evaluate, we utilize three different datasets for the extensive analysis. The experimental results show that our method consistently outperforms the gray-scale encoding technique previously proposed over the datasets employed in the evaluation. Taejoon Kim, Sang C. Suh, Hyunjoo Kim, Jonghyun Kim 0005, Jinoh Kim |
IEEE BigData | 5 |
| 2018 | Website Fingerprinting Attack Mitigation Using Traffic MorphingabstractWebsite fingerprinting attacks attempt to identify the website visited in anonymized and encrypted network traffic, that is, even if a user is using Tor and HTTPS. These attacks have been shown to be effective. Mitigations have been proposed which decreased the accuracy of the attacks from about 90% to about 20%. We propose a new mitigation technique based on traffic morphing and clustering. The intuition is that a lot of websites, by nature, are similar and can be clustered together. It is then easier and more efficient to make that whole cluster look exactly the same by using traffic morphing, rather than adding noise to make all websites look similar. All the websites in a cluster, thus, would become indistinguishable. There are many ways to perform traffic morphing. As a proof of concept, we used biggest, which means that all websites in a cluster will look exactly like the biggest website (in terms of network packet size) of that cluster. In simulating our proposed approach, the fingerprinting accuracy dropped from 70% to less than 1%. Eric Chan-Tin, Taejoon Kim, Jinoh Kim |
ICDCS | 3 |
| 2018 | An Empirical Study on Network Anomaly Detection Using Convolutional Neural NetworksabstractDeep learning has been widely applied to network anomaly detection to improve performance. In our past research, we empirically evaluated a set of deep learning models, including Fully Connected Network (FCN), Variational Auto Encoder (VAE), and Sequence to Sequence model with Long Short-Term Memory (Seq2Seq-LSTM), for network anomaly detection. Additionally, we evaluate Convolution Neural Networks (CNNs) for network anomaly detection in this paper. We set up three simple CNN models with different internal depths (shallow CNN, moderate CNN, and deep CNN) to see the impact of the depth to the performance. We evaluate the models using three different types of traffic datasets. Our experimental results show that deeper structures do not make any performance improvement. In addition, we observed that the evaluated CNN models occasionally outperform the VAE models, but do not work better than the other deep learning models based on FCN and Seq2Seq-LSTM. Donghwoon Kwon, Kathiravan Natarajan, Sang C. Suh, Hyunjoo Kim, Jinoh Kim |
ICDCS | 5 |
| 2017 | Unsupervised Labeling for Supervised Anomaly Detection in Enterprise and Cloud NetworksabstractIdentifying anomalous events in the network is one of the vital functions in enterprises, ISPs, and datacenters to protect the internal resources. With its importance, there has been a substantial body of work for network anomaly detection using supervised and unsupervised machine learning techniques with their own strengths and weaknesses. In this work, we take advantage of the both worlds of unsupervised and supervised learning methods. The basic process model we present in this paper includes (i) clustering the training data set to create referential labels, (ii) building a supervised learning model with the automatically produced labels, and (iii) testing individual data points in question using the established learning model. By doing so, it is possible to construct a supervised learning model without the provision of the associated labels, which are often not available in practice. To attain this process, we set up a new property defining anomalies in the context of clustering, based on our observations from anomalous events in network, by which the referential labels can be obtained. Through our extensive experiments with a public data set (NSL-KDD), we will show that the presented method perform very well, yielding fairly comparable performance to the traditional method running with the original labels provided in the data set, with respect to the accuracy for anomaly detection. Sunhee Baek, Donghwoon Kwon, Jinoh Kim, Sang C. Suh, Hyunjoo Kim, Ikkyun Kim |
CSCloud | 3 |
| 2017 | A New Approach to Online, Multivariate Network Traffic AnalysisabstractNetwork traffic analysis has long been a core element for effective network operations and management. While online monitoring has been studied for a while, it is still intensively challenging due to several reasons. One of the primary challenges is the heavy volume of traffic to analyze within a finite amount of time. Another important challenge to enable online monitoring is to support multivariate analysis of traffic variables to help administrators identify unexpected network events intuitively. To this end, we propose a new approach that offers a high- level summary of the network traffic with the multivariate analysis. With this approach, the current state of the network will display an abstract pattern compiled from a set of traffic variables, and the detection problems in traffic analysis (e.g., change detection and anomaly detection) can be reduced to a straightforward pattern identification problem. In this paper, we introduce our preliminary work with clustered patterns for online, multivariate traffic analysis with the challenges and limitations. We then present a grid-based model that is designed to overcome the limitations of the clustered pattern- based technique. We will discuss the potential of the new model with respect to streaming-based computation and robustness to outliers. Jinoh Kim, Alex Sim |
ICCCN | 1 |
| 2015 | Security for the scientific data services frameworkabstractScientific data is often shared among researchers and even reorganized by colleagues or third-party users. Thus, it is essential to provide an adequate degree of access control for such shared data to preserve a desired level of security requirements. In this work, we develop an essential, lightweight access control model for secure data services in a limited distributed setting such as an HPC cluster. In particular, we consider SDS (the Scientific Data Services framework) as a use case system, which is a framework offering performance-optimized data access, reorganization, and analysis. We outline the requirements and challenges for access control for effective data services, and develop an authorization service model based on the defined requirements. We also present an initial prototyping model. Jinoh Kim, Bin Dong 0002, Surendra Byna, Kesheng Wu |
IEEE BigData | 1 |
| 2015 | Incorporating multiple cluster models for network traffic classificationabstractNetwork traffic classification is one of the essential functions for local and ISP networks. With its importance, a substantial number of previous studies have explored various machine learning techniques with network flow statistics for accurate traffic classification, including the clustering-based approach. However, we obtained unacceptable results from previously proposed clustering-based techniques from our preliminary experiments. In particular, simply employing the entire flow attributes for clustering leads to unexpectedly poor accuracy (less than 70%). In this paper, we propose a new technique based on multiple trained cluster models to overcome this problem. The proposed technique utilizes multiple sets of attribute combinations in parallel for traffic classification, rather than simply merging the entire (or a subset of) attributes in a single model. Our technique also includes a selection step to reduce the results from the individual models into a single output as the final classification decision, and we explore a set of selection strategies. We present our experimental results and show that our technique significantly improves overall accuracy up to 95%. Jinoh Kim, Sang C. Suh, Ganho Choi |
LCN | 2 |
| 2015 | Exploiting Replication for Energy-Aware Scheduling in Disk Storage SystemsabstractThis paper deals with the problem of scheduling requests on disks for minimizing energy consumption. We first analyze several versions of the energy-aware disk scheduling problem based on assumptions on the arrival pattern of the requests. We show that the corresponding optimization problems are NP-complete. Then both optimal and heuristic scheduling algorithms are proposed to maximize the energy saving of a storage system. We evaluate our approach using multiple realistic I/O traces, disk simulator and energy model. The results show that we significantly reduce energy consumption up to 55 percent and achieve fewer disk spin-up/ down operations and shorter request response time as compared to other approaches. Since our approach attempts to dynamically assign each request to an energy optimized location, it can also benefit from other traditional static or semi-static solutions that rely on data placement or migration. Finally, we show that a write offloading technique can also be adapted into our solution to minimize the impact from write re quests in terms of energy consumption and request response time. Jerry Chou 0001, Ting-Hsuan Lai, Jinoh Kim, Doron Rotem |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2014 | A Security-enabled Grid System for MINDS Distributed Data Mining
Jinoh Kim, Jon B. Weissman |
J. Grid Comput. | 2 |
| 2014 | iPACS: Power-aware covering sets for energy proportionality and performance in data parallel computing clusters
Jinoh Kim, Jerry Chou 0001, Doron Rotem |
J. Parallel Distributed Comput. | 1 |
| 2012 | FREP: Energy proportionality for disk storage using replication
Jinoh Kim, Doron Rotem |
J. Parallel Distributed Comput. | 1 |
| 2011 | Energy proportionality for disk storage using replicationabstractSaving energy for storage is of major importance as storage devices (and cooling them off) may contribute over 25 percent of the total energy consumed in a datacenter. Recent work introduced the concept of energy proportionality and argued that it is a more relevant metric than just energy saving as it takes into account the tradeoff between energy consumption and performance. In this paper, we present a novel approach, called FREP (Fractional Replication for Energy Proportionality), for energy management in large datacenters. FREP includes a replication strategy and basic functions to enable flexible energy management. Specifically, our method provides performance guarantees by adaptively controlling the power states of a group of disks based on observed and predicted workloads. Our experiments, using a set of real and synthetic traces, show that FREP dramatically reduces energy requirements with a minimal response time penalty. Jinoh Kim, Doron Rotem |
EDBT | 1 |
| 2011 | Energy-Aware Scheduling in Disk Storage SystemsabstractThis paper deals with the problem of scheduling requests on disks for minimizing energy consumption. We first analyze several versions of the energy-aware disk scheduling problem based on assumptions on the arrival pattern of the requests. We show that the corresponding optimization problems are NP-complete by reduction to the set cover or the independent set problem. Then both optimal and heuristic scheduling algorithms are proposed to maximize the energy saving of a storage system. Our evaluation results using two realistic traces show that our approach significantly reduces energy consumption up to 55% and achieves fewer disk spin-up/down operations and shorter request response time as compared to other approaches. Jerry Chou 0001, Jinoh Kim, Doron Rotem |
ICDCS | 2 |
| 2011 | Energy Proportionality and Performance in Data Parallel Computing Clusters
Jinoh Kim, Jerry Chou 0001, Doron Rotem |
SSDBM | 1 |
| 2011 | Passive Network Performance Estimation for Large-Scale, Data-Intensive ComputingabstractDistributed computing applications are increasingly utilizing distributed data sources. However, the unpredictable cost of data access in large-scale computing infrastructures can lead to severe performance bottlenecks. Providing predictability in data access is, thus, essential to accommodate the large set of newly emerging large-scale, data-intensive computing applications. In this regard, accurate estimation of network performance is crucial to meeting the performance goals of such applications. Passive estimation based on past measurements is attractive for its relatively small overhead compared to relying on explicit probing. In this paper, we take a passive approach for network performance estimation. Our approach is different from existing passive techniques that rely either on past direct measurements of pairs of nodes or on topological similarities. Instead, we exploit secondhand measurements collected by other nodes without any topological restrictions. In this paper, we present Overlay Passive Estimation of Network performance (OPEN), a scalable framework providing end-to-end network performance estimation based on secondhand measurements, and discuss how OPEN achieves cost-effective estimation in a large-scale infrastructure. Our extensive experimental results show that OPEN estimation can be applicable for replica and resource selections commonly used in distributed computing. Jinoh Kim, Abhishek Chandra, Jon B. Weissman |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2009 | Using Data Accessibility for Resource Selection in Large-Scale Distributed SystemsabstractLarge-scale distributed systems provide an attractive scalable infrastructure for network applications. However, the loosely coupled nature of this environment can make data access unpredictable, and in the limit, unavailable. We introduce the notion of accessibility to capture both availability and performance. An increasing number of data-intensive applications require not only considerations of node computation power but also accessibility for adequate job allocations. For instance, selecting a node with intolerably slow connections can offset any benefit to running on a fast node. In this paper, we present accessibility-aware resource selection techniques by which it is possible to choose nodes that will have efficient data access to remote data sources. We show that the local data access observations collected from a node's neighbors are sufficient to characterize accessibility for that node. By conducting trace-based, synthetic experiments on PlanetLab, we show that the resource selection heuristics guided by this principle significantly outperform conventional techniques such as latency-based or random allocations. The suggested techniques are also shown to be stable even under churn despite the loss of prior observations. Jinoh Kim, Abhishek Chandra, Jon B. Weissman |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2008 | Accessibility-Based Resource Selection in Loosely-Coupled Distributed SystemsabstractLarge-scale distributed systems provide an attractive scalable infrastructure for network applications. However,the loosely-coupled nature of this environment can make data access unpredictable, and in the limit, unavailable. We introduce the notion of accessibility to capture both availability and performance. An increasing number of data intensive applications require not only considerations of node computation power but also accessibility for adequate job allocations. For instance, selecting a node with intolerably slow connections can offset any benefit to running on a fast node. In this paper, we present accessibility-aware resource selection techniques by which it is possible to choose nodes that will have efficient data access to remote data sources. We show that the local data access observations collected from a node's neighbors are sufficient to characterize accessibility for that node. We then present resource selection heuristics guided by this principle, and show that they significantly out perform standard techniques. The suggested techniques are also shown to be stable even under churn despite the loss of prior observations. Jinoh Kim, Abhishek Chandra, Jon B. Weissman |
ICDCS | 1 |
| 2007 | Exploiting Heterogeneity for Collective Data Downloading in Volunteer-based NetworksabstractScientific computing is being increasingly deployed over volunteer-based distributed computing environments consisting of idle resources on donated user machines. A fundamental challenge in these environments is the dissemination of data to the computation nodes, with the successful completion of jobs being driven by the efficiency of collective data download across compute nodes, and not only the individual download times. This paper considers the use of a data network consisting of data distributed across a set of data servers, and focuses on the server selection problem: how do individual nodes select a server for downloading data to minimize the communication makespan - the maximal download time for a data file. Through experiments conducted on a pastry network running on PlanetLab, we demonstrate that nodes in a volunteer-based network are heterogeneous in terms of several metrics, such as bandwidth, load, and capacity, which impact their download behavior. We propose new server selection heuristics that incorporate these metrics, and demonstrate that these heuristics outperform traditional proximity-based server selection, reducing average makespans by at least 30%. We further show that incorporating information about download concurrency avoids overloading servers, and improves performance by about 17-43% over heuristics considering only proximity and bandwidth. Jinoh Kim, Abhishek Chandra, Jon B. Weissman |
CCGRID | 1 |
| 2007 | NGS: Service Adaptation in Open Grid PlatformsabstractLarge-scale donation-based distributed infrastructures need to cope with the inherent unreliability of participant nodes. A widely-used work scheduling technique in such environments is to redundantly schedule the outsourced computations to a number of nodes. We present the design and implementation of RIDGE, a reliability-aware system which uses a node's prior performance and behavior to make more effective scheduling decisions. We have implemented RIDGE on top of the BOINC distributed computing infrastructure and have evaluated its performance on a live PlanetLab testbed. Our experimental results show that RIDGE is able to match or surpass the throughput of the best BOINC configuration by automatically adapting to the characteristics of the underlying environment. In addition, RIDGE is able to provide much lower workunit makespans compared to BOINC. RIDGE is also able to produce significantly lower communication makespans for downloading clients. Collectively, the results suggest that RIDGE has great promise for service-oriented environments with time constraints. Krishnaveni Budati, Jinoh Kim, Abhishek Chandra, Jon B. Weissman |
IPDPS | 2 |
| 2003 | Applying Data Mining Techniques to Analyze Alert Data
Moon Sun Shin, Hosung Moon, Keun Ho Ryu, Kiyoung Kim, Jinoh Kim |
APWeb | 5 |