Kazunori Kamiya

dblp:167/5154 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 4 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Malicious Log Detection Using Machine Learning to Maximize the Partial AUC
abstract
A recent trend in security log analysis is to utilize machine learning methods to detect malware. By using machine learning, we can save on labor and achieve an advanced counter-measure against constantly evolving malware. When evaluating the classification performance of malicious log detection, the true positive rate (TPR) in a low false positive rate (FPR) interval is widely recognized as important since network operators want to detect as much malware as possible while reducing the false positives of benign logs. However, the conventional supervised learning methods cannot directly maximize the TPR in a low FPR interval since they are trained to maximize accuracy. Therefore, this paper proposes a method to maximize the partial area under the receiver operating characteristic curve (pAUC), which is the mean TPR with a specific interval of the FPR. The proposed method uses the conventional supervised method as a baseline, changes the objective function of the baseline supervised learning method to maximize the pAUC, and learns on the basis of the proposed algorithm. The advantage of the proposed method is its high applicability since it can be implemented by using any conventional supervised learning method for binary classification as a baseline and modifying its objective function. We compared the proposed methods, i.e., the pAUC maximization methods on various supervised learning models, with baseline supervised learning methods by using a public dataset (NSL-KDD) and a dataset consisting of proxy logs from a real-world large enterprise network. From the results, the proposed method outperforms the baseline supervised learning method in terms of several performance measures such as the pAUC, AUC, and TPR at a low FPR. The results suggest that the proposed method is beneficial in actual operation since it can detect more mal ware when operating with the same FPR compared to the conventional supervised learning methods.
Taishi Nishiyama, Atsutoshi Kumagai, Akinori Fujino, Kazunori Kamiya
CCNC4
2024 Partial AUC Maximization for Security Log Analysis Robust to Overfitting and Noisy Labels
abstract
To mitigate the damage caused by malware, network log analysis with machine learning for detecting suspicious logs has been attracting attention. In actual security operation, the true positive rate (TPR) in settings with a low false positive rate (FPR) is important since network operators must detect as many suspicious logs as possible while suppressing false positives. This paper focuses on a partial area under the curve (pAUC) maximization method that directly maximizes the TPR in an arbitrary FPR interval. If the FPR interval is set to a very small and narrow range when using the existing pAUC maximization methods in actual operation, the amount of data of benign logs to be trained is relatively limited. Therefore, if the FPR interval you aim to learn contains malicious logs that have been mislabeled as benign, the classification performance of the existing pAUC maximization methods can significantly deteriorate. In addition, when network logs are converted into feature vectors, the number of features tends to be large, which can lead to overfitting due to the relatively limited amount of data compared to the number of features. To alleviate these problems, we propose a method that combines the AUC maximization and pAUC maximization methods in accordance with the mathematical characteristics of the features. We also demonstrate the effectiveness of proposed method through comparative experiments with proxy logs from a real-world large enterprise network.
Taishi Nishiyama, Atsutoshi Kumagai, Akinori Fujino, Kazunori Kamiya
GLOBECOM4
2022 Balanced Score Propagation for Botnet Detection
abstract
Botnets are a group of computers infected by malware, and have become a major threat to the security of the entire Internet, which enable attacks to cause serious damages to the world, such as DDoS and ransomware campaigns. To manage millions of infected computers, botnets have evolved to use layered and distributed infrastructure. To eradicate the botnet threat, we have developed a novel balanced score propagation technique on graphs to detect the entire structure of a botnet with high precision. Conventional score propagation techniques detect a botnet node close to known one based on the score that quantifies the level of the relevance between the two. Their detection coverage is limited to neighboring nodes and thus misses the macro structure. Our new technique, balanced score propagation, can reveal the entire botnet structures by enabling the detection of both neighboring and distant botnet nodes. The technique employs direct score propagation to distant nodes by adding virtual edges to change the score propagation path. We evaluate the detection precision of the proposed method by experiments using real-world network traffic. The result shows that the proposed method can also detect botnet nodes at intermediate distances effectively as well as nearby and distant nodes, and the precision is improved an average of 1.8 times more than existing methods including graph embedding whose primary focus is the detection of the overall botnet structure.
Shosuke Oba, Kazunori Kamiya, Kenji Takahashi
ICC3
2021 Multi-hop Graph Embedding for Botnet Detection
abstract
We have developed a novel multi-hop graph embedding technique for botnet detection. It can detect the entire layered architecture of a botnet in the internet backbone traffic by starting from a small set of the components of the botnet. A botnet is a group of hosts collaborating each other to launch a variety of attacks, such as distributed denial-of-service attacks and phishing campaigns. Over 20 years of their existence, botnets have been evolved to employ layered architectures for robust operation and efficient management. Several existing methods leverage graph analysis to detect malicious communications. However, they cannot detect such botnet components that communicating each other through multiple layers, which are represented at more than one hop distance in graphs. To solve this problem, our technique trains separate graph embedding models with samples at different distances and select appropriate features from multiple models to represent multi-hop adjacency for each node. By applying our proposal to real-world Internet traffic, we have confirmed that it can outperform other methods in terms of detecting collaborating botnet components with higher accuracy even if their command and control communications are cascading through multiple layers.
Kazunori Kamiya, Kenji Takahashi, Akihiro Nakao
GLOBECOM2
2021 SamE: Sampling-based Embedding for Learning Representations of the Internet
abstract
We have developed SamE, a novel sampling-based embedding technique for learning representations of the Internet. SamE can classify Internet hosts in a scalable and cost effective manner without sacrificing the classification performance. Machine learning has been applied to Internet traffic analysis for a variety of purposes, including botnet detection and application identification. For example, as a major threat on the Internet, a botnet is a group of computers that collaborate together to launch cyberattacks. To analyze related hosts such as the collaborating constituents of a botnet, graph embedding techniques seem to be promising. However, when applying existing graph embedding techniques to Internet-scale traffic data, the time and space complexities become prohibitively high for practical use. To make graph embedding applicable to Internet-scale problems, SamE only samples a subset of nodes to learn elemental representations and aggregates learned elemental representations to generate synthetic representations for all nodes. We have applied SamE to real-world Internet-scale traffic data, and the experimental results show that SamE outperforms existing methods by reducing the data samples required for representation learning by 99% while achieving the same level of classification performance in botnet detection and application identification.
Kazunori Kamiya, Kenji Takahashi, Akihiro Nakao
GLOBECOM2
2020 Piper: A Unified Machine Learning Pipeline for Internet-scale Traffic Analysis
abstract
Machine learning has been applied to network traffic analysis for a variety of purposes, including botnet detection. To improve the computational efficiency, several architectures have been proposed to consolidate processes common across multiple applications that use the same traffic data. However, when introducing conventional architectures to real-world traffic analysis at Internet scale, the amount of input traffic data and the variety of output features to represent global access patterns become new challenges. To address the challenges, we have developed Piper, a machine learning pipeline, that consolidates diversified machine learning applications in a highly efficient manner. On top of the consolidated architecture, Piper employs two novel techniques: (1) selective sampling to reduce traffic data efficiently while maintaining prediction performance, and (2) a set of enriched features to extract temporal and spatial characteristics in global traffic. For the evaluation, we have been deploying Piper to detect botnets from internet backbone traffic over nine months. The evaluation has confirmed the effectiveness of Piper in terms of computational performance, prediction performance, and lead time to detect botnets.
Kazunori Kamiya, Kenji Takahashi, Akihiro Nakao
GLOBECOM2
2020 SILU: Strategy Involving Large-scale Unlabeled Logs for Improving Malware Detector
abstract
Machine learning is becoming a key component to automatically detect malware-infected hosts by analyzing network logs in a security operations center (SOC). However, machine learning usually requires a large amount of labeled training data, which is difficult to acquire since labels are manually set by professional security analysts. On the other hand, abundant unanalyzed logs are kept stored in daily operation and stay unlabeled even though they could compensate for the lack of existing labeled training data. This paper proposes SILU, a novel semi-supervised learning method, which fully leverages unlabeled data and enhances detection capability without increasing manually labeled data. SILU learns from combined labeled and unlabeled training data to automatically augment labeled training data and then generates a classifier through the screening process. Unlike most semi-supervised learning methods used in cyber security, which use test data as unlabeled training data, SILU does not require retraining every time test data change since it can use different datasets for unlabeled training and test data. This helps SOC operation for practically suppressing detecting time. In addition, though SILU partially includes a supervised learning method, it does not require a specific supervised learning method. Therefore, SILU can be added on to any type of classifier of a supervised learning method. Moreover, SILU can suppress the deterioration of classification performance for test data through the screening process. We evaluated SILU using two types of real-world logs: proxy logs from a large enterprise and NetFlow from a large ISP. We demonstrated that by evaluating with different types of classifiers, SILU always improves detection capability for supervised learning methods. SILU also outperforms current semi-supervised methods. As a whole, SILU works as an add-on to existing supervised learning methods with little overhead and performs better than conventional supervised learning methods. Our evaluation also shows that using NetFlow from ISP as unlabeled training data works better than using only labeled proxy logs in the same enterprise. These results suggest that SILU can extend detection capability more when different organizations, e.g., SOCs and ISPs, collaborate and share unlabeled data.
Taishi Nishiyama, Atsutoshi Kumagai, Kazunori Kamiya, Kenji Takahashi
ISCC3
2019 Alchemy: Stochastic Feature Regeneration for Malicious Network Traffic Classification
abstract
As signature-based techniques have ever more difficulty detecting increasing and varying malicious activities through network traffic, machine learning has become a promising approach in network security. Many previous studies have aggregated traffic data into groups by hosts or flows for generating features and training detection models. However, two problems degrade detection performance. One is the scarcity of training sets due to the rarity of new types of malicious traffic, and the other is variations in feature values generated from incomplete data due to limited observed traffic. In this paper, we propose a stochastic method called Alchemy that regenerates a set of feature vectors by randomly resampling raw traffic data of each bag into several subsets. Alchemy can increase training sets and represent raw traffic robustly to correct the influence of variations in feature vectors, regardless of types of traffic data and classifiers. We evaluated Alchemy with real-world traffic data of network flows, passive DNS records, and HTTP logs, and demonstrated that it improves detection performance of various classifiers more effectively than the conventional methods in all three types of traffic data.
Atsutoshi Kumagai, Kazunori Kamiya, Kenji Takahashi, Daniel Dalek, Ola Söderström, Kazuya Okada, Yuji Sekiya, Akihiro Nakao
COMPSAC (1)3
2019 Subspace Clustering for Interpretable Botnet Traffic Analysis
abstract
With the growth of the Internet of Things (IoT), massively connected devices are exposed to cyber-attacks and are becoming active as bots. To protect IoT devices efficiently from increasing threats, in addition to honeypot-based approaches, we need to gain more complete understanding of IoT botnets and potential victims including purposes of attacks and events which they are participating in. In this paper, we propose a two-step subspace clustering method to cluster botnets and clarify their types (functionalities). For each target host, the proposed method separates features into several subspaces and generates a sub-label in each subspace to represent partial characteristic (e.g., low-size flows or high TCP-SYN rate). Major combinations of sub-labels presents different partial characteristic, and the proposed method classify bots and interprets the whole behaviors for each bot group. In the evaluation, we reveal the scale and variety of botnets in the wild through two real-world datasets.
Shohei Araki, Kneji Takahashi, Kazunori Kamiya, Masaki Tanikawa
ICC4