Mamoru Mimura

dblp:52/9355 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0003-4323-9911ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 10 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Oversampling method using large language model for malicious JavaScript code detection
abstract
Detecting malicious JavaScript code on websites remains a significant issue. Detection methods using machine learning models have been proposed to detect these attacks in real time. In targeted attacks, malware is often customized for a specific organization. To defend against these attacks, machine learning models have to be trained from a small number of past malicious samples. Traditional oversampling techniques applied to JavaScript do not appropriately represent the semantics and syntax of the code. Therefore, the generated samples may not properly represent the features of actual malicious JavaScript code. To address this issue, we propose a method to oversample malicious JavaScript code using CodeLlama, a large language model specialized for code generation. We constructed an imbalanced dataset consisting of 21,744 benign JavaScript code collected by crawling URLs and 8000 publicly available malicious JavaScript code categorized by collection year, and evaluated the accuracy of the proposed method. The recall of our method, which applied oversampling, was about 0.22 higher than existing methods and about 0.21 higher than the baseline without oversampling. Furthermore, we demonstrated that our method improves accuracy by applying it to both classification and feature extraction processes, and that the optimal ratio of original samples to oversampled samples is approximately 1:1 to 1:2.
Mamoru Mimura, Shinjiro Danjo
Eng. Appl. Artif. Intell.1
2024 Evasion Attempt for the Malicious PowerShell Detector Considering Feature Weights
Kou Sugiura, Mamoru Mimura
ICICS (1)2
2023 XSS Attack Detection by Attention Mechanism Based on Script Tags in URLs
Yuki Nakagawa, Mamoru Mimura
ISPEC2
2023 Detection of Malware Using Self-Attention Mechanism and Strings
Satoki Kanno, Mamoru Mimura
NSS2
2023 Impact of benign sample size on binary classification accuracy
abstract
Recently, there has been a significant increase in malware attacks and malicious traffic. Consequently, several machine learning-based detection models have been developed to detect them. However, the detection accuracy of these models is currently evaluated using different methodologies and datasets, with some studies overstating high detection rates. The lack of a common testing approach coupled with the limited datasets used for the experiments make it challenging to compare the performances of these models to identify those that provide superior detection accuracy. A few studies have focused on benign samples and their effects on detection accuracy. The datasets used in the experiments generally consist of benign and malicious samples; hence, binary classification is used in the machine learning models. In the binary classification task, the size of a benign sample affects the classification accuracy of malicious samples, that is, it can either improve or degrade detection accuracy. In this study, we propose a novel metric for evaluating accuracy degradation by increasing benign sample size. We mainly used the FFRI dataset, which consists of 11,243 malware samples and 250,000 benign samples, and evaluated the classification accuracy with extracted strings from the malware. In addition, we obtained other malware samples that we used as supplementary to the main dataset. We increased the number of benign samples for testing by tenfold, while maintaining the malicious sample and benign training sample sizes, which resulted in a decrease of 0.293 in the F1 score. Furthermore, we confirmed that using a sufficiently sized benign training sample set mitigates accuracy degradation. Our metric can be beneficial for evaluating the benign sample size needed in binary classification and comparing accuracy.
Mamoru Mimura
Expert Syst. Appl.1
2022 Evaluating the Possibility of Evasion Attacks to Machine Learning-Based Models for Malicious PowerShell Detection
Yuki Mezawa, Mamoru Mimura
ISPEC2
2021 Automating post-exploitation with deep reinforcement learning
abstract
In order to assess the risk of information systems, it is important to investigate the behavior of the attacker after successful exploitation (post-exploitation). However, the audit requires the experts, and to the best of our knowledge, there are no solutions to automate this process. This paper proposes a method of automating post-exploitation by combining deep reinforcement learning and the PowerShell Empire, which is famous as a post-exploitation framework. Our reinforcement learning agents select one of the PowerShell Empire modules as an action. The state of the agents is defined by 10 parameters such as type of account that was compromised by the agents. In the learning phase, we compared the learning progress of the 3 reinforcement learning models: A2C, Q-Learning, and SARSA. The result shows that the A2C could gain reward most efficiently. Moreover, the behavior of the trained agents are evaluated in a test domain network. The results show that the trained agent using A2C could obtain the administrative privileges to the domain controller.
Ryusei Maeda, Mamoru Mimura
Comput. Secur.2
2020 Adjusting lexical features of actual proxy logs for intrusion detection
abstract
Modern http-based malware imitates benign traffic to evade detection. To detect unseen malicious traffic, we proposed a linguistic-based detection method for proxy logs. This method extracts words as feature vectors automatically with natural language techniques, and discriminates between benign traffic and malicious traffic. The previous method generates a corpus from all the extracted words which contain trivial words. To generate discriminative feature representation, a corpus has to be effectively summarized. In actual proxy logs, benign traffic is dominant, and occupies malicious feature representation. Hence, the imbalance between benign and malicious traffic occurs. Moreover, a malicious paragraph might be mixed with some benign proxy logs. Therefore, the previous method does not perform accuracy in practical environment. This paper demonstrates that our previous method is not effective in actual proxy logs because of the imbalance. To mitigate the imbalance, our method adjusts lexical features of actual proxy logs based on the word importance. Our method does not adjust the number of each class such as the traditional sampling techniques. We performed cross-validation and timeline analysis with captured pcap files from Exploit Kit and actual proxy logs. The experimental results show our method could detect unseen malicious traffic in actual proxy logs. Moreover, we examine the effectiveness of mixing benign logs in each proportion. The best F-measure achieves 0.95 in the timeline analysis.
Mamoru Mimura
J. Inf. Secur. Appl.1
2020 Using fake text vectors to improve the sensitivity of minority class for macro malware detection
abstract
To detect new malware, machine learning approaches require many training samples. These training samples contribute to build an accurate model. To maintain the accuracy, collecting comprehensive samples continuously is very important. However, new malicious samples appear one after another, and thereby making it difficult. Hence, actual small training samples do not likely to represent the entire population adequately. Despite this gap between ideal and reality, few studies have addressed this practical problem in macro malware. To enhance small training samples, data augmentation is efficient in the field of image recognition. Data augmentation with Generative Adversarial Networks (GANs) is a reasonable approach for oversampling the minority class. A major difficulty of GANs is to generate fake samples that represent the context. This paper attempts to generate fake text vectors with Paragraph Vector to enhance small training samples. Paragraph Vector is a model to convert text into vectors, which represents the context and numerical distance. These features allow to directly vary each element of the vectors. Our method adds random noise to the vectors to generate fake text vectors which represent the context. This paper applies this technique to detect new malicious VBA (Visual Basic for Applications) macros to address the practical problem. This generic technique could be used for not only malware detection, but also any imbalanced and contextual data. To simulate small training samples, we reduce the malicious samples, and generate fake samples from the reduced ones. The experimental result shows that the fake samples enhance our model, and improve the detection rate.
Mamoru Mimura
J. Inf. Secur. Appl.1
2019 Using Sparse Composite Document Vectors to Classify VBA Macros
Mamoru Mimura
NSS1
2018 A Linguistic Approach Towards Intrusion Detection in Actual Proxy Logs
Mamoru Mimura, Hidema Tanaka
ICICS1
2018 Macros Finder: Do You Remember LOVELETTER?
Hiroya Miura, Mamoru Mimura, Hidema Tanaka
ISPEC2