EDBT 2026 Demo / reviewers in the wild / expert
Mamoru Mimura
dblp:52/9355
· DBLP profile ↗
12ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0003-4323-9911ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 10 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Oversampling method using large language model for malicious JavaScript code detectionabstractDetecting malicious JavaScript code on websites remains a significant issue. Detection methods using machine learning models have been proposed to detect these attacks in real time. In targeted attacks, malware is often customized for a specific organization. To defend against these attacks, machine learning models have to be trained from a small number of past malicious samples. Traditional oversampling techniques applied to JavaScript do not appropriately represent the semantics and syntax of the code. Therefore, the generated samples may not properly represent the features of actual malicious JavaScript code. To address this issue, we propose a method to oversample malicious JavaScript code using CodeLlama, a large language model specialized for code generation. We constructed an imbalanced dataset consisting of 21,744 benign JavaScript code collected by crawling URLs and 8000 publicly available malicious JavaScript code categorized by collection year, and evaluated the accuracy of the proposed method. The recall of our method, which applied oversampling, was about 0.22 higher than existing methods and about 0.21 higher than the baseline without oversampling. Furthermore, we demonstrated that our method improves accuracy by applying it to both classification and feature extraction processes, and that the optimal ratio of original samples to oversampled samples is approximately 1:1 to 1:2. Mamoru Mimura, Shinjiro Danjo |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Evasion Attempt for the Malicious PowerShell Detector Considering Feature Weights
Kou Sugiura, Mamoru Mimura |
ICICS (1) | 2 |
| 2023 | XSS Attack Detection by Attention Mechanism Based on Script Tags in URLs
Yuki Nakagawa, Mamoru Mimura |
ISPEC | 2 |
| 2023 | Detection of Malware Using Self-Attention Mechanism and Strings
Satoki Kanno, Mamoru Mimura |
NSS | 2 |
| 2023 | Impact of benign sample size on binary classification accuracyabstractRecently, there has been a significant increase in malware attacks and malicious traffic. Consequently, several machine learning-based detection models have been developed to detect them. However, the detection accuracy of these models is currently evaluated using different methodologies and datasets, with some studies overstating high detection rates. The lack of a common testing approach coupled with the limited datasets used for the experiments make it challenging to compare the performances of these models to identify those that provide superior detection accuracy. A few studies have focused on benign samples and their effects on detection accuracy. The datasets used in the experiments generally consist of benign and malicious samples; hence, binary classification is used in the machine learning models. In the binary classification task, the size of a benign sample affects the classification accuracy of malicious samples, that is, it can either improve or degrade detection accuracy. In this study, we propose a novel metric for evaluating accuracy degradation by increasing benign sample size. We mainly used the FFRI dataset, which consists of 11,243 malware samples and 250,000 benign samples, and evaluated the classification accuracy with extracted strings from the malware. In addition, we obtained other malware samples that we used as supplementary to the main dataset. We increased the number of benign samples for testing by tenfold, while maintaining the malicious sample and benign training sample sizes, which resulted in a decrease of 0.293 in the F1 score. Furthermore, we confirmed that using a sufficiently sized benign training sample set mitigates accuracy degradation. Our metric can be beneficial for evaluating the benign sample size needed in binary classification and comparing accuracy. Mamoru Mimura |
Expert Syst. Appl. | 1 |
| 2022 | Evaluating the Possibility of Evasion Attacks to Machine Learning-Based Models for Malicious PowerShell Detection
Yuki Mezawa, Mamoru Mimura |
ISPEC | 2 |
| 2021 | Automating post-exploitation with deep reinforcement learningabstractIn order to assess the risk of information systems, it is important to investigate the behavior of the attacker after successful exploitation (post-exploitation). However, the audit requires the experts, and to the best of our knowledge, there are no solutions to automate this process. This paper proposes a method of automating post-exploitation by combining deep reinforcement learning and the PowerShell Empire, which is famous as a post-exploitation framework. Our reinforcement learning agents select one of the PowerShell Empire modules as an action. The state of the agents is defined by 10 parameters such as type of account that was compromised by the agents. In the learning phase, we compared the learning progress of the 3 reinforcement learning models: A2C, Q-Learning, and SARSA. The result shows that the A2C could gain reward most efficiently. Moreover, the behavior of the trained agents are evaluated in a test domain network. The results show that the trained agent using A2C could obtain the administrative privileges to the domain controller. Ryusei Maeda, Mamoru Mimura |
Comput. Secur. | 2 |
| 2020 | Adjusting lexical features of actual proxy logs for intrusion detectionabstractModern http-based malware imitates benign traffic to evade detection. To detect unseen malicious traffic, we proposed a linguistic-based detection method for proxy logs. This method extracts words as feature vectors automatically with natural language techniques, and discriminates between benign traffic and malicious traffic. The previous method generates a corpus from all the extracted words which contain trivial words. To generate discriminative feature representation, a corpus has to be effectively summarized. In actual proxy logs, benign traffic is dominant, and occupies malicious feature representation. Hence, the imbalance between benign and malicious traffic occurs. Moreover, a malicious paragraph might be mixed with some benign proxy logs. Therefore, the previous method does not perform accuracy in practical environment. This paper demonstrates that our previous method is not effective in actual proxy logs because of the imbalance. To mitigate the imbalance, our method adjusts lexical features of actual proxy logs based on the word importance. Our method does not adjust the number of each class such as the traditional sampling techniques. We performed cross-validation and timeline analysis with captured pcap files from Exploit Kit and actual proxy logs. The experimental results show our method could detect unseen malicious traffic in actual proxy logs. Moreover, we examine the effectiveness of mixing benign logs in each proportion. The best F-measure achieves 0.95 in the timeline analysis. Mamoru Mimura |
J. Inf. Secur. Appl. | 1 |
| 2020 | Using fake text vectors to improve the sensitivity of minority class for macro malware detectionabstractTo detect new malware, machine learning approaches require many training samples. These training samples contribute to build an accurate model. To maintain the accuracy, collecting comprehensive samples continuously is very important. However, new malicious samples appear one after another, and thereby making it difficult. Hence, actual small training samples do not likely to represent the entire population adequately. Despite this gap between ideal and reality, few studies have addressed this practical problem in macro malware. To enhance small training samples, data augmentation is efficient in the field of image recognition. Data augmentation with Generative Adversarial Networks (GANs) is a reasonable approach for oversampling the minority class. A major difficulty of GANs is to generate fake samples that represent the context. This paper attempts to generate fake text vectors with Paragraph Vector to enhance small training samples. Paragraph Vector is a model to convert text into vectors, which represents the context and numerical distance. These features allow to directly vary each element of the vectors. Our method adds random noise to the vectors to generate fake text vectors which represent the context. This paper applies this technique to detect new malicious VBA (Visual Basic for Applications) macros to address the practical problem. This generic technique could be used for not only malware detection, but also any imbalanced and contextual data. To simulate small training samples, we reduce the malicious samples, and generate fake samples from the reduced ones. The experimental result shows that the fake samples enhance our model, and improve the detection rate. Mamoru Mimura |
J. Inf. Secur. Appl. | 1 |
| 2019 | Using Sparse Composite Document Vectors to Classify VBA Macros
Mamoru Mimura |
NSS | 1 |
| 2018 | A Linguistic Approach Towards Intrusion Detection in Actual Proxy Logs
Mamoru Mimura, Hidema Tanaka |
ICICS | 1 |
| 2018 | Macros Finder: Do You Remember LOVELETTER?
Hiroya Miura, Mamoru Mimura, Hidema Tanaka |
ISPEC | 2 |