Toshiki Shibahara

dblp:168/9586 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0002-2192-4355ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 6 · 1 first-author · 3 since 2021Computer networks · 5 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Membership Inference Attack Against Bayesian Neural Network
abstract
Membership Inference Attacks (MIAs) have been actively studied to evaluate the privacy risks of training data. However, existing MIAs focus on deterministic deep neural networks (DNNs). In this paper, we extend MIAs against deterministic DNNs to be applicable to Bayesian NNs (BNNs) and evaluate the privacy risks of BNNs. Specifically, we propose four MIAs, each differing in the extent to which detailed information on the posterior predictive distribution is exploited. Additionally, considering the trait of BNNs that produce different outputs for the same input, we also propose four multiple query attacks where attackers query BNNs multiple times using the same data and conduct MIAs with aggregated outputs. We conducted experiments using two tabular datasets for regression tasks and three representative BNNs. Our experiments show that outputting more detailed information on the posterior predictive distribution poses a higher privacy risk. Additionally, we found that the privacy risks may be underestimated if attackers exploiting multiple queries are not assumed.
Toshiki Shibahara, Takayuki Miura, Masanobu Kii, Atsunori Ichikawa
ICC1
2024 Noisy Label Detection for Multi-labeled Malware
abstract
Malware attacks have become increasingly prevalent, and accurate and reliable malware detection is essential for combating them. However, mislabeling, where data is given a different/noisy label than its true label, can significantly affect the accuracy and reliability of malware detection. In this paper, we propose a new method for detecting noisy labels in multi-labeled malware datasets. Our approach involves a new transformation method that allows malware datasets with multiple labels to be treated as data with a single label without losing any essential information. We also introduce a new machine learning model for detecting mislabeling, based on this transformation method. We conducted experiments on a real-world malware dataset to evaluate the effectiveness of our proposed method, and our findings indicate that our method can detect mislabels with high accuracy, up to 94.7%. Our research aims to improve the quality of labeling and reduce the factors contributing to mislabeling in malware datasets, leading to more accurate and reliable malware detection.
Naoki Fukushi, Toshiki Shibahara, Hiroki Nakano, Takashi Koide, Daiki Chiba 0001
CCNC2
2024 Efficiently Calculating Stronger Lower Bound for Differentially Private SGD in Black-Box Setting
abstract
Differentially private stochastic gradient descent (DP-SGD) is widely used to protect the privacy of training datasets for deep neural networks. Recently, a lower bound of ∊ for DP-SGD has been gaining attention to know whether the upper bound is tight or not. The lower bound is empirically calculated by repeating a game of an attacker and trainer. In the game, the trainer builds a model using one of the neighboring datasets differing by only a target sample. Then the attacker pre-dicts whether the target sample is used in training. In this paper, we focus on a black-box and realistic setting and propose methods for efficiently calculating stronger lower bounds by solving two challenges of lower bound calculation: computational cost and vulnerable neighboring datasets. To reduce the computational cost, we propose a multiple-sample game where an attacker predicts whether multiple target samples are used in training. To make vulnerable neighboring datasets, we propose three methods based on the analysis of vulnerable samples: vulnerability-based selection, label manipulation, and perturbation. We evaluated our methods using three realistic datasets: MNIST, CIFAR-10, and CIFAR-100, and two neural networks: a six-layer convolutional neural network and ResNetl8. Regarding the multiple-sample game, we confirmed that it produced lower bounds similar to those calculated with the prior game while reducing the computational cost by a factor of 1/100. Regarding the neighboring datasets, we compared our datasets with three existing ones and found that we can obtain 1.3 to 57.9 times stronger lower bounds.
Toshiki Shibahara, Takayuki Miura, Masanobu Kii, Atsunori Ichikawa
COMPSAC1
2024 Reusability Evaluation of Reports in Security Operation Centers for IoT with Sentence ALBERT and Jaccard Similarity
Masaru Matsubayashi, Toshiki Shibahara, Takuma Koyama, Masashi Tanaka
SecureComm (3)2
2023 Do Backdoors Assist Membership Inference Attacks?
Yumeki Goto, Nami Ashizawa, Toshiki Shibahara, Naoto Yanai
SecureComm (2)3
2023 Interpreting Graph-Based Sybil Detection Methods as Low-Pass Filtering
abstract
Online social networks (OSNs) are threatened by Sybil attacks, which create fake accounts (also called Sybils) on OSNs and use them for various malicious activities. Therefore, Sybil detection is a fundamental task for OSN security. Most existing Sybil detection methods are based on the graph structure of OSNs, and various methods have been proposed recently. However, although almost all methods have been compared experimentally in terms of detection performance and noise robustness, theoretical understanding of them is still lacking. In this study, we show that existing graph-based Sybil detection methods can be interpreted in a unified framework of low-pass filtering. This framework enables us to theoretically compare and analyze each method from two perspectives: filter kernel properties and the spectrum of shift matrices. Our analysis reveals that the detection performance of each method depends on the effectiveness of the low-pass filtering. Furthermore, on the basis of the analysis, we propose a novel Sybil detection method called SybilHeat. Numerical experiments on synthetic graphs and real social networks demonstrate that SybilHeat performs consistently well on graphs with various structural properties. This study lays a theoretical foundation for graph-based Sybil detection and leads to a better understanding of Sybil detection methods.
Satoshi Furutani, Toshiki Shibahara, Mitsuaki Akiyama, Masaki Aida
IEEE Trans. Inf. Forensics Secur.2
2022 Objection!: Identifying Misclassified Malicious Activities with XAI
abstract
Many studies have been conducted to detect various malicious activities in cyberspace using classifiers built by machine learning. However, it is natural for any classifier to make mistakes, and hence, human verification is necessary. One method to address this issue is eXplainable AI (XAI), which provides a reason for the classification result. However, when the number of classification results to be verified is large, it is not realistic to check the output of the XAI for all cases. In addition, it is sometimes difficult to interpret the output of XAI. In this study, we propose a machine learning model called classification verifier that verifies the classification results by using the output of XAI as a feature and raises objections when there is doubt about the reliability of the classification results. The results of experiments on malicious website detection and malware detection show that the proposed classification verifier can efficiently identify misclassified malicious activities.
Koji Fujita, Toshiki Shibahara, Daiki Chiba 0001, Mitsuaki Akiyama, Masato Uchida
ICC2
2020 Detecting Malware-infected Hosts Using Templates of Multiple HTTP Requests
abstract
In this paper, we propose a method for detecting malware-infected hosts with a high rate of detection and a low rate of false positives without using any data on benign communication. Based on the fact that many malware-infected hosts generate multiple HTTP requests, we propose a method using the templates of sets of those HTTP requests. For each malware, this method generates a template that comprises the set of templates of the HTTP requests that the malware generates. We call the set of templates group template. It then detects malware-infected hosts by comparing the set of monitored HTTP requests with the group templates.
Taiga Hokaguchi, Yuichi Ohsita, Toshiki Shibahara, Daiki Chiba 0001, Mitsuaki Akiyama, Masayuki Murata 0001
CCNC3
2020 Sybil Detection as Graph Filtering
abstract
Sybils are users created for carrying out nefarious actions in online social networks (OSNs) and threaten the security of OSNs. Therefore, Sybil detection is an urgent security task, and various detection methods have been proposed. Existing Sybil detection methods are based on the relationship (i.e., graph structure) of users in OSNs. Structure-based methods can be classified into two categories: Random Walk (RW)-based and Belief Propagation (BP)-based. However, although almost all methods have been experimentally evaluated in terms of their performance and robustness to noise, the theoretical understanding of them is insufficient. In this paper, we interpret the Sybil detection problem from the viewpoint of graph signal processing and provide a framework to formulate RW- and BPbased methods as low-pass filtering. This framework enables us to theoretically compare RW- and BP-based methods and explain why BP-based methods perform well for scale-free graphs, unlike RW-based methods. Furthermore, by this framework, we relate RW- and BP-based methods and Graph Neural Networks (GNNs) and discuss the difference among these methods. Finally, we evaluate the validity of this framework through numerical experiments.
Satoshi Furutani, Toshiki Shibahara, Kunio Hato, Mitsuaki Akiyama, Masaki Aida
GLOBECOM2
2020 Special-purpose Model Extraction Attacks: Stealing Coarse Model with Fewer Queries
abstract
Model extraction (ME) attacks have been shown to cause financial losses for Machine-Learning-as-a-Service (MLaaS) providers. Attackers steal ML models on MLaaS platforms by building substitute models using queries to and responses from MLaaS platforms. The ML models targeted by attackers are called targeted models. In previous studies, researchers have assumed that attackers build substitute models that classify the same number of classes as targeted ones, which classify thousands or millions of classes to meet users' diverse expectations. We call such models general-purpose models. In fact, attackers can monetize stolen models if they accurately distinguish some classes from others. We call such models special-purpose models. For instance, a model that detects vehicles is useful for collision avoidance systems, and a model that detects wild animals is useful to drive them away from agricultural land. In this work, we investigate a threat of special-purpose ME attacks that steal special-purpose models. Our experimental results show that attackers can build an accurate special-purpose model, which achieves an 80% f-measure, with as few as 100 queries in the worst case. We discuss the difficulty in preventing the attacks with previously proposed defense methods and point out the necessity of a new defense method.
Rina Okada, Zen Ishikura, Toshiki Shibahara, Satoshi Hasegawa
TrustCom3
2019 Graph Signal Processing for Directed Graphs Based on the Hermitian Laplacian
Satoshi Furutani, Toshiki Shibahara, Mitsuaki Akiyama, Kunio Hato, Masaki Aida
ECML/PKDD (1)2
2017 Detecting Malicious Websites by Integrating Malicious, Benign, and Compromised Redirection Subgraph Similarities
abstract
To expose more users to threats of drive-by download attacks, attackers compromise vulnerable websites discovered by search engines and redirect clients to malicious websites created with exploit kits. Security researchers and vendors have tried to prevent the attacks by detecting malicious data, i.e., malicious URLs, web content, and redirections. However, attackers conceal a part of malicious data with evasion techniques to circumvent detection systems. In this paper, we propose a system for detecting malicious websites without collecting all malicious data. Even if we cannot observe a part of malicious data, we can always observe compromised websites. Since vulnerable websites are discovered by search engines, compromised websites have similar traits. Therefore, we built a classifier by leveraging not only malicious websites but also compromised websites. More precisely, we convert all websites observed at the time of access into a redirection graph and classify it by integrating similarities between its subgraphs and redirection subgraphs shared across malicious, benign, and compromised websites. As a result of evaluating our system with crawling data of 455,860 websites, we found that the system achieved a 91.7% true positive rate for malicious websites containing exploit URLs at a low false positive rate of 0.1%. Moreover, it detected 143 more evasive malicious websites than conventional systems.
Toshiki Shibahara, Yuta Takata, Mitsuaki Akiyama, Takeshi Yagi, Takeshi Yada
COMPSAC (1)1
2017 Malicious URL sequence detection using event de-noising convolutional neural network
abstract
Attackers have increased the number of infected hosts by redirecting users of compromised popular websites toward websites that exploit vulnerabilities of a browser and its plugins. To prevent damage, detecting infected hosts based on proxy logs, which are generally recorded on enterprise networks, is gaining attention rather than blacklist-based filtering because creating blacklists has become difficult due to the short lifetime of malicious domains and concealment of exploit code. Since information extracted from one URL is limited, we focus on a sequence of URLs that includes artifacts of malicious redirections. We propose a system for detecting malicious URL sequences from proxy logs with a low false positive rate. To elucidate an effective approach of malicious URL sequence detection, we compared three approaches: individual-based approach, convolutional neural network (CNN), and our newly developed event de-noising CNN (EDCNN). Our EDCNN is a new CNN to reduce the negative effect of benign URLs redirected from compromised websites included in malicious URL sequences. Our evaluation shows that the EDCNN lowers the operation cost of malware infection by reducing 47% of false alerts compared with a CNN when users access compromised websites but do not obtain exploit code due to browser fingerprinting.
Toshiki Shibahara, Kohei Yamanishi, Yuta Takata, Daiki Chiba 0001, Mitsuaki Akiyama, Takeshi Yagi, Yuichi Ohsita, Masayuki Murata 0001
ICC1
2017 Understanding the origins of mobile app vulnerabilities: a large-scale measurement study of free and paid apps
abstract
This paper reports a large-scale study that aims to understand how mobile application (app) vulnerabilities are associated with software libraries. We analyze both free and paid apps. Studying paid apps was quite meaningful because it helped us understand how differences in app development/maintenance affect the vulnerabilities associated with libraries. We analyzed 30k free and paid apps collected from the official Android marketplace. Our extensive analyses revealed that approximately 70%/50% of vulnerabilities of free/paid apps stem from software libraries, particularly from third-party libraries. Somewhat paradoxically, we found that more expensive/popular paid apps tend to have more vulnerabilities. This comes from the fact that more expensive/popular paid apps tend to have more functionality, i.e., more code and libraries, which increases the probability of vulnerabilities. Based on our findings, we provide suggestions to stakeholders of mobile app distribution ecosystems.
Takuya Watanabe 0001, Mitsuaki Akiyama, Fumihiro Kanei, Eitaro Shioji, Yuta Takata, Yuta Ishii, Toshiki Shibahara, Takeshi Yagi, Tatsuya Mori 0003
MSR8
2016 DomainProfiler: Discovering Domain Names Abused in Future
abstract
Cyber attackers abuse the domain name system (DNS) to mystify their attack ecosystems, they systematically generate a huge volume of distinct domain names to make it infeasible for blacklisting approaches to keep up with newly generated malicious domain names. As a solution to this problem, we propose a system for discovering malicious domain names that will likely be abused in future. The key idea with our system is to exploit temporal variation patterns (TVPs) of domain names. The TVPs of domain names include information about how and when a domain name has been listed in legitimate/popular and/or malicious domain name lists. On the basis of this idea, our system actively collects DNS logs, analyzes their TVPs, and predicts whether a given domain name will be used for malicious purposes. Our evaluation revealed that our system can predict malicious domain names 220 days beforehand with a true positive rate of 0.985.
Daiki Chiba 0001, Takeshi Yagi, Mitsuaki Akiyama, Toshiki Shibahara, Takeshi Yada, Tatsuya Mori 0003, Shigeki Goto
DSN4
2016 Efficient Dynamic Malware Analysis Based on Network Behavior Using Deep Learning
abstract
Malware authors or attackers always try to evade detection methods to accomplish their mission. Such detection methods are broadly divided into three types: static feature, host-behavior, and network-behavior based. Static feature-based methods are evaded using packing techniques. Host- behavior-based methods also can be evaded using some code injection methods, such as API hook and dynamic link library hook. This arms race regarding static feature-based and host-behavior- based methods increases the importance of network-behavior-based methods. The necessity of communication between infected hosts and attackers makes it difficult to evade network-behavior- based methods. The effectiveness of such methods depends on how we collect a variety of communications by using malware samples. However, analyzing all new malware samples for a long period is infeasible. Therefore, we propose a method for determining whether dynamic analysis should be suspended based on network behavior to collect malware communications efficiently and exhaustively. The key idea behind our proposed method is focused on two characteristics of malware communication: the change in the communication purpose and the common latent function. These characteristics of malware communications resemble those of natural language from the viewpoint of data structure, and sophisticated analysis methods have been proposed in the field of natural language processing. For this reason, we applied the recursive neural network, which has recently exhibited high classification performance, to our proposed method. In the evaluation with 29,562 malware samples, our proposed method reduced 67.1% of analysis time while keeping the coverage of collected URLs to 97.9% of the method that continues full analyses.
Toshiki Shibahara, Takeshi Yagi, Mitsuaki Akiyama, Daiki Chiba 0001, Takeshi Yada
GLOBECOM1
2015 POSTER: Detecting Malicious Web Pages based on Structural Similarity of Redirection Chains
abstract
Detecting malicious web pages used in attacks and building blacklists and signatures from them are done to protect users against drive-by download attacks. Gathering the content on web pages by crawling and evaluating it to check if it is malicious can help in detecting malicious web pages. Methods that apply supervised machine learning to this evaluation are proposed for detecting malicious web pages from a massive amount of web pages. However, these methods need manual inspections for preparing training data when classifiers are retrained in accordance with changes in the content on malicious web pages. In this paper, we propose a method that evaluates whether web pages are malicious and needs only the discrimination results of web pages identified by high-interaction honeyclients to prepare training data. This method evaluates maliciousness on the basis of the structural similarity of redirection chains arising from drive-by download attacks. The results of our experiments with two years of data showed that the accuracy of our method was about 20\% higher than that of the previous method.
Toshiki Shibahara, Takeshi Yagi, Mitsuaki Akiyama, Yuta Takata, Takeshi Yada
CCS1