EDBT 2026 Demo / reviewers in the wild / expert
Naoki Fukushi
dblp:248/3108
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2024
0009-0007-8103-8244ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Noisy Label Detection for Multi-labeled MalwareabstractMalware attacks have become increasingly prevalent, and accurate and reliable malware detection is essential for combating them. However, mislabeling, where data is given a different/noisy label than its true label, can significantly affect the accuracy and reliability of malware detection. In this paper, we propose a new method for detecting noisy labels in multi-labeled malware datasets. Our approach involves a new transformation method that allows malware datasets with multiple labels to be treated as data with a single label without losing any essential information. We also introduce a new machine learning model for detecting mislabeling, based on this transformation method. We conducted experiments on a real-world malware dataset to evaluate the effectiveness of our proposed method, and our findings indicate that our method can detect mislabels with high accuracy, up to 94.7%. Our research aims to improve the quality of labeling and reduce the factors contributing to mislabeling in malware datasets, leading to more accurate and reliable malware detection. Naoki Fukushi, Toshiki Shibahara, Hiroki Nakano, Takashi Koide, Daiki Chiba 0001 |
CCNC | 1 |
| 2024 | ChatSpamDetector: Leveraging Large Language Models for Effective Phishing Email Detection
Takashi Koide, Naoki Fukushi, Hiroki Nakano, Daiki Chiba 0001 |
SecureComm (3) | 2 |
| 2023 | Canary in Twitter Mine: Collecting Phishing Reports from Experts and Non-expertsabstractThe rise in phishing attacks via e-mail and short message service (SMS) has not slowed down at all. The first thing we need to do to combat the ever-increasing number of phishing attacks is to collect and characterize more phishing cases that reach end users. Without understanding these characteristics, anti-phishing countermeasures cannot evolve. In this study, we propose an approach using Twitter as a new observation point to immediately collect and characterize phishing cases via e-mail and SMS that evade countermeasures and reach users. Specifically, we propose CrowdCanary, a system capable of structurally and accurately extracting phishing information (e.g., URLs and domains) from tweets about phishing by users who have actually discovered or encountered it. In our three months of live operation, CrowdCanary identified 35,432 phishing URLs out of 38,935 phishing reports, 31,960 (90.2%) of these phishing URLs were later detected by the anti-virus engine. We analyzed users who shared phishing threats by categorizing them into two groups: experts and non-experts. As a results, we discovered that CrowdCanary extracts non-expert report-specific information, like company brand name in tweets, phishing attack details from tweet images, and pre-redirect landing page information. Hiroki Nakano, Daiki Chiba 0001, Takashi Koide, Naoki Fukushi, Takeshi Yagi, Takeo Hariu, Katsunari Yoshioka, Tsutomu Matsumoto |
ARES | 4 |
| 2023 | PhishReplicant: A Language Model-based Approach to Detect Generated Squatting Domain NamesabstractDomain squatting is a technique used by attackers to create domain names for phishing sites. In recent phishing attempts, we have observed many domain names that use multiple techniques to evade existing methods for domain squatting. These domain names, which we call generated squatting domains (GSDs), are quite different in appearance from legitimate domain names and do not contain brand names, making them difficult to associate with phishing. In this paper, we propose a system called PhishReplicant that detects GSDs by focusing on the linguistic similarity of domain names. We analyzed newly registered and observed domain names extracted from certificate transparency logs, passive DNS, and DNS zone files. We detected 3,498 domain names acquired by attackers in a four-week experiment, of which 2,821 were used for phishing sites within a month of detection. We also confirmed that our proposed system outperformed existing systems in both detection accuracy and number of domain names detected. As an in-depth analysis, we examined 205k GSDs collected over 150 days and found that phishing using GSDs was distributed globally. However, attackers intensively targeted brands in specific regions and industries. By analyzing GSDs in real time, we can block phishing sites before or immediately after they appear. Takashi Koide, Naoki Fukushi, Hiroki Nakano, Daiki Chiba 0001 |
ACSAC | 2 |
| 2021 | Analyzing Security Risks of Ad-Based URL Shortening Services Caused by Users' Behaviors
Naoki Fukushi, Takashi Koide, Daiki Chiba 0001, Hiroki Nakano, Mitsuaki Akiyama |
SecureComm (2) | 1 |
| 2019 | Exploration into Gray Area: Efficient Labeling for Malicious Domain Name DetectionabstractThis paper presents a method to reduce the labeling cost when acquiring training data for a system that detects malicious domain names by supervised machine learning. The conventional system requires large quantities of both benign and malicious domain names to be prepared as training data to obtain a classifier with high classification accuracy. In general, malicious domain names are observed less frequently than benign domain names. Therefore, it is difficult to acquire a large number of malicious domain names without a dedicated labeling method. We propose a method based on active learning that labels data around the decision boundary of classification, i.e., in the gray area, and we show that the classification accuracy can be improved by only using approximately 2.5% of the training data used by the conventional system. An additional disadvantage of the conventional system is that, if the classifier is trained with a small amount of training data, its generalization ability cannot be guaranteed. We propose a method based on ensemble learning that integrates multiple classifiers, and we show that the classification accuracy can be stabilized and improved. Naoki Fukushi, Daiki Chiba 0001, Mitsuaki Akiyama, Masato Uchida |
COMPSAC (1) | 1 |