EDBT 2026 Demo / reviewers in the wild / expert
Takashi Koide
dblp:169/7325
· DBLP profile ↗
18ranked-venue papers
4as first author
14since 2021 · last 2026
0009-0008-1942-0335ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 10 · 4 first-author · 6 since 2021Computer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PhishLumos: An Adaptive Multi-Agent System for Proactive Phishing Campaign Mitigation
Daiki Chiba 0001, Hiroki Nakano, Takashi Koide |
ICC | 3 |
| 2025 | DomainLynx: Leveraging Large Language Models for Enhanced Domain Squatting DetectionabstractDomain squatting poses a significant threat to Internet security, with attackers employing increasingly sophisticated techniques. This study introduces DomainLynx, an innovative compound AI system leveraging Large Language Models (LLMs) for enhanced domain squatting detection. Unlike existing methods focusing on predefined patterns for top-ranked domains, DomainLynx excels in identifying novel squatting techniques and protecting less prominent brands. The system's architecture integrates advanced data processing, intelligent domain pairing, and LLM-powered threat assessment. Crucially, DomainLynx incorporates specialized components that mitigate LLM hallucinations, ensuring reliable and context-aware detection. This approach enables efficient analysis of vast security data from diverse sources, including Certificate Transparency logs, Passive DNS records, and zone files. Evaluated on a curated dataset of 1,649 squatting domains, DomainLynx achieved 94.7% accuracy using Llama-3-70B. In a month-long real-world test, it detected 34,359 squatting domains from 2.09 million new domains, outperforming baseline methods by 2.5 times. This research advances Internet security by providing a versatile, accurate, and adaptable tool for combating evolving domain squatting threats. DomainLynx's approach paves the way for more robust, AI-driven cybersecurity solutions, enhancing protection for a broader range of online entities and contributing to a safer digital ecosystem. Daiki Chiba 0001, Hiroki Nakano, Takashi Koide |
CCNC | 3 |
| 2025 | DomainDynamics: Lifecycle-Aware Risk Timeline Construction for Domain NamesabstractThe persistent threat posed by malicious domain names in cyber-attacks underscores the urgent need for effective detection mechanisms. Traditional machine learning methods, while capable of identifying such domains, often suffer from high false positive and false negative rates due to their extensive reliance on historical data. Conventional approaches often overlook the dynamic nature of domain names, the purposes and ownership of which may evolve, potentially rendering risk assessments outdated or irrelevant. To address these shortcomings, we introduce DomainDynamics, a novel system designed to predict domain name risks by considering their lifecycle stages. DomainDynamics constructs a timeline for each domain, evaluating the characteristics of each domain at various points in time to make informed, temporal risk determinations. In an evaluation experiment involving over 85,000 actual malicious domains from malware and phishing incidents, DomainDynamics demonstrated a significant improvement in detection rates, achieving an 82.58% detection rate with a low false positive rate of 0.41%. This performance surpasses that of previous studies and commercial services, improving detection capability substantially. Daiki Chiba 0001, Hiroki Nakano, Takashi Koide |
CCNC | 3 |
| 2025 | DomainHarvester: Harvesting Infrequently Visited Yet Trustworthy Domain NamesabstractIn cybersecurity, allow lists play a crucial role in distinguishing safe websites from potential threats. Conventional methods for compiling allow lists, focusing heavily on website popularity, often overlook infrequently visited legitimate domains. This paper introduces DomainHarvester, a system aimed at generating allow lists that include trustworthy yet infrequently visited domains. By adopting an innovative bottom-up methodology that leverages the web's hyperlink structure, DomainHarvester identifies legitimate yet underrepresented domains. The system uses seed URLs to gather domain names, employing machine learning with a Transformer-based approach to assess their trustworthiness. DomainHarvester has developed two distinct allow lists: one with a global focus and another emphasizing local relevance. Compared to six existing top lists, DomainHarvester's allow lists show minimal overlaps, 4% globally and 0.1 % locally, while significantly reducing the risk of including malicious domains, thereby enhancing security. The contributions of this research are substantial, illuminating the overlooked aspect of trustworthy yet underrepresented domains and introducing DomainHarvester, a system that goes beyond traditional popularity-based metrics. Our methodology enhances the inclusivity and precision of allow lists, offering significant advantages to users and businesses worldwide, especially in non-English speaking regions. Daiki Chiba 0001, Hiroki Nakano, Takashi Koide |
CCNC | 3 |
| 2025 | ScamFerret: Detecting Scam Websites Autonomously with Large Language Models
Hiroki Nakano, Takashi Koide, Daiki Chiba 0001 |
DIMVA (1) | 2 |
| 2025 | PhishParrot: LLM-Driven Adaptive Crawling to Unveil Cloaked Phishing SitesabstractPhishing attacks continue to evolve, with cloaking techniques posing a significant challenge to detection efforts. Cloaking allows attackers to display phishing sites only to specific users while presenting legitimate pages to security crawlers, rendering traditional detection systems ineffective. This research proposes PhishParrot, a novel crawling environment optimization system designed to counter cloaking techniques. PhishParrot leverages the contextual analysis capabilities of Large Language Models (LLMs) to identify potential patterns in crawling information, enabling the construction of optimal user profiles capable of bypassing cloaking mechanisms. The system accumulates information on phishing sites collected from diverse environments. It then adapts browser settings and network configurations to match the attacker’s target user conditions based on information extracted from similar cases. A 21-day evaluation showed that PhishParrot improved detection accuracy by up to 33.8% over standard analysis systems, yielding 91 distinct crawling environments for diverse conditions targeted by attackers. The findings confirm that the combination of similar-case extraction and LLM-based context analysis is an effective approach for detecting cloaked phishing attacks. Hiroki Nakano, Takashi Koide, Daiki Chiba 0001 |
GLOBECOM | 2 |
| 2025 | Uncovering Suspicious Posts on Abandoned BlogsabstractAs the Internet evolves, the diversification of information dissemination has led to the growth of various social media platforms and an increase in abandoned blogs. These abandoned blogs, if not properly managed, present a significant security risk. In this study, we introduce “Zombified Blogs,” defined as blogs deserted by their administrators and targeted for continuous submissions of unsolicited malicious content. We explore the causes and effects of this global phenomenon on a large scale, using data from multiple blog services over the past 18 years for the first time. To this end, we propose an innovative method for the continuous identification and collection of Zombified Blogs using a search engine. We analyzed data spanning 18 years, encompassing approximately 1.5 million blog posts, including 258, 547 malicious posts, across 1,141 Zombified Blogs on six different blog services. Our findings show a significant rise in malicious posts, primarily social engineering attacks, over the past four years on previously active blogs. We identify the email posting function as a critical area exploited by attackers, leading to the unintended posting of content. Finally, we provide an overview of the current challenges facing blog services and administrators, and discuss potential content management solutions. Hiroki Nakano, Takashi Koide, Daiki Chiba 0001, Katsunari Yoshioka, Tsutomu Matsumoto |
ICC | 2 |
| 2025 | DomainDynamics: Advancing lifecycle-based risk assessment of domain namesabstractThe persistent threat of malicious domains in cybersecurity necessitates robust detection systems. Traditional machine learning approaches often struggle to accurately assess domain name risks due to their static analysis methods and lack of consideration for temporal changes in domain attributes. To address these limitations, we developed DomainDynamics, a novel system that evaluates domain name risks by analyzing their lifecycle phases. This study provides a comprehensive evaluation and refinement of the DomainDynamics framework. The system creates temporal profiles for domains and assesses their attributes at various stages, enabling informed, time-sensitive risk assessments. Our initial evaluation, involving over 85,000 malicious domains, achieved an 82.58% detection rate with a low 0.41% false positive rate. We expanded our research to include benchmarking against commercial services, feature significance analysis using interpretable AI techniques, and detailed case studies. This investigation not only validates the effectiveness of DomainDynamics but also reveals temporal indicators of malicious intent. Our findings demonstrate the advantages of lifecycle-based analysis over static methodologies, providing valuable insights for practical cybersecurity applications. • DomainDynamics: lifecycle-based detection (82.58% Detection Rate, 0.41% FPR). • Temporal analysis outperforms static methods on 85,000+ malicious domains. • Explainable AI offers deep insights into domain risk through real case studies. • Benchmarks against commercial tools confirm lifecycle indicators’ significance. • Enhances malicious domain detection via temporal context and dynamic changes. Daiki Chiba 0001, Hiroki Nakano, Takashi Koide |
Comput. Secur. | 3 |
| 2024 | Noisy Label Detection for Multi-labeled MalwareabstractMalware attacks have become increasingly prevalent, and accurate and reliable malware detection is essential for combating them. However, mislabeling, where data is given a different/noisy label than its true label, can significantly affect the accuracy and reliability of malware detection. In this paper, we propose a new method for detecting noisy labels in multi-labeled malware datasets. Our approach involves a new transformation method that allows malware datasets with multiple labels to be treated as data with a single label without losing any essential information. We also introduce a new machine learning model for detecting mislabeling, based on this transformation method. We conducted experiments on a real-world malware dataset to evaluate the effectiveness of our proposed method, and our findings indicate that our method can detect mislabels with high accuracy, up to 94.7%. Our research aims to improve the quality of labeling and reduce the factors contributing to mislabeling in malware datasets, leading to more accurate and reliable malware detection. Naoki Fukushi, Toshiki Shibahara, Hiroki Nakano, Takashi Koide, Daiki Chiba 0001 |
CCNC | 4 |
| 2024 | ChatSpamDetector: Leveraging Large Language Models for Effective Phishing Email Detection
Takashi Koide, Naoki Fukushi, Hiroki Nakano, Daiki Chiba 0001 |
SecureComm (3) | 1 |
| 2023 | Canary in Twitter Mine: Collecting Phishing Reports from Experts and Non-expertsabstractThe rise in phishing attacks via e-mail and short message service (SMS) has not slowed down at all. The first thing we need to do to combat the ever-increasing number of phishing attacks is to collect and characterize more phishing cases that reach end users. Without understanding these characteristics, anti-phishing countermeasures cannot evolve. In this study, we propose an approach using Twitter as a new observation point to immediately collect and characterize phishing cases via e-mail and SMS that evade countermeasures and reach users. Specifically, we propose CrowdCanary, a system capable of structurally and accurately extracting phishing information (e.g., URLs and domains) from tweets about phishing by users who have actually discovered or encountered it. In our three months of live operation, CrowdCanary identified 35,432 phishing URLs out of 38,935 phishing reports, 31,960 (90.2%) of these phishing URLs were later detected by the anti-virus engine. We analyzed users who shared phishing threats by categorizing them into two groups: experts and non-experts. As a results, we discovered that CrowdCanary extracts non-expert report-specific information, like company brand name in tweets, phishing attack details from tweet images, and pre-redirect landing page information. Hiroki Nakano, Daiki Chiba 0001, Takashi Koide, Naoki Fukushi, Takeshi Yagi, Takeo Hariu, Katsunari Yoshioka, Tsutomu Matsumoto |
ARES | 3 |
| 2023 | PhishReplicant: A Language Model-based Approach to Detect Generated Squatting Domain NamesabstractDomain squatting is a technique used by attackers to create domain names for phishing sites. In recent phishing attempts, we have observed many domain names that use multiple techniques to evade existing methods for domain squatting. These domain names, which we call generated squatting domains (GSDs), are quite different in appearance from legitimate domain names and do not contain brand names, making them difficult to associate with phishing. In this paper, we propose a system called PhishReplicant that detects GSDs by focusing on the linguistic similarity of domain names. We analyzed newly registered and observed domain names extracted from certificate transparency logs, passive DNS, and DNS zone files. We detected 3,498 domain names acquired by attackers in a four-week experiment, of which 2,821 were used for phishing sites within a month of detection. We also confirmed that our proposed system outperformed existing systems in both detection accuracy and number of domain names detected. As an in-depth analysis, we examined 205k GSDs collected over 150 days and found that phishing using GSDs was distributed globally. However, attackers intensively targeted brands in specific regions and industries. By analyzing GSDs in real time, we can block phishing sites before or immediately after they appear. Takashi Koide, Naoki Fukushi, Hiroki Nakano, Daiki Chiba 0001 |
ACSAC | 1 |
| 2021 | Detecting Event-synced Navigation Attacks across User-generated Content PlatformsabstractWith the spread of service platforms that enable users to generate content, people use user-generated content (UGC) to search for and access information on the web instead of search engines. Attackers can also use UGC on a service platform (UGC platform) to spread web-based social engineering (SE) attacks to a large number of people. In this paper, we focus on a type of web-based SE attack, called an event-synced navigation attack, which generates UGC with links navigating users to malicious websites and distribute it synced with a real-life event at a specific time. To understand the attacks in the wild, we propose a new system for detecting event-synced navigation attacks in real time by capturing the inevitable footprints left by attacks that affect a large number of users. We evaluate each of the three steps of the proposed system and finally find that the system can classify malicious and non-malicious UGC with 97% accuracy. Furthermore, we perform a comprehensive measurement study on event-synced navigation attacks spread from popular UGC platforms (Twitter, Facebook, YouTube, and Reddit) and confirm that many event-synced navigation attacks are deployed in the wild. Hiroki Nakano, Daiki Chiba 0001, Takashi Koide, Mitsuaki Akiyama |
COMPSAC | 3 |
| 2021 | Analyzing Security Risks of Ad-Based URL Shortening Services Caused by Users' Behaviors
Naoki Fukushi, Takashi Koide, Daiki Chiba 0001, Hiroki Nakano, Mitsuaki Akiyama |
SecureComm (2) | 2 |
| 2020 | To Get Lost is to Learn the Way: Automatically Collecting Multi-step Social Engineering Attacks on the WebabstractBy exploiting people's psychological vulnerabilities, modern web-based social engineering (SE) attacks manipulate victims to download malware and expose personal information. To effectively lure users, some SE attacks constitute a sequence of web pages starting from a landing page and require browser interactions at each web page, which we call multi-step SE attacks. Also, different browser interactions executed on a web page often branch to multiple sequences to redirect users to different SE attacks. Although common systems analyze only landing pages or conduct browser interactions limited to a specific attack, little effort has been made to follow such sequences of web pages to collect multi-step SE attacks. Takashi Koide, Daiki Chiba 0001, Mitsuaki Akiyama |
AsiaCCS | 1 |
| 2020 | It Never Rains but It Pours: Analyzing and Detecting Fake Removal Information Advertisement Sites
Takashi Koide, Daiki Chiba 0001, Mitsuaki Akiyama, Katsunari Yoshioka, Tsutomu Matsumoto |
DIMVA | 1 |
| 2019 | DomainScouter: Understanding the Risks of Deceptive IDNs
Daiki Chiba 0001, Ayako Akiyama Hasegawa, Takashi Koide, Yuta Sawabe, Shigeki Goto, Mitsuaki Akiyama |
RAID | 3 |
| 2015 | AmpPot: Monitoring and Defending Against Amplification DDoS Attacks
Lukas Krämer, Johannes Krupp, Daisuke Makita, Tomomi Nishizoe, Takashi Koide, Katsunari Yoshioka, Christian Rossow |
RAID | 5 |