VLDB 2026 Research / reviewers in the wild / expert
Daiki Chiba 0001
dblp:53/9778
· DBLP profile ↗
34ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0002-7532-6633ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 14 · 4 first-author · 8 since 2021Computer networks · 10 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PhishLumos: An Adaptive Multi-Agent System for Proactive Phishing Campaign Mitigation
Daiki Chiba 0001, Hiroki Nakano, Takashi Koide |
ICC | 1 |
| 2025 | DomainLynx: Leveraging Large Language Models for Enhanced Domain Squatting DetectionabstractDomain squatting poses a significant threat to Internet security, with attackers employing increasingly sophisticated techniques. This study introduces DomainLynx, an innovative compound AI system leveraging Large Language Models (LLMs) for enhanced domain squatting detection. Unlike existing methods focusing on predefined patterns for top-ranked domains, DomainLynx excels in identifying novel squatting techniques and protecting less prominent brands. The system's architecture integrates advanced data processing, intelligent domain pairing, and LLM-powered threat assessment. Crucially, DomainLynx incorporates specialized components that mitigate LLM hallucinations, ensuring reliable and context-aware detection. This approach enables efficient analysis of vast security data from diverse sources, including Certificate Transparency logs, Passive DNS records, and zone files. Evaluated on a curated dataset of 1,649 squatting domains, DomainLynx achieved 94.7% accuracy using Llama-3-70B. In a month-long real-world test, it detected 34,359 squatting domains from 2.09 million new domains, outperforming baseline methods by 2.5 times. This research advances Internet security by providing a versatile, accurate, and adaptable tool for combating evolving domain squatting threats. DomainLynx's approach paves the way for more robust, AI-driven cybersecurity solutions, enhancing protection for a broader range of online entities and contributing to a safer digital ecosystem. Daiki Chiba 0001, Hiroki Nakano, Takashi Koide |
CCNC | 1 |
| 2025 | DomainDynamics: Lifecycle-Aware Risk Timeline Construction for Domain NamesabstractThe persistent threat posed by malicious domain names in cyber-attacks underscores the urgent need for effective detection mechanisms. Traditional machine learning methods, while capable of identifying such domains, often suffer from high false positive and false negative rates due to their extensive reliance on historical data. Conventional approaches often overlook the dynamic nature of domain names, the purposes and ownership of which may evolve, potentially rendering risk assessments outdated or irrelevant. To address these shortcomings, we introduce DomainDynamics, a novel system designed to predict domain name risks by considering their lifecycle stages. DomainDynamics constructs a timeline for each domain, evaluating the characteristics of each domain at various points in time to make informed, temporal risk determinations. In an evaluation experiment involving over 85,000 actual malicious domains from malware and phishing incidents, DomainDynamics demonstrated a significant improvement in detection rates, achieving an 82.58% detection rate with a low false positive rate of 0.41%. This performance surpasses that of previous studies and commercial services, improving detection capability substantially. Daiki Chiba 0001, Hiroki Nakano, Takashi Koide |
CCNC | 1 |
| 2025 | DomainHarvester: Harvesting Infrequently Visited Yet Trustworthy Domain NamesabstractIn cybersecurity, allow lists play a crucial role in distinguishing safe websites from potential threats. Conventional methods for compiling allow lists, focusing heavily on website popularity, often overlook infrequently visited legitimate domains. This paper introduces DomainHarvester, a system aimed at generating allow lists that include trustworthy yet infrequently visited domains. By adopting an innovative bottom-up methodology that leverages the web's hyperlink structure, DomainHarvester identifies legitimate yet underrepresented domains. The system uses seed URLs to gather domain names, employing machine learning with a Transformer-based approach to assess their trustworthiness. DomainHarvester has developed two distinct allow lists: one with a global focus and another emphasizing local relevance. Compared to six existing top lists, DomainHarvester's allow lists show minimal overlaps, 4% globally and 0.1 % locally, while significantly reducing the risk of including malicious domains, thereby enhancing security. The contributions of this research are substantial, illuminating the overlooked aspect of trustworthy yet underrepresented domains and introducing DomainHarvester, a system that goes beyond traditional popularity-based metrics. Our methodology enhances the inclusivity and precision of allow lists, offering significant advantages to users and businesses worldwide, especially in non-English speaking regions. Daiki Chiba 0001, Hiroki Nakano, Takashi Koide |
CCNC | 1 |
| 2025 | ScamFerret: Detecting Scam Websites Autonomously with Large Language Models
Hiroki Nakano, Takashi Koide, Daiki Chiba 0001 |
DIMVA (1) | 3 |
| 2025 | PhishParrot: LLM-Driven Adaptive Crawling to Unveil Cloaked Phishing SitesabstractPhishing attacks continue to evolve, with cloaking techniques posing a significant challenge to detection efforts. Cloaking allows attackers to display phishing sites only to specific users while presenting legitimate pages to security crawlers, rendering traditional detection systems ineffective. This research proposes PhishParrot, a novel crawling environment optimization system designed to counter cloaking techniques. PhishParrot leverages the contextual analysis capabilities of Large Language Models (LLMs) to identify potential patterns in crawling information, enabling the construction of optimal user profiles capable of bypassing cloaking mechanisms. The system accumulates information on phishing sites collected from diverse environments. It then adapts browser settings and network configurations to match the attacker’s target user conditions based on information extracted from similar cases. A 21-day evaluation showed that PhishParrot improved detection accuracy by up to 33.8% over standard analysis systems, yielding 91 distinct crawling environments for diverse conditions targeted by attackers. The findings confirm that the combination of similar-case extraction and LLM-based context analysis is an effective approach for detecting cloaked phishing attacks. Hiroki Nakano, Takashi Koide, Daiki Chiba 0001 |
GLOBECOM | 3 |
| 2025 | Uncovering Suspicious Posts on Abandoned BlogsabstractAs the Internet evolves, the diversification of information dissemination has led to the growth of various social media platforms and an increase in abandoned blogs. These abandoned blogs, if not properly managed, present a significant security risk. In this study, we introduce “Zombified Blogs,” defined as blogs deserted by their administrators and targeted for continuous submissions of unsolicited malicious content. We explore the causes and effects of this global phenomenon on a large scale, using data from multiple blog services over the past 18 years for the first time. To this end, we propose an innovative method for the continuous identification and collection of Zombified Blogs using a search engine. We analyzed data spanning 18 years, encompassing approximately 1.5 million blog posts, including 258, 547 malicious posts, across 1,141 Zombified Blogs on six different blog services. Our findings show a significant rise in malicious posts, primarily social engineering attacks, over the past four years on previously active blogs. We identify the email posting function as a critical area exploited by attackers, leading to the unintended posting of content. Finally, we provide an overview of the current challenges facing blog services and administrators, and discuss potential content management solutions. Hiroki Nakano, Takashi Koide, Daiki Chiba 0001, Katsunari Yoshioka, Tsutomu Matsumoto |
ICC | 3 |
| 2025 | DomainDynamics: Advancing lifecycle-based risk assessment of domain namesabstractThe persistent threat of malicious domains in cybersecurity necessitates robust detection systems. Traditional machine learning approaches often struggle to accurately assess domain name risks due to their static analysis methods and lack of consideration for temporal changes in domain attributes. To address these limitations, we developed DomainDynamics, a novel system that evaluates domain name risks by analyzing their lifecycle phases. This study provides a comprehensive evaluation and refinement of the DomainDynamics framework. The system creates temporal profiles for domains and assesses their attributes at various stages, enabling informed, time-sensitive risk assessments. Our initial evaluation, involving over 85,000 malicious domains, achieved an 82.58% detection rate with a low 0.41% false positive rate. We expanded our research to include benchmarking against commercial services, feature significance analysis using interpretable AI techniques, and detailed case studies. This investigation not only validates the effectiveness of DomainDynamics but also reveals temporal indicators of malicious intent. Our findings demonstrate the advantages of lifecycle-based analysis over static methodologies, providing valuable insights for practical cybersecurity applications. • DomainDynamics: lifecycle-based detection (82.58% Detection Rate, 0.41% FPR). • Temporal analysis outperforms static methods on 85,000+ malicious domains. • Explainable AI offers deep insights into domain risk through real case studies. • Benchmarks against commercial tools confirm lifecycle indicators’ significance. • Enhances malicious domain detection via temporal context and dynamic changes. Daiki Chiba 0001, Hiroki Nakano, Takashi Koide |
Comput. Secur. | 1 |
| 2024 | Noisy Label Detection for Multi-labeled MalwareabstractMalware attacks have become increasingly prevalent, and accurate and reliable malware detection is essential for combating them. However, mislabeling, where data is given a different/noisy label than its true label, can significantly affect the accuracy and reliability of malware detection. In this paper, we propose a new method for detecting noisy labels in multi-labeled malware datasets. Our approach involves a new transformation method that allows malware datasets with multiple labels to be treated as data with a single label without losing any essential information. We also introduce a new machine learning model for detecting mislabeling, based on this transformation method. We conducted experiments on a real-world malware dataset to evaluate the effectiveness of our proposed method, and our findings indicate that our method can detect mislabels with high accuracy, up to 94.7%. Our research aims to improve the quality of labeling and reduce the factors contributing to mislabeling in malware datasets, leading to more accurate and reliable malware detection. Naoki Fukushi, Toshiki Shibahara, Hiroki Nakano, Takashi Koide, Daiki Chiba 0001 |
CCNC | 5 |
| 2024 | ChatSpamDetector: Leveraging Large Language Models for Effective Phishing Email Detection
Takashi Koide, Naoki Fukushi, Hiroki Nakano, Daiki Chiba 0001 |
SecureComm (3) | 4 |
| 2023 | Canary in Twitter Mine: Collecting Phishing Reports from Experts and Non-expertsabstractThe rise in phishing attacks via e-mail and short message service (SMS) has not slowed down at all. The first thing we need to do to combat the ever-increasing number of phishing attacks is to collect and characterize more phishing cases that reach end users. Without understanding these characteristics, anti-phishing countermeasures cannot evolve. In this study, we propose an approach using Twitter as a new observation point to immediately collect and characterize phishing cases via e-mail and SMS that evade countermeasures and reach users. Specifically, we propose CrowdCanary, a system capable of structurally and accurately extracting phishing information (e.g., URLs and domains) from tweets about phishing by users who have actually discovered or encountered it. In our three months of live operation, CrowdCanary identified 35,432 phishing URLs out of 38,935 phishing reports, 31,960 (90.2%) of these phishing URLs were later detected by the anti-virus engine. We analyzed users who shared phishing threats by categorizing them into two groups: experts and non-experts. As a results, we discovered that CrowdCanary extracts non-expert report-specific information, like company brand name in tweets, phishing attack details from tweet images, and pre-redirect landing page information. Hiroki Nakano, Daiki Chiba 0001, Takashi Koide, Naoki Fukushi, Takeshi Yagi, Takeo Hariu, Katsunari Yoshioka, Tsutomu Matsumoto |
ARES | 2 |
| 2023 | PhishReplicant: A Language Model-based Approach to Detect Generated Squatting Domain NamesabstractDomain squatting is a technique used by attackers to create domain names for phishing sites. In recent phishing attempts, we have observed many domain names that use multiple techniques to evade existing methods for domain squatting. These domain names, which we call generated squatting domains (GSDs), are quite different in appearance from legitimate domain names and do not contain brand names, making them difficult to associate with phishing. In this paper, we propose a system called PhishReplicant that detects GSDs by focusing on the linguistic similarity of domain names. We analyzed newly registered and observed domain names extracted from certificate transparency logs, passive DNS, and DNS zone files. We detected 3,498 domain names acquired by attackers in a four-week experiment, of which 2,821 were used for phishing sites within a month of detection. We also confirmed that our proposed system outperformed existing systems in both detection accuracy and number of domain names detected. As an in-depth analysis, we examined 205k GSDs collected over 150 days and found that phishing using GSDs was distributed globally. However, attackers intensively targeted brands in specific regions and industries. By analyzing GSDs in real time, we can block phishing sites before or immediately after they appear. Takashi Koide, Naoki Fukushi, Hiroki Nakano, Daiki Chiba 0001 |
ACSAC | 4 |
| 2023 | A First Look at Brand Indicators for Message Identification (BIMI)abstractAbstract As promising approaches to thwarting the damage caused by phishing emails, DNS-based email security mechanisms, such as the Sender Policy Framework (SPF), Domain-based Message Authentication, Reporting & Conformance (DMARC) and DNS-based Authentication of Named Entities (DANE), have been proposed and widely adopted. Nevertheless, the number of victims of phishing emails continues to increase, suggesting that there should be a mechanism for supporting end-users in correctly distinguishing such emails from legitimate emails. To address this problem, the standardization of Brand Indicators for Message Identification (BIMI) is underway. BIMI is a mechanism that helps an email recipient visually distinguish between legitimate and phishing emails. With Google officially supporting BIMI in July 2021, the approach shows signs of spreading worldwide. With these backgrounds, we conduct an extensive measurement of the adoption of BIMI and its configuration. The results of our measurement study revealed that, as of November 2022, 3,538 out of the one million most popular domain names have a set BIMI record, whereas only 396 (11%) of the BIMI-enabled domain names had valid logo images and verified mark certificates. The study also revealed the existence of several misconfigurations in such logo images and certificates. Masanori Yajima, Daiki Chiba 0001, Yoshiro Yoneya, Tatsuya Mori 0003 |
PAM | 2 |
| 2022 | Objection!: Identifying Misclassified Malicious Activities with XAIabstractMany studies have been conducted to detect various malicious activities in cyberspace using classifiers built by machine learning. However, it is natural for any classifier to make mistakes, and hence, human verification is necessary. One method to address this issue is eXplainable AI (XAI), which provides a reason for the classification result. However, when the number of classification results to be verified is large, it is not realistic to check the output of the XAI for all cases. In addition, it is sometimes difficult to interpret the output of XAI. In this study, we propose a machine learning model called classification verifier that verifies the classification results by using the output of XAI as a feature and raises objections when there is doubt about the reliability of the classification results. The results of experiments on malicious website detection and malware detection show that the proposed classification verifier can efficiently identify misclassified malicious activities. Koji Fujita, Toshiki Shibahara, Daiki Chiba 0001, Mitsuaki Akiyama, Masato Uchida |
ICC | 3 |
| 2021 | Detecting Event-synced Navigation Attacks across User-generated Content PlatformsabstractWith the spread of service platforms that enable users to generate content, people use user-generated content (UGC) to search for and access information on the web instead of search engines. Attackers can also use UGC on a service platform (UGC platform) to spread web-based social engineering (SE) attacks to a large number of people. In this paper, we focus on a type of web-based SE attack, called an event-synced navigation attack, which generates UGC with links navigating users to malicious websites and distribute it synced with a real-life event at a specific time. To understand the attacks in the wild, we propose a new system for detecting event-synced navigation attacks in real time by capturing the inevitable footprints left by attacks that affect a large number of users. We evaluate each of the three steps of the proposed system and finally find that the system can classify malicious and non-malicious UGC with 97% accuracy. Furthermore, we perform a comprehensive measurement study on event-synced navigation attacks spread from popular UGC platforms (Twitter, Facebook, YouTube, and Reddit) and confirm that many event-synced navigation attacks are deployed in the wild. Hiroki Nakano, Daiki Chiba 0001, Takashi Koide, Mitsuaki Akiyama |
COMPSAC | 2 |
| 2021 | Measuring Adoption of DNS Security Mechanisms with Cross-Sectional ApproachabstractThe threat of attacks targeting a DNS, such as DNS cache poisoning attacks and DNS amplification attacks, continues unabated. In addition, attacks that exploit the difficulty in deter-mining the authenticity of domain names, such as phishing sites and fraudulent emails, continue to be a significant threat. Various DNS security mechanisms have been proposed, standardized, and implemented as effective countermeasures against DNS-related attacks. However, it is not clear how widespread these security mechanisms are in the DNS ecosystem and how effectively they work in the wild. With this background, this study targets the major DNS security mechanisms deployed for the DNS name servers, DNSSEC, DNS Cookies, CAA, SPF, DMARC, MTA-STS, DANE, and TLSRPT, and a large-scale measurement analysis of their deployment is conducted. Our results quantitatively reveal that, as of 2021, the adoption rate of most DNS security mechanisms, except SPF, remains low, and the adoption rate is lower for mechanisms that are more difficult to configure. These findings suggest the importance of developing easy-to-deploy tools to promote the adoption of security mechanisms. Masanori Yajima, Daiki Chiba 0001, Yoshiro Yoneya, Tatsuya Mori 0003 |
GLOBECOM | 2 |
| 2021 | Auto-creation of Android Malware Family TreeabstractAndroid malware has been a growing threat. For an effective countermeasure against Android malware, we need to not only detect the malware at a certain point in time but also analyze its time-series changes of malware, taking into account that the family of Android malware will increase in number over time. In this paper, we propose a new method for automatically creating a "family tree" of Android malware that can represent how the newly detected Android malware is related to existing Android malware and its families, and how they have changed over time. Our evaluation using 24,474 actual Android malware APKs shows that our proposed family tree is able to accurately represent time-series changes between malware families. Kazuya Nomura, Daiki Chiba 0001, Mitsuaki Akiyama, Masato Uchida |
ICC | 2 |
| 2021 | A First Look at COVID-19 Domain Names: Origin and Implications
Ryo Kawaoka, Daiki Chiba 0001, Takuya Watanabe 0001, Mitsuaki Akiyama, Tatsuya Mori 0003 |
PAM | 2 |
| 2021 | Analyzing Security Risks of Ad-Based URL Shortening Services Caused by Users' Behaviors
Naoki Fukushi, Takashi Koide, Daiki Chiba 0001, Hiroki Nakano, Mitsuaki Akiyama |
SecureComm (2) | 3 |
| 2020 | Detecting Malware-infected Hosts Using Templates of Multiple HTTP RequestsabstractIn this paper, we propose a method for detecting malware-infected hosts with a high rate of detection and a low rate of false positives without using any data on benign communication. Based on the fact that many malware-infected hosts generate multiple HTTP requests, we propose a method using the templates of sets of those HTTP requests. For each malware, this method generates a template that comprises the set of templates of the HTTP requests that the malware generates. We call the set of templates group template. It then detects malware-infected hosts by comparing the set of monitored HTTP requests with the group templates. Taiga Hokaguchi, Yuichi Ohsita, Toshiki Shibahara, Daiki Chiba 0001, Mitsuaki Akiyama, Masayuki Murata 0001 |
CCNC | 4 |
| 2020 | To Get Lost is to Learn the Way: Automatically Collecting Multi-step Social Engineering Attacks on the WebabstractBy exploiting people's psychological vulnerabilities, modern web-based social engineering (SE) attacks manipulate victims to download malware and expose personal information. To effectively lure users, some SE attacks constitute a sequence of web pages starting from a landing page and require browser interactions at each web page, which we call multi-step SE attacks. Also, different browser interactions executed on a web page often branch to multiple sequences to redirect users to different SE attacks. Although common systems analyze only landing pages or conduct browser interactions limited to a specific attack, little effort has been made to follow such sequences of web pages to collect multi-step SE attacks. Takashi Koide, Daiki Chiba 0001, Mitsuaki Akiyama |
AsiaCCS | 2 |
| 2020 | It Never Rains but It Pours: Analyzing and Detecting Fake Removal Information Advertisement Sites
Takashi Koide, Daiki Chiba 0001, Mitsuaki Akiyama, Katsunari Yoshioka, Tsutomu Matsumoto |
DIMVA | 2 |
| 2020 | Time-series Measurement of Parked Domain NamesabstractDomain parking is a monetization mechanism for displaying online advertisements in unused domain names. Some domain names used in cyber attacks are known to leverage domain parking services after the attack. However, the temporal relationships between domain parking services and malicious domain names have not been studied well. In this study, we investigated how malicious domain names using domain parking services change over time. We conducted a large-scale measurement study of more than 66.8 million domain names that have used domain parking services in the past 19 months. We reveal the existence of 3,964 domain names that have been malicious after using domain parking. We also reveal the existence of 3.02 million domain names that utilized multiple parking services simultaneously or while switching between them. Our study can contribute to the efficient analysis of malicious domain names using domain parking services. Takayuki Tomatsuri, Daiki Chiba 0001, Mitsuaki Akiyama, Masato Uchida |
GLOBECOM | 2 |
| 2019 | Exploration into Gray Area: Efficient Labeling for Malicious Domain Name DetectionabstractThis paper presents a method to reduce the labeling cost when acquiring training data for a system that detects malicious domain names by supervised machine learning. The conventional system requires large quantities of both benign and malicious domain names to be prepared as training data to obtain a classifier with high classification accuracy. In general, malicious domain names are observed less frequently than benign domain names. Therefore, it is difficult to acquire a large number of malicious domain names without a dedicated labeling method. We propose a method based on active learning that labels data around the decision boundary of classification, i.e., in the gray area, and we show that the classification accuracy can be improved by only using approximately 2.5% of the training data used by the conventional system. An additional disadvantage of the conventional system is that, if the classifier is trained with a small amount of training data, its generalization ability cannot be guaranteed. We propose a method based on ensemble learning that integrates multiple classifiers, and we show that the classification accuracy can be stabilized and improved. Naoki Fukushi, Daiki Chiba 0001, Mitsuaki Akiyama, Masato Uchida |
COMPSAC (1) | 2 |
| 2019 | Precise and Robust Detection of Advertising FraudabstractAs the online advertising market has grown, advertising frauds (ad frauds) have become a serious problem. Countermeasures against ad frauds are evaded since they rely on noticeable features (e.g., burstiness of ad requests) that attackers can easily change. We propose an ad-fraud-detection method that leverages robust features against attacker evasion. We designed novel features on the basis of the statistics observed in an ad network calculated from a large amount of ad requests from legitimate users, such as the popularity of publisher websites and the tendencies of client environments. We assume that attackers cannot know of or manipulate these statistics and that features extracted from fraudulent ad requests tend to be outliers. These features are used to construct a machine-learning model for detecting fraudulent ad requests. We evaluated our proposed method by using ad-request logs observed within an actual ad network. The results revealed that our designed features improved the recall rate by 10% and had about 100,000 - 160,000 fewer false negatives per day than conventional features based on the burstiness of ad requests. In addition, by evaluating detection performance with long-term dataset, we confirmed that the proposed method is robust against performance degradation over time. Finally, we applied our proposed method to a large dataset constructed on an ad network and found several characteristics of the latest ad frauds in the wild, for example, a large amount of fraudulent ad requests is sent from cloud servers. Fumihiro Kanei, Daiki Chiba 0001, Kunio Hato, Mitsuaki Akiyama |
COMPSAC (1) | 2 |
| 2019 | ShamFinder: An Automated Framework for Detecting IDN HomographsabstractThe internationalized domain name (IDN) is a mechanism that enables us to use Unicode characters in domain names. The set of Unicode characters contains several pairs of characters that are visually identical with each other; e.g., the Latin character 'a' (U+0061) and Cyrillic character 'a' (U+0430). Visually identical characters such as these are generally known as homoglyphs. IDN homograph attacks, which are widely known, abuse Unicode homoglyphs to create lookalike URLs. Although the threat posed by IDN homograph attacks is not new, the recent rise of IDN adoption in both domain name registries and web browsers has resulted in the threat of these attacks becoming increasingly widespread, leading to large-scale phishing attacks such as those targeting cryptocurrency exchange companies. In this work, we developed a framework named "ShamFinder," which is an automated scheme to detect IDN homographs. Our key contribution is the automatic construction of a homoglyph database, which can be used for direct countermeasures against the attack and to inform users about the context of an IDN homograph. Using the ShamFinder framework, we perform a large-scale measurement study that aims to understand the IDN homographs that exist in the wild. On the basis of our approach, we provide insights into an effective countermeasure against the threats caused by the IDN homograph attack. Hiroaki Suzuki, Daiki Chiba 0001, Yoshiro Yoneya, Tatsuya Mori 0003, Shigeki Goto |
Internet Measurement Conference | 2 |
| 2019 | DomainScouter: Understanding the Risks of Deceptive IDNs
Daiki Chiba 0001, Ayako Akiyama Hasegawa, Takashi Koide, Yuta Sawabe, Shigeki Goto, Mitsuaki Akiyama |
RAID | 1 |
| 2018 | Don't throw me away: Threats Caused by the Abandoned Internet Resources Used by Android AppsabstractThis study aims to understand the threats caused by abandoned Internet resources used by Android apps. By abandoned, we mean Internet resources that support apps that were published and are still available on the mobile app marketplace, but have not been maintained and hence are at risk for abuse by an outsider. Internet resources include domain names and hard-coded IP addresses, which could be used for nefarious purposes, e.g., stealing sensitive private information, scamming and phishing, click fraud, and injecting malware distribution URL. As a result of the analysis of 1.1 M Android apps published in the official marketplace, we uncovered 3,628 of abandoned Internet resources associated with 7,331 available mobile apps. These resources are subject to hijack by outsiders. Of these apps, 13 apps have been installed more than a million of times, a measure of the breadth of the threat. Based on the findings of empirical experiments, we discuss potential threats caused by abandoned Internet resources and propose countermeasures against these threats. Elkana Pariwono, Daiki Chiba 0001, Mitsuaki Akiyama, Tatsuya Mori 0003 |
AsiaCCS | 2 |
| 2018 | DomainChroma: Building actionable threat intelligence from malicious domain namesabstractSince the 1980s, domain names and the domain name system (DNS) have been used and abused. Although legitimate Internet users rely on domain names as indispensable infrastructures for using the Internet, attackers use or abuse them as reliable, instantaneous, and distributed attack infrastructures. However, there is a lack of complete understanding of such domain-name abuses and methods for coping with them. In this study, we designed and implemented a unified analysis system combining current defense solutions to build actionable threat intelligence from malicious domain names. The basic concept underlying our system is malicious domain name chromatography. Our analysis system can distinguish among mixtures of malicious domain names for websites. On the basis of this concept, we do not create a hodgepodge of current solutions but design separation of abused domain names and offer actionable threat intelligence or defense information by considering the characteristics of malicious domain names as well as the possible defense solutions and points of defense. Finally, we evaluated our analysis system and defense-information output using a large real dataset to show the effectiveness and validity of our system. Daiki Chiba 0001, Mitsuaki Akiyama, Takeshi Yagi, Kunio Hato, Tatsuya Mori 0003, Shigeki Goto |
Comput. Secur. | 1 |
| 2017 | DomainChroma: Providing Optimal Countermeasures against Malicious Domain NamesabstractDomain names and domain name system (DNS) have been used and abused for over 30 years since the 1980s. Although legitimate Internet users rely on domain names as their indispensable infrastructures for using the Internet, attackers use or abuse them as reliable, instantaneous, and distributed attack infrastructure. However, there is a lack of complete understanding of such domain name abuses and the methods for coping with them. In this paper, we design and implement a unified and objective analysis pipeline combining the existing defense solutions to realize practical and optimal defenses against today's malicious domain names. The basic concept underlying our novel analytical approach is malicious domain names' chromatography. Our new analysis pipeline can distinguish among mixtures of malicious domain names for websites. On the basis of this concept, we do not create a hodgepodge of existing solutions but design separation of abused domain names and offer defense information by considering the characteristics of malicious domain names as well as the possible defense solutions and points of defense. Finally, we evaluate our analysis pipeline and output defense information using a large and real dataset to show the effectiveness and validity of our proposed approach. Daiki Chiba 0001, Mitsuaki Akiyama, Takeshi Yagi, Takeshi Yada, Tatsuya Mori 0003, Shigeki Goto |
COMPSAC (1) | 1 |
| 2017 | Malicious URL sequence detection using event de-noising convolutional neural networkabstractAttackers have increased the number of infected hosts by redirecting users of compromised popular websites toward websites that exploit vulnerabilities of a browser and its plugins. To prevent damage, detecting infected hosts based on proxy logs, which are generally recorded on enterprise networks, is gaining attention rather than blacklist-based filtering because creating blacklists has become difficult due to the short lifetime of malicious domains and concealment of exploit code. Since information extracted from one URL is limited, we focus on a sequence of URLs that includes artifacts of malicious redirections. We propose a system for detecting malicious URL sequences from proxy logs with a low false positive rate. To elucidate an effective approach of malicious URL sequence detection, we compared three approaches: individual-based approach, convolutional neural network (CNN), and our newly developed event de-noising CNN (EDCNN). Our EDCNN is a new CNN to reduce the negative effect of benign URLs redirected from compromised websites included in malicious URL sequences. Our evaluation shows that the EDCNN lowers the operation cost of malware infection by reducing 47% of false alerts compared with a CNN when users access compromised websites but do not obtain exploit code due to browser fingerprinting. Toshiki Shibahara, Kohei Yamanishi, Yuta Takata, Daiki Chiba 0001, Mitsuaki Akiyama, Takeshi Yagi, Yuichi Ohsita, Masayuki Murata 0001 |
ICC | 4 |
| 2016 | DomainProfiler: Discovering Domain Names Abused in FutureabstractCyber attackers abuse the domain name system (DNS) to mystify their attack ecosystems, they systematically generate a huge volume of distinct domain names to make it infeasible for blacklisting approaches to keep up with newly generated malicious domain names. As a solution to this problem, we propose a system for discovering malicious domain names that will likely be abused in future. The key idea with our system is to exploit temporal variation patterns (TVPs) of domain names. The TVPs of domain names include information about how and when a domain name has been listed in legitimate/popular and/or malicious domain name lists. On the basis of this idea, our system actively collects DNS logs, analyzes their TVPs, and predicts whether a given domain name will be used for malicious purposes. Our evaluation revealed that our system can predict malicious domain names 220 days beforehand with a true positive rate of 0.985. Daiki Chiba 0001, Takeshi Yagi, Mitsuaki Akiyama, Toshiki Shibahara, Takeshi Yada, Tatsuya Mori 0003, Shigeki Goto |
DSN | 1 |
| 2016 | Efficient Dynamic Malware Analysis Based on Network Behavior Using Deep LearningabstractMalware authors or attackers always try to evade detection methods to accomplish their mission. Such detection methods are broadly divided into three types: static feature, host-behavior, and network-behavior based. Static feature-based methods are evaded using packing techniques. Host- behavior-based methods also can be evaded using some code injection methods, such as API hook and dynamic link library hook. This arms race regarding static feature-based and host-behavior- based methods increases the importance of network-behavior-based methods. The necessity of communication between infected hosts and attackers makes it difficult to evade network-behavior- based methods. The effectiveness of such methods depends on how we collect a variety of communications by using malware samples. However, analyzing all new malware samples for a long period is infeasible. Therefore, we propose a method for determining whether dynamic analysis should be suspended based on network behavior to collect malware communications efficiently and exhaustively. The key idea behind our proposed method is focused on two characteristics of malware communication: the change in the communication purpose and the common latent function. These characteristics of malware communications resemble those of natural language from the viewpoint of data structure, and sophisticated analysis methods have been proposed in the field of natural language processing. For this reason, we applied the recursive neural network, which has recently exhibited high classification performance, to our proposed method. In the evaluation with 29,562 malware samples, our proposed method reduced 67.1% of analysis time while keeping the coverage of collected URLs to 97.9% of the method that continues full analyses. Toshiki Shibahara, Takeshi Yagi, Mitsuaki Akiyama, Daiki Chiba 0001, Takeshi Yada |
GLOBECOM | 4 |
| 2016 | Detection of vulnerability scanning using features of collective accesses based on information collected from multiple honeypotsabstractAttacks against websites are increasing rapidly with the expansion of web services. An increasing number of diversified web services make it difficult to prevent such attacks due to many known vulnerabilities in websites. To overcome this problem, it is necessary to collect the most recent attacks using decoy web honeypots and to implement countermeasures against malicious threats. Web honeypots collect not only malicious accesses by attackers but also benign accesses such as those by web search crawlers. Thus, it is essential to develop a means of automatically identifying malicious accesses from mixed collected data including both malicious and benign accesses. Specifically, detecting vulnerability scanning, which is a preliminary process, is important for preventing attacks. In this study, we focused on classification of accesses for web crawling and vulnerability scanning since these accesses are too similar to be identified. We propose a feature vector including features of collective accesses, e.g., intervals of request arrivals and the dispersion of source port numbers, obtained with multiple honeypots deployed in different networks for classification. Through evaluation using data collected from 37 honeypots in a real network, we show that features of collective accesses are advantageous for vulnerability scanning and crawler classification. Naomi Kuze, Shu Ishikura, Takeshi Yagi, Daiki Chiba 0001, Masayuki Murata 0001 |
NOMS | 4 |