Weifeng Li 0002

dblp:97/69-2 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0002-2105-3596ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 10 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 An interpretable wide and deep model for online disinformation detection
Yidong Chai, Weifeng Li 0002, Bin Zhu 0007, Hongyan Liu 0002, Yuan-Chun Jiang
Expert Syst. Appl.3
2023 Evading Deep Learning-Based Malware Detectors via Obfuscation: A Deep Reinforcement Learning Approach
abstract
Adversarial Malware Generation (AMG), the generation of adversarial malware variants to strengthen Deep Learning (DL)-based malware detectors has emerged as a crucial tool in the development of proactive cyberdefense. However, the majority of extant works offer subtle perturbations or additions to executable files and do not explore full-file obfuscation. In this study, we show that an open-source encryption tool coupled with a Reinforcement Learning (RL) framework can successfully obfuscate malware to evade state-of-the-art malware detection engines and outperform techniques that use advanced modification methods. Our results show that the proposed method improves the evasion rate from 27%-49% compared to widely-used state-of-the-art reinforcement learning-based methods.
Brian Etter, James Lee Hu, Reza Ebrahimi 0001, Weifeng Li 0002, Xin Li 0108, Hsinchun Chen
ICDM4
2022 An Explainable Multi-Modal Hierarchical Attention Model for Developing Phishing Threat Intelligence
abstract
Phishing website attack, as one of the most persistent forms of cyber threats, evolves and remains a major cyber threat. Various detection methods (e.g., lookup systems, fraud cue-based methods) have been proposed to identify phishing websites. The limitations of lookup systems (e.g., failing to address newly created attacks) and the fraud cue-based methods (e.g., relying on feature engineering) motivated the development of deep representation-based methods capable of learning deep fraud cues for enhanced anti-phishing capacity. Focusing mostly on URLs, these methods fail to analyze other two important modalities of website content: textual information and visual design. Moreover, the interpretability of these deep learning based methods is limited, reducing model trustworthiness and preventing relevant and actionable intelligence. We propose a multi-modal hierarchical attention model (MMHAM) which jointly learns the deep fraud cues from the three major modalities of website content for phishing website detection. Specifically, MMHAM features an innovative shared dictionary learning approach for aligning representations from different modalities in the attention mechanism. In evaluation experiments, the proposed MMHAM not only learned improved deep cues for enhanced phishing detection, but provided a hierarchical interpretability system from which we could develop phishing threat intelligence to inform phishing websites detection at different levels.
Yidong Chai, Yonghang Zhou, Weifeng Li 0002, Yuan-Chun Jiang
IEEE Trans. Dependable Secur. Comput.3
2021 Exploring Differences Among Darknet and Surface Internet Hacking Communities
abstract
Cyber-threat intelligence (CTI) has matured into its own industry within recent years. CTI efforts frequently involve scrutinizing data within Darknet communities to understand emerging threats. Many hackers within the Darknet share knowledge and other information through a variety of formats, including video. At the same time, many hackers are also making use of the “surface” Internet and traditional video-sharing platforms to disseminate hacking knowledge. Gleaning intelligence from the Darknet can be a very laborious and costly task, raising the question of how meaningful and valuable are the hacker patterns that can be observed on the surface Internet. Extant research contains no studies that compare and contrast hacking videos uploaded to the Darknet versus those uploaded to traditional Internet communities. In this research-in-progress, a testbed of hacking videos is constructed by sourcing videos from a popular video-sharing website, as well as several Darknet forums. The testbed is scrutinized to understand differences in how the populations of users watching such videos respond to them, and whether there are any unique engagement patterns that emerge within the Darknet and surface Internet populations. The results of this work serve to justify further investigations into the hacker knowledge gap between the Darknet and the traditional Internet.
Zhiyuan Ding, Victor A. Benjamin, Weifeng Li 0002, Xueyan Yin
ISI3
2021 Automated PII Extraction from Social Media for Raising Privacy Awareness: A Deep Transfer Learning Approach
abstract
Internet users have been exposing an increasing amount of Personally Identifiable Information (PII) on social media. Such exposed PII can be exploited by cybercriminals and cause severe losses to the users. Informing users of their PII exposure in social media is crucial to raise their privacy awareness and encourage them to take protective measures. To this end, advanced techniques are needed to extract users’ exposed PII in social media automatically, whereas most existing studies remain manual. While Information Extraction (IE) techniques can be used to extract the PII automatically, Deep Learning (DL)-based IE models alleviate the need for feature engineering and further improve the efficiency. However, DL-based IE models often require large-scale labeled data for training, but PII-labeled social media posts are difficult to obtain due to privacy concerns. Also, these models rely heavily on pre-trained word embeddings, while PII in social media often varies in forms and thus has no fixed representations in pre-trained word embeddings. In this study, we propose the Deep Transfer Learning for PII Extraction (DTL-PIIE) framework to address these two limitations. DTL-PIIE transfers knowledge learned from publicly available PII data to social media in order to address the problem of rare PII-labeled data. Moreover, our framework leverages Graph Convolutional Networks (GCNs) to incorporate syntactic patterns to guide PIIE without relying on pre-trained word embeddings. Evaluation against benchmark IE models indicates that our approach outperforms state-of-the-art DL-based IE models. An ablation analysis further confirms the efficacy of each component in our model. Our proposed framework can facilitate various applications, such as PII misuse prediction and privacy risk assessment, thereby protecting the privacy of internet users.
Fang Yu Lin, Reza Ebrahimi 0001, Weifeng Li 0002, Hsinchun Chen
ISI4
2020 Detecting Cyber-Adversarial Videos in Traditional Social media
abstract
Cyber-threat intelligence (CTI) has matured and grown into its own industry within recent years. Many CTI efforts involve scrutinizing text-based conversations in DarkNet forums and markets. However, hackers commonly share knowledge and other information through video formats that have been largely ignored. Further, cybercriminals are increasingly making use of mainstream social media to transmit hacking knowledge and assets, but this has gone unexplored in literature. In this research-in-progress, a video classifier to detect cybercriminal content in mainstream social media is designed and implemented. A collection of hacking and non-hacking videos was retrieved from a popular social media website to serve as a testbed. Feature sets included video metadata as well as features engineered from the videos themselves, including object detection and aesthetic qualities. This study demonstrates a methodological proof-of-concept that can enable future research that further investigates cyber-adversarial video contents, which have remained largely unexplored to this day. This study also contributes to literature regarding cyber-adversarial contents in mainstream social media.
Bingyan Du, Pranay Singhal, Victor A. Benjamin, Weifeng Li 0002
ISI4
2020 Identifying, Collecting, and Monitoring Personally Identifiable Information: From the Dark Web to the Surface Web
abstract
Personally identifiable information (PII) has become a major target of cyber-attacks, causing severe losses to data breach victims. To protect data breach victims, researchers focus on collecting exposed PII to assess privacy risk and identify at-risk individuals. However, existing studies mostly rely on exposed PII collected from either the dark web or the surface web. Due to the wide exposure of PII on both the dark web and surface web, collecting from only the dark web or the surface web could result in an underestimation of privacy risk. Despite its research and practical value, jointly collecting PII from both sources is a non-trivial task. In this paper, we summarize our effort to systematically identify, collect, and monitor a total of 1,212,004,819 exposed PII records across both the dark web and surface web. Our effort resulted in 5.8 million stolen SSNs, 845,000 stolen credit/debit cards, and 1.2 billion stolen account credentials. From the surface web, we identified and collected over 1.3 million PII records of the victims whose PII is exposed on the dark web. To the best of our knowledge, this is the largest academic collection of exposed PII, which, if properly anonymized, enables various privacy research inquiries, including assessing privacy risk and identifying at-risk populations.
Fang Yu Lin, Zara Ahmad-Post, Reza Ebrahimi 0001, James Lee Hu, Jingyu Xin, Weifeng Li 0002, Hsinchun Chen
ISI8
2020 A Generative Adversarial Learning Framework for Breaking Text-Based CAPTCHA in the Dark Web
abstract
Cyber threat intelligence (CTI) necessitates automated monitoring of dark web platforms (e.g., Dark Net Markets and carding shops) on a large scale. While there are existing methods for collecting data from the surface web, large-scale dark web data collection is commonly hindered by anti-crawling measures. Text-based CAPTCHA serves as the most prohibitive type of these measures. Text-based CAPTCHA requires the user to recognize a combination of hard-to-read characters. Dark web CAPTCHA patterns are intentionally designed to have additional background noise and variable character length to prevent automated CAPTCHA breaking. Existing CAPTCHA breaking methods cannot remedy these challenges and are therefore not applicable to the dark web. In this study, we propose a novel framework for breaking text-based CAPTCHA in the dark web. The proposed framework utilizes Generative Adversarial Network (GAN) to counteract dark web-specific background noise and leverages an enhanced character segmentation algorithm. Our proposed method was evaluated on both benchmark and dark web CAPTCHA testbeds. The proposed method significantly outperformed the state-of-the-art baseline methods on all datasets, achieving over 92.08% success rate on dark web testbeds. Our research enables the CTI community to develop advanced capabilities of large-scale dark web monitoring.
Reza Ebrahimi 0001, Weifeng Li 0002, Hsinchun Chen
ISI3
2018 Supervised Topic Modeling Using Hierarchical Dirichlet Process-Based Inverse Regression: Experiments on E-Commerce Applications
abstract
The proliferation of e-commerce calls for mining consumer preferences and opinions from user-generated text. To this end, topic models have been widely adopted to discover the underlying semantic themes (i.e., topics). Supervised topic models have emerged to leverage discovered topics for predicting the response of interest (e.g., product quality and sales). However, supervised topic modeling remains a challenging problem because of the need to prespecify the number of topics, the lack of predictive information in topics, and limited scalability. In this paper, we propose a novel supervised topic model, Hierarchical Dirichlet Process-based Inverse Regression (HDP-IR). HDP-IR characterizes the corpus with a flexible number of topics, which prove to retain as much predictive information as the original corpus. Moreover, we develop an efficient inference algorithm capable of examining large-scale corpora (millions of documents or more). Three experiments were conducted to evaluate the predictive performance over major e-commerce benchmark testbeds of online reviews. Overall, HDP-IR outperformed existing state-of-the-art supervised topic models. Particularly, retaining sufficient predictive information improved predictive R-squared by over 17.6 percent; having topic structure flexibility contributed to predictive R-squared by at least 4.1 percent. HDP-IR provides an important step for future study on user-generated texts from a topic perspective.
Weifeng Li 0002, Junming Yin, Hsinchun Chen
IEEE Trans. Knowl. Data Eng.1
2016 Exploring key hackers and cybersecurity threats in Chinese hacker communities
abstract
Chinese hacker communities are of interest to cybersecurity researchers and investigators. When examining Chinese hacker communities, researchers and investigators face many challenges, including understanding the Chinese language, detecting variations in topic evolution, and identifying key hackers with their specialty areas. Therefore, we are motivated to develop a framework for analyzing key hackers and emerging threats in Chinese hacker communities. Specifically, we develop a set of topic models for extracting popular topics, tracking topic evolution, and identifying key hackers with their specialty topics. We applied our framework to 19 major Chinese hacker communities. As a result, we identified five major popular topics, including trading, fraud prevention & identification, calling for cooperation, casual chat, and monetizing. Moreover, we found several trends related to new communication channels, new stolen cards of interest, and new operating mechanism. Further, we also found the key hackers in each extracted area. Our work contributes to the cybersecurity literature by providing an advanced and scalable framework for analyzing Chinese hacker communities.
Qiang Wei 0001, Yong Zhang 0002, Chunxiao Xing, Weifeng Li 0002, Hsinchun Chen
ISI7
2016 Targeting key data breach services in underground supply chain
abstract
Over the past decade, a growing body of cybercriminals have founded an underground supply chain to facilitate data breaches, leading to the leak of personal information for hundreds of millions of individuals. As many service providers in the supply chain are rippers, cybercriminals tend to rely on a few key services. Identifying key services is of great interest to both cybersecurity researchers and practitioners. This study presents a text-mining framework for identifying key data breach services based on analysis of service reviews. The framework includes crawlers with counter anti-crawling measures, text preprocessing, and supervised topic models. In our experiment, more than 70% of the key services were identified by our framework.
Weifeng Li 0002, Junming Yin, Hsinchun Chen
ISI1
2016 Chinese underground market jargon analysis based on unsupervised learning
abstract
With the rapid growth of online population, China has become the world's largest online market. This also gives rise to the Chinese underground market, which has facilitated many of the cybercrimes in China. Consequently, there is a need for research scrutinizing Chinese underground markets. One major challenge facing cybersecurity researchers is to understand the unfamiliar cybercriminal jargons. To this end, we are motivated to analyze jargons in Chinese underground market. Particularly, we utilize the recent advancements in unsupervised machine learning methods, word embedding and Latent Dirichlet Allocation. We evaluate our work on a research testbed encompassing 29 exclusive underground market QQ groups with 23,000 members. Specifically, we test the ability of the proposed approach to learn semantically similar words of known cybersecurity-related jargons. Results suggest the state-of-the-art unsupervised learning approaches can help better understand cybercriminal language, providing promising insights for future research on Chinese underground markets.
Kangzhi Zhao, Yong Zhang 0002, Chunxiao Xing, Weifeng Li 0002, Hsinchun Chen
ISI4
2015 Exploring threats and vulnerabilities in hacker web: Forums, IRC and carding shops
abstract
Cybersecurity is a problem of growing relevance that impacts all facets of society. As a result, many researchers have become interested in studying cybercriminals and online hacker communities in order to develop more effective cyber defenses. In particular, analysis of hacker community contents may reveal existing and emerging threats that pose great risk to individuals, businesses, and government. Thus, we are interested in developing an automated methodology for identifying tangible and verifiable evidence of potential threats within hacker forums, IRC channels, and carding shops. To identify threats, we couple machine learning methodology with information retrieval techniques. Our approach allows us to distill potential threats from the entirety of collected hacker contents. We present several examples of identified threats found through our analysis techniques. Results suggest that hacker communities can be analyzed to aid in cyber threat detection, thus providing promising direction for future work.
Victor A. Benjamin, Weifeng Li 0002, Thomas Holt, Hsinchun Chen
ISI2