Fan Shi 0003

dblp:96/8708-3 · DBLP profile ↗
← Back
30ranked-venue papers
1as first author
28since 2021 · last 2026
0000-0003-4533-2706ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 17 · 1 first-author · 15 since 2021Computer networks · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Device Type Identification with Deep Metric Learning
Xinyu Yin, Fan Shi 0003, Chengxi Xu, Jinfeng Peng, Jiatang Zhao
ICIC (4)2
2026 Beyond Up and Down: Analyzing Temporal Phenomena in IPv6 Probing Responses
Yudong Lian, Jinfeng Peng, Chengxi Xu, Fan Shi 0003, Min Zhang 0054, Jiatang Zhao, Quming Peng
IWQoS5
2026 CoT-DPG: A Co-Training based Dynamic Password Guessing Method
Fan Shi 0003, Min Zhang 0054, Chengxi Xu, Shasha Guo 0001, Jinghua Zheng
NDSS2
2026 TNet: Efficient IPv6 active network discovery
Jiatang Zhao, Fan Shi 0003, Chengxi Xu, Jinfeng Peng, Mingyi Ge, Min Zhang 0054
Comput. Networks2
2026 Darkness at dawn: understanding illicit websites in newly registered domain names
abstract
Abstract Illicit website represents a significant challenge on the Internet. Miscreants exploit the inherent flexibility and invisibility of the Internet to promote illicit activities, particularly online gambling and pornography, intending to generate substantial profits. Previous studies have primarily focused on illicit website detection techniques and analyzed illicit activities using passive datasets. However, constrained by the limitations of passive dataset perspectives, the security community lacks a global understanding of illicit website deployment and operational behavior patterns, particularly during the early stages of website activation. In this paper, we conduct an in-depth analysis of the activities of illicit websites through the advantageous lens of newly registered domains (NRDs). The NRD dataset’s key strength is its broad coverage of emerging illicit activities during observation, complementing previous studies. Specifically, we designed and implemented a framework, NRDMiner, for tracking and analyzing illicit activities associated with large-scale NRDs. This framework supports long-term monitoring of vast quantities of domains and enables accurate identification of illicit websites. Over a 133-day period (July 1–Nov 10, 2024), we collected 27,623,326 NRDs across 481 top-level domains (e.g., and ), and identified 910,794 abusive domains. Our analysis highlights several important patterns. First, illicit activity shows a consistent and steady pattern, with an average of 3.3% of NRDs flagged for illicit website. Moreover, 98% of these domains are first-time registrations. Second, 60% of abusive domains are activated on the same day they are registered, indicating mature automated domain abuse techniques. Third, from a global NRD perspective, we observed regional tendencies in illicit activities, like Asia identified as the primary concentration area, with over 70% of illicit website pages being in Asian languages. Furthermore, we analyzed the deployment and operation of illicit websites. Our work provides a large-scale empirical study of the early-stage activities of illicit websites from the perspective of NRDs, offering valuable evidence that contribute to the timely mitigation of illicit activities.
Bingyang Guo, Fan Shi 0003, Min Zhang 0054, Chengxi Xu, Yi Shen 0012
Cybersecur.3
2026 Trilink: discovering embedded siblings using a novel approach
abstract
Abstract Due to the bucket effect, dual-stack hosts face more severe security risks than single-stack hosts, making the discovery and identification of dual-stack hosts particularly important. Traditional studies employ methods such as domain name association and service fingerprinting for dual-stack identification; however, these methods suffer from incomplete identification and limited dual-stack scale. To solve this issue, we introduce the Trilink algorithm, which performs dual-stack host discovery and identification, as well as conducts security analysis, by verifying whether IPv6 addresses conform to the potential dual-stack address format standards, comparing the consistency of port fingerprints between IPv4 and IPv6, and utilizing IP geolocation and IP address ASN matching. The results show that we have discovered a total of 204,825 dual-stack devices across 118 countries and 269 autonomous systems. Meanwhile, our research reveals that dual-stack devices have 27% higher asset exposure across common service types than single-stack devices.
Fan Shi 0003, Mingyi Ge, Chengxi Xu, Jiatang Zhao, Xinyu Yin
Cybersecur.1
2026 CyMapNER: a named entity recognition model for cyberspace surveying and mapping domain
abstract
Abstract Cyberspace Surveying and Mapping (CSM) involves the identification and analysis of digital assets to support network management and security, yet its domain-specific named entity recognition (NER) remains underexplored. A key challenge is the semantic gap between general-domain corpora and CSM domain texts, the suboptimal performance of existing named entity recognition (NER) models in accurately identifying entities within CSM data. To tackle obstacles, we proposed a NER model CyMapNER for the CSM domain. A clear definition of named entity categories pertinent to the CSM domain was established initially, followed by the creation of a dedicated NER dataset tailored to this domain. Subsequently, we present a domain adaptation training framework that integrates large language models. It combines with data augments, pseudo-labeling and domain-adaptive pretraining to enhance the adaptability of the NER model. The comparative experimental results demonstrate that CyMapNER models outperforms traditional NER models in CSM datasets. The results reveal that by domain adaptation training framework, the recognition accuracy of CyMapNER model reaches 97%, which achieves an improvement from 5.6% to 18.3% over the state-of-the-art NER models, and it performs well in recognizing complex and sparse entities, highlighting its effectiveness in handling the intricacies of CSM data.
Fan Shi 0003, Chengxi Xu, Xinyu Yin, Mingyi Ge
Cybersecur.2
2026 PGMaP: Password generation based on mask prediction
abstract
Numerous studies have focused on data-driven password guessing methods in recent years, aiming to reduce the use of weak passwords by users and improve password security. Existing password generation models learn the distribution of password datasets and generate candidate guesses by fitting sequential conditional probabilities. These methods are based on a key assumption: users construct passwords in one direction from left to right. However, with the more complex password policy requirements of authentication systems and the increasing security awareness of people, users construct passwords by modifying existing or popular passwords. At this point, users consider global and bi-directional information of passwords. This breaks the key assumption of uni-directional construction and leads to omissions when generating passwords by existing methods. Motivated by this, we propose a password generation method based on mask prediction, named PGMaP, which captures this large number of omitted passwords. First, we design a password construction template extraction algorithm to cluster the templates used by users for constructing and modifying passwords. Then we construct a transformer-based masked language model to learn password bi-directional features. The extracted templates are fed into the model to generate password guesses by means of mask prediction. Different from existing auto-regressive model based methods that generate in one direction, PGMaP uses the auto-encoding model to generate passwords based on the bidirectional information. Finally, through password guessing experiments all eight real-world datasets, we demonstrate that PGMaP can effectively generate a large number of omitted passwords, and its password guessing performance outperforms existing methods.
Fan Shi 0003, Shasha Guo 0001, Min Zhang 0054, Yi Shen 0012, Chengxi Xu
Expert Syst. Appl.2
2026 Intelligent Penetration Testing Through Integrated Knowledge Graph and Historical Decision Enhancement
abstract
Penetration Testing (PT), a key network security assessment technique that simulates real cyber attacks to identify vulnerabilities, is traditionally manual and expert-dependent, leading to low efficiency and high costs. Automating and intelligentizing PT has thus become a critical research focus, yet current technologies face two core challenges: lack of standardized, reusable simulated network scenarios (hindering unified experiments and result comparison) and intelligent models' failure to integrate historical decision experience or utilize attack chain temporal correlations (restricting adaptability). To address these, this study proposes an intelligent PT method integrating knowledge graph-driven automated scenario construction and historical decision enhancement. Two innovations are introduced: a network knowledge graph-based mechanism to generate standardized, real-characteristic testing environments; and a historical decision enhancement scheme with a collaborative state temporal processing and action filtering architecture. Experimental results show the method reduces average iterations by 69%, eliminates redundant executions, and enhances decision rationality, offering a new path for automated PT advancement.
Qianyu Li 0001, Anupam Chattopadhyay, Cheng Tu, Fan Shi 0003, Min Zhang 0054, Zulie Pan
IEEE Trans. Dependable Secur. Comput.6
2026 Password Guessing Based on Hidden Weak Password Analysis
abstract
Password has become the mainstream method of authentication today. To improve password security, researchers evaluate the strength of target password datasets through early brute-force attacks to current password guessing methods, aiming to help users reduce the use of weak passwords. With users becoming more aware of security, they make local variations on weak passwords to improve the password strength while being easy to remember. These transformations render passwords more complex and enhance the score in password strength meter. However, such variations do not genuinely enhance password security, as human habits tend to converge. This allows attackers to deduce the modification patterns and consequently crack these passwords. Motivated by this, this paper defines the hidden weak passwords, a local variant of explicit weak passwords, which appear to enhance password security yet remain vulnerable. We systematically analyze transformation behavior between explicit and hidden weak passwords. Then we design an automated rule generation algorithm to identify hidden weak passwords and generate transformation rules. Based on automatically mined rules, we generate a large number of password guesses and fuses them with existing methods to improve password guessing performance. Finally, we demonstrate the effectiveness of the proposed method through password guessing experiments on eight real-world datasets, where the cracking rate improves on all five state-of-the-art methods.
Min Zhang 0054, Zhijie Xie, Shasha Guo 0001, Yuliang Lu, Fan Shi 0003, Yi Shen 0012
IEEE Trans. Dependable Secur. Comput.6
2026 MetaRAG: Identifying Website Owner Using Meta-Path-Guided Dynamic Graph Retrieval-Augmented Generation
abstract
Website owner identification aims to link websites to their real-world owners, which is crucial for credibility assessment and information provenance in information retrieval and vital for applications in cybersecurity, Internet governance, and digital regulation. Existing approaches for website owner identification primarily rely on querying infrastructure registration records or analyzing webpage content. However, these methods often fail due to incomplete or outdated registration records and sparse webpage content. We observe that inter-website relationships, derived from shared infrastructure data such as primary domains, IP blocks, and geolocations, can provide valuable but underutilized ownership cues. To exploit this insight, we propose MetaRAG, a meta-path-guided dynamic graph retrieval-augmented generation framework that performs reasoning using large language models over ownership-relevant paths in a website-centric knowledge graph. MetaRAG consists of three components: (1) a knowledge graph construction module that integrates infrastructure data and crawled webpage content into a unified representation; (2) a meta-path-guided dynamic reasoning module that constrains retrieval to ownership-relevant meta-paths and adaptively decides whether to retrieve more information or perform inference based on evidence completeness; and (3) a multi-path evidence refinement module that aggregates and scores retrieved paths to suppress noise and distill high-confidence ownership signals. We evaluate MetaRAG on two constructed real-world datasets, achieving up to 6.82% improvement over strong baselines. The results demonstrate the effectiveness of our approach in combining structured web knowledge with large language model-based reasoning for more accurate website owner identification.
Cheng Tu, Yunshan Ma 0002, Bingyang Guo, Qianyu Li 0001, Yang Li 0215, Min Zhang 0054, Fan Shi 0003, Xiang Wang 0010
ACM Trans. Inf. Syst.7
2025 Email Cloaking: Deceiving Users and Spam Email Detectors with Invisible HTML Settings
Bingyang Guo, Mingxuan Liu 0006, Yihui Ma, Ruixuan Li 0008, Fan Shi 0003, Min Zhang 0054, Baojun Liu 0002, Chengxi Xu, Hai-Xin Duan, Geng Hong, Min Yang 0002, Qingfeng Pan
ESORICS (4)5
2025 NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search Engines
Mingxuan Liu 0006, Lijie Wu, Baojun Liu 0002, Geng Hong, Yiming Zhang 0009, Jia Zhang 0004, Hai-Xin Duan, Min Zhang 0054, Fan Shi 0003, Min Yang 0002
USENIX Security Symposium12
2025 Misty Registry: An Empirical Study of Flawed Domain Registry Operation
Mingming Zhang 0010, Baojun Liu 0002, Hai-Xin Duan, Min Zhang 0054, Fan Shi 0003, Chengxi Xu
USENIX Security Symposium6
2025 A Causal Adjustment Module for Debiasing Scene Graph Generation
abstract
While recent debiasing methods for Scene Graph Generation (SGG) have shown impressive performance, these efforts often attribute model bias solely to the long-tail distribution of relationships, overlooking the more profound causes stemming from skewed object and object pair distributions. In this paper, we employ causal inference techniques to model the causality among these observed skewed distributions. Our insight lies in the ability of causal inference to capture the unobservable causal effects between complex distributions, which is crucial for tracing the roots of model bias. Specifically, we introduce the Mediator-based Causal Chain Model (MCCM), which, in addition to modeling causality among objects, object pairs, and relationships, incorporates mediator variables, i.e., cooccurrence distribution, for complementing the causality. Following this, we propose the Causal Adjustment Module (CAModule) to estimate the modeled causal structure, using variables from MCCM as inputs to produce a set of adjustment factors aimed at correcting biased model predictions. Moreover, our method enables the composition of zero-shot relationships, thereby enhancing the model's ability to recognize such relationships. Experiments conducted across various SGG backbones and popular benchmarks demonstrate that CAModule achieves state-of-the-art mean recall rates, with significant improvements also observed on the challenging zero-shot recall rate metric.
Li Liu 0002, Shuzhou Sun, Shuaifeng Zhi, Fan Shi 0003, Zhen Liu 0004, Janne Heikkilä, Yongxiang Liu
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Understanding and Characterizing the Adoption of Internationalized Domain Names in Practice
abstract
Internationalized Domain Names (IDNs) allow users to access the internet using domain names in their native languages. This technology provides significant convenience for non-English speaking users. However, despite the widespread acceptance and use of IDNs, the risks associated with using IDNs remain unclear in practice, such as the IDN homograph problem. To address this issue, we conduct a systematic analysis of the IDN homograph problem and explore the adoption characteristics of IDNs in practice. Specifically, we design and implement an effective IDN analysis framework, named as IDNMon. We perform a large-scale measurement study covering 863 top-level domain zone files and historical top lists based on IDNMon. Our findings indicate that the IDN registration and usage in Europe exceeds that in East Asia. Our results confirm that the IDN homograph problem is universal (12.32% of 2,623,161 IDNs face this problem), which raises serious challenges when designing protection strategies for browsers. Our work provides new insights into the adoption of IDNs in practice, contributes to a better understanding, and promotes the development of IDNs.
Chengxi Xu, Fan Shi 0003, Min Zhang 0054, Yuwei Li 0002, Zhijie Xie
IEEE Trans. Dependable Secur. Comput.3
2025 Heterogeneous Binary Pixel Difference Networks for Remote Sensing Object Detection
abstract
Recent research in remote sensing object detection (RSOD) has significantly advanced the development of vision foundation models. However, deploying these models on resource-constrained edge devices is challenging due to their high computational demands. Binarized detectors utilize binary neural networks (BNNs) to achieve extreme compression by quantizing weights and activations to +1 or −1, which have been extensively studied for generic object detection tasks. In remote sensing images, the objects of interest typically exhibit weak responses, and the images often contain numerous unique local areas. Feature binarization in these images can lead to substantial loss of object contrast and scale prior information, which exacerbates performance issues, particularly for small objects, resulting in significant performance degradation. To address these challenges, we propose a novel binarized detector for RSOD named the heterogeneous binary pixel difference network (HBiPiDiNet). Initially, we developed a binary pixel difference convolution (BiPDC) that integrates local binary patterns (LBPs) to capture local contrast information with traditional binary convolution, thereby enhancing the representation of small objects. Subsequently, we constructed heterogeneous kernel fusion convolution blocks (HKFCB) based on BiPDC and standard binary convolution. The HKFCB comprises multiple BiPDCs at different scales, effectively representing BiPDC under multiscale LBP and multiscale binary convolutions. Extensive experiments demonstrate that our proposed method significantly enhances the performance of state-of-the-art binary detection methods across three remote sensing datasets: AI-TOD, VisDrone2019, and DIOR. We have released our code and models athttps://github.com/yuhua666/HBiPiDiNet/tree/main.
Jialei Zhan, Liang Bai 0003, Tianpeng Liu, Fan Shi 0003, Yongxiang Liu, Li Liu 0002
IEEE Trans. Geosci. Remote. Sens.5
2025 Website Owner Identification through Multi-level Contrastive Representation Learning
abstract
Website owner identification aims to recognize the organization or individual who owns a given website that is served on the web. It is a crucial step for cyberspace surveying and mapping, playing a significant role in cyberspace administration and governance. Existing widely employed solutions for website owner identification mainly fall into two paradigms: (1) querying the public information databases such as WHOIS, which store the Internet resource’s registered users or assignees; and (2) directly extracting the organization or individual name of the website owner from the webpage using the technique of named entity recognition. However, the former is less reliable due to the incomplete, encrypted, and outdated records in the public information databases. Meanwhile, the latter requires that the webpages explicitly and precisely present their owner names without ambiguity, which is often hard to guarantee in practice. To address these limitations, we propose to formulate website owner identification as a problem of webpage representation learning, thereby introducing a novel representation learning framework empowered by large language model-based text Rewriting and Multi-level contrastive learning, named ReMon. First, we devise a prompt to rewrite the webpages using large language models, which effectively filters out noise from the original webpages. Second, we model website–website, website–owner, and owner–owner interactions through multi-level contrastive learning, fully utilizing the self-supervision signals on long-tail items to learn the multi-level constraints. Third, we design a retrieval-based prediction framework and a clustering-based framework to apply websites’ and owners’ representations for different scenarios of the website owner identification task. To evaluate ReMon under our formulation, we construct two datasets based on real-world data. Compared to existing approaches, our ReMon can address the challenging scenarios when valid information cannot be found in public information databases and the owner’s name does not appear on the webpage. Meanwhile, the experimental results show that ReMon outperforms all representation learning-based baselines and significantly enhances training efficiency. The code is available at https://github.com/tuchen9/ReMon .
Cheng Tu, Yunshan Ma 0002, Yang Li 0215, Min Zhang 0054, Fan Shi 0003, Xiang Wang 0010
ACM Trans. Knowl. Discov. Data6
2024 CAKGC: A Clustering Method of Cybercrime Assets Knowledge Graph Based on Feature Fusion
Fan Shi 0003, Chengxi Xu, Jiankun Sun
ICIC (9)2
2024 Poster: A Fistful of Queries: Accurate and Lightweight Anycast Enumeration of Public DNS
Chengxi Xu, Fan Shi 0003
IMC4
2024 Poster: Capture the List: Ranking Manipulation Leveraging Open Forwarders
Chengxi Xu, Fan Shi 0003
IMC3
2024 Rethinking the Security Threats of Stale DNS Glue Records
Baojun Liu 0002, Hai-Xin Duan, Min Zhang 0054, Xiang Li 0108, Fan Shi 0003, Chengxi Xu, Eihal Alowaisheq
USENIX Security Symposium6
2024 Into the Dark: Unveiling Internal Site Search Abused for Black Hat SEO
Mingxuan Liu 0006, Baojun Liu 0002, Yiming Zhang 0009, Hai-Xin Duan, Min Zhang 0054, Fan Shi 0003
USENIX Security Symposium9
2024 Cross the Zone: Toward a Covert Domain Hijacking via Shared DNS Infrastructure
Mingming Zhang 0010, Baojun Liu 0002, Jia Zhang 0004, Hai-Xin Duan, Min Zhang 0054, Fan Shi 0003, Chengxi Xu
USENIX Security Symposium8
2024 DynPen: Automated Penetration Testing in Dynamic Network Scenarios Using Deep Reinforcement Learning
abstract
Penetration testing, a crucial industrial practice for securing networked systems and infrastructures, has traditionally depended on the extensive expertise of human professionals. Addressing the scarcity of human experts, the development of automated penetration testing tools emerges as a promising avenue. Against the backdrop of rapid advancements in artificial intelligence technologies, reinforcement learning has demonstrated considerable potential for realizing automated penetration testing. However, existing research predominantly concentrates on reinforcement learning-based automated penetration testing tools within static scenarios, with limited exploration in dynamic network environments. This paper addresses a noteworthy challenge in developing autonomous agents for real-world applications, particularly focusing on scenarios marked by environmental changes. Such alterations necessitate autonomous agents to continuously monitor environmental characteristics, and adapt, and adjust learned actions to ensure the system’s effective operation. Consequently, the paper proposes an automated reinforcement learning-based penetration testing scheme tailored for dynamic network scenarios, named DynPen. DynPen captures observed changes in the scenario, aiding the penetration testing agent in decision-making based on historical experiences. Simulation results demonstrate the proposed scheme’s efficacy in significantly expediting the convergence speed of the penetration testing agent using reinforcement learning algorithms. Furthermore, the scheme successfully maintains the learning agility and adaptability of the agent in dynamic network scenarios.
Qianyu Li 0001, Dong Li 0054, Fan Shi 0003, Min Zhang 0054, Anupam Chattopadhyay, Yi Shen 0012, Yang Li 0215
IEEE Trans. Inf. Forensics Secur.4
2024 GuessFuse: Hybrid Password Guessing With Multi-View
abstract
Password guessing is a primary method for password strength evaluation. Despite various password guessing models have been proposed, there is still a significant gap between their guessing effectiveness and the actual cracking capabilities of attackers. Integrating multiple models for password guessing, also known as hybrid password guessing, could better capture the cracking capabilities of real attackers. However, the reason why hybrid password guessing can enhance cracking capabilities, and how to effectively integrate multiple heterogeneous password guessing models, are still not well understood. To address these issues, this paper draws inspiration from the concept of multi-view learning. We regard the guess lists generated by various password guessing models as multiple views of the data. Through a comprehensive analysis of these guess lists, we have identified the key reason why hybrid password guessing can enhance the cracking capabilities: integrating more diverse views allows for the coverage of a wider range of heterogeneous password characteristics, and provides more detailed information on effective password distributions. Based on the these findings, we propose a new hybrid password guessing framework, namedGuessFuse.GuessFuseemploys the multi-view subset extraction module and segment splitting selection module to accurately extract and reorganize the effective password from multiple guess lists. Experimental results on six large-scale datasets demonstrate the effectiveness ofGuessFuse. By combining two (resp. five) guess lists,GuessFuseoutperforms its foremost counterparts by an average of 11.00% ~ 59.62% (resp. 4.70% ~ 17.66%) within 107guesses.GuessFusecan effectively improve the cracking success rate under a limited number of guesses, approaching the actual cracking capabilities of attackers.
Zhijie Xie, Fan Shi 0003, Min Zhang 0054, Huimin Ma 0004, Huaixi Wang, Zhenhan Li
IEEE Trans. Inf. Forensics Secur.2
2024 An Intelligent Penetration Testing Method Using Human Feedback
abstract
Penetration testing is widely acknowledged as the foremost method for evaluating network security. However, three challenges impede the generation of strategies that align with human expectations. In this article, we present, for the first time, a method based on human feedback to enhance strategy generation. Our approach comprises two components: agent training and decision-making. During agent training, we establish a hierarchical framework to decompose tasks and a knowledge base to offer advice for improving data efficiency. We then impose constraints on the action space to mitigate ineffective exploration. Finally, we train a reward model based on human feedback and fine tune the model guided by this reward model. In decision-making, we process the model output to enhance decision accuracy. We crafted scenarios based on real-world networks, and the results demonstrate the effectiveness of our method in generating penetration testing strategies that align more closely with human intentions.
Qianyu Li 0001, Min Zhang 0054, Fan Shi 0003, Yi Shen 0012, Bingyang Guo, Chengxi Xu
IEEE Trans. Ind. Informatics4
2021 Mining Centralization of Internet Service Infrastructure in the Wild
abstract
The last decade has witnessed the rapid evolution of the Internet structure, one of which is centralization, that is, Internet core infrastructure has been constantly transferred into the hands of a few popular market participants. Researchers are trying to measure centrality and analyze its security impact from the perspective of traffic analysis. But the underlying distribution of service providers is still enveloped in mysterious veils. In order to address this problem and assess the security risk associated with such centralization. Firstly, we performed linear regression on the data of each kind of service provider in the Alexa Top 1M domains to study the current underlying distribution of various services for the first time. The results show that Zipf’s law is universal in various service providers’ market share, which proves that Internet service infrastructures are centralized. Secondly, we explored the security impacts of centralized infrastructures on the Internet. we conducted attack simulations on providers. Results show that intentional attacks on core providers can greatly downgrade the performance of the Internet. To make matters worse, the quantitative analysis of the provider’s infrastructures found that a considerable number of provider’s infrastructures have low diversity. In addition, we proposed an algorithm to calculate the dependencies between different types of service providers and carried out an evaluation of our datasets, and found the tendency for different services to depend on each other. Our results indicate that the Internet is facing huge security challenges, because the centralized infrastructure will impair service redundancy, and at the same time, it will also cause dependence between infrastructures, which in turn strengthens its centralization.
Bingyang Guo, Fan Shi 0003, Chengxi Xu, Min Zhang 0054, Yang Li 0215
MSN2
2020 Covert timing channel detection method based on time interval and payload length analysis
Jiaxuan Han, Cheng Huang 0003, Fan Shi 0003
Comput. Secur.3
2020 Binary File's Visualization and Entropy Features Analysis Combined with Multiple Deep Learning Networks for Malware Classification
abstract
In recent years, the research on malware variant classification has attracted much more attention. However, there are still many challenges, including the low accuracy of classification of samples of similar malware families, high time, and resource consumption. This paper proposes a new method of malware classification based on multiple visual features of malware and deep learning algorithms. In prior research, visualization techniques and entropy demonstrated exemplary performance in many areas. This paper extracts numerous visual features from the raw bytes and entropy sequence of the malware, which makes it more sensitive to malware samples of similar families and endows it the ability to classify malware variants more accurately. To evaluate the proposed method, this paper conducted a series of experiments on two malware datasets with a total of more than 20,000 samples provided by the Malware Research Lab and Microsoft Research. Through experiments, the method showed its superiority compared with some leading malware visual classification methods, achieving good performance on the accuracy with at least 1% improvement. The accuracy of the method even could reach 99.73% and 99.54%, respectively, on the two datasets.
Shuguang Huang, Cheng Huang 0003, Fan Shi 0003, Min Zhang 0054, Zulie Pan
Secur. Commun. Networks4