EDBT 2026 Demo / reviewers in the wild / expert
Min Zhang 0054
dblp:83/5342-54
· DBLP profile ↗
34ranked-venue papers
1as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 20 · 19 since 2021Computer networks · 5 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Up and Down: Analyzing Temporal Phenomena in IPv6 Probing Responses
Yudong Lian, Jinfeng Peng, Chengxi Xu, Fan Shi 0003, Min Zhang 0054, Jiatang Zhao, Quming Peng |
IWQoS | 6 |
| 2026 | Demystifying the Access Control Mechanism of ESXi VMKernel
Zexiang Zhang, Jiaxun Zhu, Jiaqing Huang, Wenbo Shen, Yuliang Lu, Min Zhang 0054, Zulie Pan |
NDSS | 9 |
| 2026 | CoT-DPG: A Co-Training based Dynamic Password Guessing Method
Fan Shi 0003, Min Zhang 0054, Chengxi Xu, Shasha Guo 0001, Jinghua Zheng |
NDSS | 3 |
| 2026 | TNet: Efficient IPv6 active network discovery
Jiatang Zhao, Fan Shi 0003, Chengxi Xu, Jinfeng Peng, Mingyi Ge, Min Zhang 0054 |
Comput. Networks | 7 |
| 2026 | Darkness at dawn: understanding illicit websites in newly registered domain namesabstractAbstract Illicit website represents a significant challenge on the Internet. Miscreants exploit the inherent flexibility and invisibility of the Internet to promote illicit activities, particularly online gambling and pornography, intending to generate substantial profits. Previous studies have primarily focused on illicit website detection techniques and analyzed illicit activities using passive datasets. However, constrained by the limitations of passive dataset perspectives, the security community lacks a global understanding of illicit website deployment and operational behavior patterns, particularly during the early stages of website activation. In this paper, we conduct an in-depth analysis of the activities of illicit websites through the advantageous lens of newly registered domains (NRDs). The NRD dataset’s key strength is its broad coverage of emerging illicit activities during observation, complementing previous studies. Specifically, we designed and implemented a framework, NRDMiner, for tracking and analyzing illicit activities associated with large-scale NRDs. This framework supports long-term monitoring of vast quantities of domains and enables accurate identification of illicit websites. Over a 133-day period (July 1–Nov 10, 2024), we collected 27,623,326 NRDs across 481 top-level domains (e.g., and ), and identified 910,794 abusive domains. Our analysis highlights several important patterns. First, illicit activity shows a consistent and steady pattern, with an average of 3.3% of NRDs flagged for illicit website. Moreover, 98% of these domains are first-time registrations. Second, 60% of abusive domains are activated on the same day they are registered, indicating mature automated domain abuse techniques. Third, from a global NRD perspective, we observed regional tendencies in illicit activities, like Asia identified as the primary concentration area, with over 70% of illicit website pages being in Asian languages. Furthermore, we analyzed the deployment and operation of illicit websites. Our work provides a large-scale empirical study of the early-stage activities of illicit websites from the perspective of NRDs, offering valuable evidence that contribute to the timely mitigation of illicit activities. Bingyang Guo, Fan Shi 0003, Min Zhang 0054, Chengxi Xu, Yi Shen 0012 |
Cybersecur. | 4 |
| 2026 | PGMaP: Password generation based on mask predictionabstractNumerous studies have focused on data-driven password guessing methods in recent years, aiming to reduce the use of weak passwords by users and improve password security. Existing password generation models learn the distribution of password datasets and generate candidate guesses by fitting sequential conditional probabilities. These methods are based on a key assumption: users construct passwords in one direction from left to right. However, with the more complex password policy requirements of authentication systems and the increasing security awareness of people, users construct passwords by modifying existing or popular passwords. At this point, users consider global and bi-directional information of passwords. This breaks the key assumption of uni-directional construction and leads to omissions when generating passwords by existing methods. Motivated by this, we propose a password generation method based on mask prediction, named PGMaP, which captures this large number of omitted passwords. First, we design a password construction template extraction algorithm to cluster the templates used by users for constructing and modifying passwords. Then we construct a transformer-based masked language model to learn password bi-directional features. The extracted templates are fed into the model to generate password guesses by means of mask prediction. Different from existing auto-regressive model based methods that generate in one direction, PGMaP uses the auto-encoding model to generate passwords based on the bidirectional information. Finally, through password guessing experiments all eight real-world datasets, we demonstrate that PGMaP can effectively generate a large number of omitted passwords, and its password guessing performance outperforms existing methods. Fan Shi 0003, Shasha Guo 0001, Min Zhang 0054, Yi Shen 0012, Chengxi Xu |
Expert Syst. Appl. | 4 |
| 2026 | Intelligent Penetration Testing Through Integrated Knowledge Graph and Historical Decision EnhancementabstractPenetration Testing (PT), a key network security assessment technique that simulates real cyber attacks to identify vulnerabilities, is traditionally manual and expert-dependent, leading to low efficiency and high costs. Automating and intelligentizing PT has thus become a critical research focus, yet current technologies face two core challenges: lack of standardized, reusable simulated network scenarios (hindering unified experiments and result comparison) and intelligent models' failure to integrate historical decision experience or utilize attack chain temporal correlations (restricting adaptability). To address these, this study proposes an intelligent PT method integrating knowledge graph-driven automated scenario construction and historical decision enhancement. Two innovations are introduced: a network knowledge graph-based mechanism to generate standardized, real-characteristic testing environments; and a historical decision enhancement scheme with a collaborative state temporal processing and action filtering architecture. Experimental results show the method reduces average iterations by 69%, eliminates redundant executions, and enhances decision rationality, offering a new path for automated PT advancement. Qianyu Li 0001, Anupam Chattopadhyay, Cheng Tu, Fan Shi 0003, Min Zhang 0054, Zulie Pan |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2026 | Password Guessing Based on Hidden Weak Password AnalysisabstractPassword has become the mainstream method of authentication today. To improve password security, researchers evaluate the strength of target password datasets through early brute-force attacks to current password guessing methods, aiming to help users reduce the use of weak passwords. With users becoming more aware of security, they make local variations on weak passwords to improve the password strength while being easy to remember. These transformations render passwords more complex and enhance the score in password strength meter. However, such variations do not genuinely enhance password security, as human habits tend to converge. This allows attackers to deduce the modification patterns and consequently crack these passwords. Motivated by this, this paper defines the hidden weak passwords, a local variant of explicit weak passwords, which appear to enhance password security yet remain vulnerable. We systematically analyze transformation behavior between explicit and hidden weak passwords. Then we design an automated rule generation algorithm to identify hidden weak passwords and generate transformation rules. Based on automatically mined rules, we generate a large number of password guesses and fuses them with existing methods to improve password guessing performance. Finally, we demonstrate the effectiveness of the proposed method through password guessing experiments on eight real-world datasets, where the cracking rate improves on all five state-of-the-art methods. Min Zhang 0054, Zhijie Xie, Shasha Guo 0001, Yuliang Lu, Fan Shi 0003, Yi Shen 0012 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2026 | OwnerHunter: Multilingual Website Owner Identification Powered by Large Language ModelabstractAs cyberspace continues to expand, identifying the organization or individual behind a website has become increasingly vital in security incident response, phishing website detection, and other cybersecurity subfields. An existing solution for it involves analyzing webpage content and extracting owner names using named entity recognition techniques. However, since these techniques operate on a sentence-by-sentence basis, they struggle to identify the true owner when multiple individual or organizational names appear on a webpage. Moreover, they often perform poorly on non-English websites. To address these limitations, we propose OwnerHunter, a novel multilingual framework powered by large language models, which formulates website owner identification as a multilingual document-level information extraction task and utilizes global information from webpages to identify the owner. In OwnerHunter, we first craft prompts that fully leverage the capabilities of large language models to effectively recognize potential owners on webpages in different languages with minimal examples. To enhance the comprehensiveness and accuracy of recognition, we further design a multimodal augmentation strategy, an example pool strategy, and a self-verification strategy. Then, we devise a semantic and string similarity aggregation-based entity disambiguation technique to eliminate ambiguities among multiple potential owners recognized by large language models and a position-based hybrid ranking technique to exactly select the true owner. To evaluate OwnerHunter, we refine the publicly available English dataset ONER and construct the Chinese dataset WOI-cn with 16,036 real websites. Experimental results show that OwnerHunter achieves F1 scores of 0.9505 on ONER and 0.9621 on WOI-cn, setting new state-of-the-art performance on both datasets. Cheng Tu, Enhuan Dong, Zexiang Zhang, Min Zhang 0054, Yang Li 0215, Jiahai Yang 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | MetaRAG: Identifying Website Owner Using Meta-Path-Guided Dynamic Graph Retrieval-Augmented GenerationabstractWebsite owner identification aims to link websites to their real-world owners, which is crucial for credibility assessment and information provenance in information retrieval and vital for applications in cybersecurity, Internet governance, and digital regulation. Existing approaches for website owner identification primarily rely on querying infrastructure registration records or analyzing webpage content. However, these methods often fail due to incomplete or outdated registration records and sparse webpage content. We observe that inter-website relationships, derived from shared infrastructure data such as primary domains, IP blocks, and geolocations, can provide valuable but underutilized ownership cues. To exploit this insight, we propose MetaRAG, a meta-path-guided dynamic graph retrieval-augmented generation framework that performs reasoning using large language models over ownership-relevant paths in a website-centric knowledge graph. MetaRAG consists of three components: (1) a knowledge graph construction module that integrates infrastructure data and crawled webpage content into a unified representation; (2) a meta-path-guided dynamic reasoning module that constrains retrieval to ownership-relevant meta-paths and adaptively decides whether to retrieve more information or perform inference based on evidence completeness; and (3) a multi-path evidence refinement module that aggregates and scores retrieved paths to suppress noise and distill high-confidence ownership signals. We evaluate MetaRAG on two constructed real-world datasets, achieving up to 6.82% improvement over strong baselines. The results demonstrate the effectiveness of our approach in combining structured web knowledge with large language model-based reasoning for more accurate website owner identification. Cheng Tu, Yunshan Ma 0002, Bingyang Guo, Qianyu Li 0001, Yang Li 0215, Min Zhang 0054, Fan Shi 0003, Xiang Wang 0010 |
ACM Trans. Inf. Syst. | 6 |
| 2025 | Email Cloaking: Deceiving Users and Spam Email Detectors with Invisible HTML Settings
Bingyang Guo, Mingxuan Liu 0006, Yihui Ma, Ruixuan Li 0008, Fan Shi 0003, Min Zhang 0054, Baojun Liu 0002, Chengxi Xu, Hai-Xin Duan, Geng Hong, Min Yang 0002, Qingfeng Pan |
ESORICS (4) | 6 |
| 2025 | Insvdf: Interface-State-Aware Virtual Device FuzzingabstractHypervisor is the core technology of virtualization for emulating independent hardware resources for each virtual machine. Virtual devices serve as the main interface of the hypervisor, making the security of virtual devices crucial, as any vulnerabilities can impact the entire virtualization environment and pose a threat to the host machine's security. Direct Memory Access (DMA) is the interface of virtual devices, enabling communication with the host machine. Recently, many efforts have focused on fuzzing against DMA to discover the hypervisor's vulnerabilities. However, the lack of sensitivity to the DMA state causes these efforts to be hindered in efficiency during fuzzing. Specifically, there are two main issues: the uncertain interaction moment and the unclear interaction depth. In this paper, we introduce InSVDF, a DMA interface stateaware fuzzing engine. InSVDF first models the intra-interface state of the DMA interface and incorporates an asynchronyaware state snapshot mechanism along with a depth-aware seed preservation mechanism. To validate our approach, we compare InSVDF with a state-of-the-art fuzzer. The results demonstrate that InSVDF significantly enhances vulnerability discovery speed, with improvements of up to 24.2 x in the best case. Furthermore, InSVDF has identified 2 new vulnerabilities, one of which has been assigned a CVE ID. Zexiang Zhang, Yiming Tao, Zulie Pan, Cheng Tu, Min Zhang 0054, Yang Li 0215, Yi Shen 0012, Chunming Wu 0001 |
ICSE | 7 |
| 2025 | Fusing Multimodal Binary Code Representations for Enhanced Similarity DetectionabstractAs software reuse has become increasingly prevalent in the modern era, binary code similarity detection plays a critical role in program analysis. Numerous machine learning methods have been introduced to this field, with the primary challenge lying in the effective representation of binary code. While existing approaches leverage multimodal features for embedding, they still adopt relatively simple fusion methods to combine different modalities. Therefore, their performance can be limited by inadequate modality interaction during the feature fusion step and an over-reliance on a single modality in the final embedding step. To address these issues, we propose FuseBinRepr, a method that fuses text modality and graph modality representation techniques using a fusion model architecture and three specialized learning tasks to enhance binary code similarity detection. We design a fusion architecture integrating both text and graph embeddings via cross-attention and self-attention mechanisms. For model training, we developed three tasks: text-graph alignment (TGA), graph masking recovery (GMR), and contrastive learning (CL), to capture and align high-level semantics across modalities. Through evaluation, FuseBinRepr demonstrates improvements in Mean Reciprocal Rank (MRR) and Recall compared to state-of-the-art baseline methods, achieving increases of up to 40.8 % and 42 %, respectively. Ablation studies confirm the robustness of our model design, as well as the effectiveness of pretraining tasks. In the downstream software vulnerability detection task, FuseBinRepr achieves the best MRR and Recall performance in ranking CVE binary functions. Taiyan Wang, Yu Chen 0053, Zulie Pan, Min Zhang 0054 |
SRDS | 6 |
| 2025 | NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search Engines
Mingxuan Liu 0006, Lijie Wu, Baojun Liu 0002, Geng Hong, Yiming Zhang 0009, Jia Zhang 0004, Hai-Xin Duan, Min Zhang 0054, Fan Shi 0003, Min Yang 0002 |
USENIX Security Symposium | 10 |
| 2025 | Misty Registry: An Empirical Study of Flawed Domain Registry Operation
Mingming Zhang 0010, Baojun Liu 0002, Hai-Xin Duan, Min Zhang 0054, Fan Shi 0003, Chengxi Xu |
USENIX Security Symposium | 5 |
| 2025 | A survey of binary code representation technologyabstractBinary analysis, as an important foundational technology, provides support for numerous applications in the fields of software engineering and security research. With the continuous expansion of software scale and the complex evolution of software architecture, binary analysis technology is facing new challenges. To break through existing bottlenecks, researchers have applied artificial intelligence (AI) technology to the understanding and analysis of binary code. The core lies in characterizing binary code, i.e., how to use intelligent methods to generate representation vectors containing semantic information for binary code, and apply them to multiple downstream tasks of binary analysis. In this paper, we provide a comprehensive survey of recent advances in binary code representation technology, and introduce the workflow of existing research in two parts, i.e., binary code feature selection methods and binary code feature embedding methods. The feature selection section includes mainly two parts: definition and classification of features, and feature construction. First, the abstract definition and classification of features are systematically explained, and second, the process of constructing specific representations of features is introduced in detail. In the feature embedding section, based on the different intelligent semantic understanding models used, the embedding methods are classified into four categories based on the usage of text-embedding models and graph-embedding models. Finally, we summarize the overall development of existing research and provide prospects for some potential research directions related to binary code representation technology. Taiyan Wang, Qingsong Xie, Zulie Pan, Min Zhang 0054 |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2025 | Understanding and Characterizing the Adoption of Internationalized Domain Names in PracticeabstractInternationalized Domain Names (IDNs) allow users to access the internet using domain names in their native languages. This technology provides significant convenience for non-English speaking users. However, despite the widespread acceptance and use of IDNs, the risks associated with using IDNs remain unclear in practice, such as the IDN homograph problem. To address this issue, we conduct a systematic analysis of the IDN homograph problem and explore the adoption characteristics of IDNs in practice. Specifically, we design and implement an effective IDN analysis framework, named as IDNMon. We perform a large-scale measurement study covering 863 top-level domain zone files and historical top lists based on IDNMon. Our findings indicate that the IDN registration and usage in Europe exceeds that in East Asia. Our results confirm that the IDN homograph problem is universal (12.32% of 2,623,161 IDNs face this problem), which raises serious challenges when designing protection strategies for browsers. Our work provides new insights into the adoption of IDNs in practice, contributes to a better understanding, and promotes the development of IDNs. Chengxi Xu, Fan Shi 0003, Min Zhang 0054, Yuwei Li 0002, Zhijie Xie |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2025 | Website Owner Identification through Multi-level Contrastive Representation LearningabstractWebsite owner identification aims to recognize the organization or individual who owns a given website that is served on the web. It is a crucial step for cyberspace surveying and mapping, playing a significant role in cyberspace administration and governance. Existing widely employed solutions for website owner identification mainly fall into two paradigms: (1) querying the public information databases such as WHOIS, which store the Internet resource’s registered users or assignees; and (2) directly extracting the organization or individual name of the website owner from the webpage using the technique of named entity recognition. However, the former is less reliable due to the incomplete, encrypted, and outdated records in the public information databases. Meanwhile, the latter requires that the webpages explicitly and precisely present their owner names without ambiguity, which is often hard to guarantee in practice. To address these limitations, we propose to formulate website owner identification as a problem of webpage representation learning, thereby introducing a novel representation learning framework empowered by large language model-based text Rewriting and Multi-level contrastive learning, named ReMon. First, we devise a prompt to rewrite the webpages using large language models, which effectively filters out noise from the original webpages. Second, we model website–website, website–owner, and owner–owner interactions through multi-level contrastive learning, fully utilizing the self-supervision signals on long-tail items to learn the multi-level constraints. Third, we design a retrieval-based prediction framework and a clustering-based framework to apply websites’ and owners’ representations for different scenarios of the website owner identification task. To evaluate ReMon under our formulation, we construct two datasets based on real-world data. Compared to existing approaches, our ReMon can address the challenging scenarios when valid information cannot be found in public information databases and the owner’s name does not appear on the webpage. Meanwhile, the experimental results show that ReMon outperforms all representation learning-based baselines and significantly enhances training efficiency. The code is available at https://github.com/tuchen9/ReMon . Cheng Tu, Yunshan Ma 0002, Yang Li 0215, Min Zhang 0054, Fan Shi 0003, Xiang Wang 0010 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2024 | Enhancing Black-box Compiler Option Fuzzing with LLM through Command FeedbackabstractSince the compiler acts as a core component in software building, it is essential to ensure its availability and reliability through software testing and security analysis. Most research has focused on compiler robustness when compiling various test cases, while the reliability of compiler options lacks attention, especially since each option can activate a specific compiler function. Although some researchers have made efforts in testing it, the insufficient utilization of compiler command feedback messages leads to the poor efficiency, which hinders more diverse and in-depth testing.In this paper, we propose a novel solution to enhance black-box compiler option fuzzing by utilizing command feedback, such as error messages, standard output and compiled files, to guide the error fixing and option pruning via prompting large language models for suggestions. We have implemented the prototype and evaluated it on 4 versions of LLVM. Experiments show that our method significantly improves the detection of crashes, reduces false negatives, and even increase the success rate of compilation when compared to the baseline. To date, our method has identified hundreds of unique bugs, and 9 of them are previously unknown. Among these, 8 have been assigned CVE numbers, and 1 has been fixed following our report. Taiyan Wang, Yu Chen 0053, Zulie Pan, Min Zhang 0054, Huimin Ma 0004, Jinghua Zheng |
ISSRE | 6 |
| 2024 | Rethinking the Security Threats of Stale DNS Glue Records
Baojun Liu 0002, Hai-Xin Duan, Min Zhang 0054, Xiang Li 0108, Fan Shi 0003, Chengxi Xu, Eihal Alowaisheq |
USENIX Security Symposium | 4 |
| 2024 | Into the Dark: Unveiling Internal Site Search Abused for Black Hat SEO
Mingxuan Liu 0006, Baojun Liu 0002, Yiming Zhang 0009, Hai-Xin Duan, Min Zhang 0054, Fan Shi 0003 |
USENIX Security Symposium | 6 |
| 2024 | Cross the Zone: Toward a Covert Domain Hijacking via Shared DNS Infrastructure
Mingming Zhang 0010, Baojun Liu 0002, Jia Zhang 0004, Hai-Xin Duan, Min Zhang 0054, Fan Shi 0003, Chengxi Xu |
USENIX Security Symposium | 7 |
| 2024 | DynPen: Automated Penetration Testing in Dynamic Network Scenarios Using Deep Reinforcement LearningabstractPenetration testing, a crucial industrial practice for securing networked systems and infrastructures, has traditionally depended on the extensive expertise of human professionals. Addressing the scarcity of human experts, the development of automated penetration testing tools emerges as a promising avenue. Against the backdrop of rapid advancements in artificial intelligence technologies, reinforcement learning has demonstrated considerable potential for realizing automated penetration testing. However, existing research predominantly concentrates on reinforcement learning-based automated penetration testing tools within static scenarios, with limited exploration in dynamic network environments. This paper addresses a noteworthy challenge in developing autonomous agents for real-world applications, particularly focusing on scenarios marked by environmental changes. Such alterations necessitate autonomous agents to continuously monitor environmental characteristics, and adapt, and adjust learned actions to ensure the system’s effective operation. Consequently, the paper proposes an automated reinforcement learning-based penetration testing scheme tailored for dynamic network scenarios, named DynPen. DynPen captures observed changes in the scenario, aiding the penetration testing agent in decision-making based on historical experiences. Simulation results demonstrate the proposed scheme’s efficacy in significantly expediting the convergence speed of the penetration testing agent using reinforcement learning algorithms. Furthermore, the scheme successfully maintains the learning agility and adaptability of the agent in dynamic network scenarios. Qianyu Li 0001, Dong Li 0054, Fan Shi 0003, Min Zhang 0054, Anupam Chattopadhyay, Yi Shen 0012, Yang Li 0215 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | GuessFuse: Hybrid Password Guessing With Multi-ViewabstractPassword guessing is a primary method for password strength evaluation. Despite various password guessing models have been proposed, there is still a significant gap between their guessing effectiveness and the actual cracking capabilities of attackers. Integrating multiple models for password guessing, also known as hybrid password guessing, could better capture the cracking capabilities of real attackers. However, the reason why hybrid password guessing can enhance cracking capabilities, and how to effectively integrate multiple heterogeneous password guessing models, are still not well understood. To address these issues, this paper draws inspiration from the concept of multi-view learning. We regard the guess lists generated by various password guessing models as multiple views of the data. Through a comprehensive analysis of these guess lists, we have identified the key reason why hybrid password guessing can enhance the cracking capabilities: integrating more diverse views allows for the coverage of a wider range of heterogeneous password characteristics, and provides more detailed information on effective password distributions. Based on the these findings, we propose a new hybrid password guessing framework, namedGuessFuse.GuessFuseemploys the multi-view subset extraction module and segment splitting selection module to accurately extract and reorganize the effective password from multiple guess lists. Experimental results on six large-scale datasets demonstrate the effectiveness ofGuessFuse. By combining two (resp. five) guess lists,GuessFuseoutperforms its foremost counterparts by an average of 11.00% ~ 59.62% (resp. 4.70% ~ 17.66%) within 107guesses.GuessFusecan effectively improve the cracking success rate under a limited number of guesses, approaching the actual cracking capabilities of attackers. Zhijie Xie, Fan Shi 0003, Min Zhang 0054, Huimin Ma 0004, Huaixi Wang, Zhenhan Li |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | An Intelligent Penetration Testing Method Using Human FeedbackabstractPenetration testing is widely acknowledged as the foremost method for evaluating network security. However, three challenges impede the generation of strategies that align with human expectations. In this article, we present, for the first time, a method based on human feedback to enhance strategy generation. Our approach comprises two components: agent training and decision-making. During agent training, we establish a hierarchical framework to decompose tasks and a knowledge base to offer advice for improving data efficiency. We then impose constraints on the action space to mitigate ineffective exploration. Finally, we train a reward model based on human feedback and fine tune the model guided by this reward model. In decision-making, we process the model output to enhance decision accuracy. We crafted scenarios based on real-world networks, and the results demonstrate the effectiveness of our method in generating penetration testing strategies that align more closely with human intentions. Qianyu Li 0001, Min Zhang 0054, Fan Shi 0003, Yi Shen 0012, Bingyang Guo, Chengxi Xu |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Wolf in Sheep's Clothing: Evaluating Security Risks of the Undelegated Record on DNS Hosting ServicesabstractLeveraging DNS for covert communications is appealing since most networks allow DNS traffic, especially the ones directed toward renowned DNS hosting services. Unfortunately, most DNS hosting services overlook domain ownership verification, enabling miscreants to host undelegated DNS records of a domain they do not own. Consequently, miscreants can conduct covert communication through such undelegated records for whitelisted domains on reputable hosting providers. In this paper, we shed light on the emerging threat posed by undelegated records and demonstrate their exploitation in the wild. To the best of our knowledge, this security risk has not been studied before. Fenglu Zhang, Baojun Liu 0002, Eihal Alowaisheq, Lingyun Ying, Xiang Li 0108, Zaifeng Zhang, Ying Liu 0024, Hai-Xin Duan, Min Zhang 0054 |
IMC | 10 |
| 2023 | AlphaEXP: An Expert System for Identifying Security-Sensitive Kernel Objects
Kaixiang Chen, Chao Zhang 0008, Zulie Pan, Qianyu Li 0001, Siliang Qin, Shenglin Xu, Min Zhang 0054, Yang Li 0215 |
USENIX Security Symposium | 8 |
| 2023 | INNES: An intelligent network penetration testing model based on deep reinforcement learning
Qianyu Li 0001, Min Zhang 0054, Yang Li 0215 |
Appl. Intell. | 4 |
| 2023 | A hierarchical deep reinforcement learning model with expert prior knowledge for intelligent penetration testing
Qianyu Li 0001, Min Zhang 0054, Yi Shen 0012, Yang Li 0215 |
Comput. Secur. | 2 |
| 2023 | Tunter: Assessing Exploitability of Vulnerabilities with Taint-Guided Exploitable States Exploration
Kaixiang Chen, Zulie Pan, Yuwei Li 0002, Qianyu Li 0001, Yang Li 0215, Min Zhang 0054, Chao Zhang 0008 |
Comput. Secur. | 7 |
| 2023 | BD-CVSA: A Broadband Direction Finding Method Based on Constructing Virtual Sparse Arrays
Min Zhang 0054, Cheng Tu, Wenli Zhu, Qianyu Li 0001, Hongjun Wang 0010 |
Signal Process. | 1 |
| 2021 | Mining Centralization of Internet Service Infrastructure in the WildabstractThe last decade has witnessed the rapid evolution of the Internet structure, one of which is centralization, that is, Internet core infrastructure has been constantly transferred into the hands of a few popular market participants. Researchers are trying to measure centrality and analyze its security impact from the perspective of traffic analysis. But the underlying distribution of service providers is still enveloped in mysterious veils. In order to address this problem and assess the security risk associated with such centralization. Firstly, we performed linear regression on the data of each kind of service provider in the Alexa Top 1M domains to study the current underlying distribution of various services for the first time. The results show that Zipf’s law is universal in various service providers’ market share, which proves that Internet service infrastructures are centralized. Secondly, we explored the security impacts of centralized infrastructures on the Internet. we conducted attack simulations on providers. Results show that intentional attacks on core providers can greatly downgrade the performance of the Internet. To make matters worse, the quantitative analysis of the provider’s infrastructures found that a considerable number of provider’s infrastructures have low diversity. In addition, we proposed an algorithm to calculate the dependencies between different types of service providers and carried out an evaluation of our datasets, and found the tendency for different services to depend on each other. Our results indicate that the Internet is facing huge security challenges, because the centralized infrastructure will impair service redundancy, and at the same time, it will also cause dependence between infrastructures, which in turn strengthens its centralization. Bingyang Guo, Fan Shi 0003, Chengxi Xu, Min Zhang 0054, Yang Li 0215 |
MSN | 4 |
| 2020 | Binary File's Visualization and Entropy Features Analysis Combined with Multiple Deep Learning Networks for Malware ClassificationabstractIn recent years, the research on malware variant classification has attracted much more attention. However, there are still many challenges, including the low accuracy of classification of samples of similar malware families, high time, and resource consumption. This paper proposes a new method of malware classification based on multiple visual features of malware and deep learning algorithms. In prior research, visualization techniques and entropy demonstrated exemplary performance in many areas. This paper extracts numerous visual features from the raw bytes and entropy sequence of the malware, which makes it more sensitive to malware samples of similar families and endows it the ability to classify malware variants more accurately. To evaluate the proposed method, this paper conducted a series of experiments on two malware datasets with a total of more than 20,000 samples provided by the Malware Research Lab and Microsoft Research. Through experiments, the method showed its superiority compared with some leading malware visual classification methods, achieving good performance on the accuracy with at least 1% improvement. The accuracy of the method even could reach 99.73% and 99.54%, respectively, on the two datasets. Shuguang Huang, Cheng Huang 0003, Fan Shi 0003, Min Zhang 0054, Zulie Pan |
Secur. Commun. Networks | 5 |
| 2020 | Modified Password Guessing Methods Based on TarGuess-IabstractTarGuess − I is a leading online targeted password guessing model using users’ personally identifiable information (PII) proposed at ACM CCS 2016 by Wang et al. It has attracted widespread attention in password security owing to its superior guessing performance. Yet, after analyzing the users’ vulnerable behaviors of using popular passwords and constructing passwords with users’ PII, we find that this model does not take into account popular passwords, keyboard patterns, and the special strings. The special strings are the strings related to users but do not appear in the users’ demographic information. Thus, we propose TarGuess − I + K P X , a modified password guessing model with three semantic methods, including (1) identifying popular passwords by generating top-300 lists from similar websites, (2) recognizing keyboard patterns by relative position, and (3) catching the special strings by extracting continuous characters from user-generated PII. We conduct a series of evaluations on six large-scale real-world leaked password datasets. The experimental results show that our modified model outperforms TarGuess − I by 2.62% within 100 guesses. Zhijie Xie, Min Zhang 0054, Yuqi Guo 0002, Zhenhan Li, Hongjun Wang 0010 |
Wirel. Commun. Mob. Comput. | 2 |