EDBT 2026 Demo / reviewers in the wild / expert
Zulie Pan
dblp:270/5744
· DBLP profile ↗
24ranked-venue papers
1as first author
23since 2021 · last 2026
0000-0001-5775-5824ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 13 · 12 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Demystifying the Access Control Mechanism of ESXi VMKernel
Zexiang Zhang, Jiaxun Zhu, Jiaqing Huang, Wenbo Shen, Yuliang Lu, Min Zhang 0054, Zulie Pan |
NDSS | 10 |
| 2026 | Software defect detection using large language models: a literature reviewabstractAbstract As software systems grow in complexity, the importance of efficient defect detection escalates, becoming vital to maintain software quality. In recent years, artificial intelligence technology has boomed. In particular, with the proposal of Large Language Models (LLMs), researchers have found the huge potential of LLMs to enhance the performance of software defect detection. This review aims to elucidate the relationship between LLMs and software defect detection. We categorize and summarize existing research based on the distinct applications of LLMs in dynamic and static detection scenarios. Dynamic detection methods are categorized based on the different phases in which they employ LLMs, such as using them for test case generation, providing feedback guidance, and conducting output assessment. Static detection methods are classified according to whether they analyze the source code or the binary of the software under test. Furthermore, we investigate the prompt engineering and model fine-tuning strategies adopted within these studies. Finally, we summarize the emerging trend of integrating LLMs into software defect detection, identify challenges to be addressed and prospect for some potential research directions. Yu Chen 0053, Yi Shen 0012, Taiyan Wang, Shiwen Ou, Yuwei Li 0002, Zulie Pan |
Frontiers Comput. Sci. | 7 |
| 2026 | Unreachable Features? Exposing the Security Risks of Invisible Interfaces in Embedded Web Services of IoT DevicesabstractIoT devices, now integral to our daily routines, offer unparalleled convenience but also face mounting security threats. Embedded web services, prevalent in public networks, pose a major risk to these devices. While research has focused on detecting vulnerabilities in IoT embedded web services, it has overlooked the presence of invisible interfaces, which have emerged as significant security threats. In this paper, we propose InvRadar, a novel framework for detecting vulnerabilities in invisible interfaces of embedded web services in IoT devices. Specifically, InvRadar identifies invisible interfaces by analyzing the differences between the front-end visible interface keywords and the back-end interface keywords through a correlation analysis method. Subsequently, InvRadar uses a static taint analysis method to detect the vulnerabilities that can be triggered by the invisible interfaces. To validate the performance of InvRadar, we conduct extensive experiments and compare InvRadar with the state-of-the-art methods. In testing 13 device firmware, InvRadar identifies 1,793 invisible interfaces and detects 124 vulnerabilities, including 53 newly discovered ones, with 34 receiving new CVE/CNVD IDs. Additionally, InvRadar outperforms the state-of-the-art methods in interface keyword extraction, border binary and data ingestion function identification. Yuanchao Chen, Yuwei Li 0002, Yi Shen 0012, Yu Chen 0053, Yang Li 0215, Taiyan Wang, Yuliang Lu, Zulie Pan, Shouling Ji |
IEEE Internet Things J. | 8 |
| 2026 | Intelligent Penetration Testing Through Integrated Knowledge Graph and Historical Decision EnhancementabstractPenetration Testing (PT), a key network security assessment technique that simulates real cyber attacks to identify vulnerabilities, is traditionally manual and expert-dependent, leading to low efficiency and high costs. Automating and intelligentizing PT has thus become a critical research focus, yet current technologies face two core challenges: lack of standardized, reusable simulated network scenarios (hindering unified experiments and result comparison) and intelligent models' failure to integrate historical decision experience or utilize attack chain temporal correlations (restricting adaptability). To address these, this study proposes an intelligent PT method integrating knowledge graph-driven automated scenario construction and historical decision enhancement. Two innovations are introduced: a network knowledge graph-based mechanism to generate standardized, real-characteristic testing environments; and a historical decision enhancement scheme with a collaborative state temporal processing and action filtering architecture. Experimental results show the method reduces average iterations by 69%, eliminates redundant executions, and enhances decision rationality, offering a new path for automated PT advancement. Qianyu Li 0001, Anupam Chattopadhyay, Cheng Tu, Fan Shi 0003, Min Zhang 0054, Zulie Pan |
IEEE Trans. Dependable Secur. Comput. | 9 |
| 2026 | MirrorFuzz: Leveraging LLM and Shared Bugs for Deep Learning Framework APIs FuzzingabstractDeep learning (DL) frameworks serve as the backbone for a wide range of artificial intelligence applications. However, bugs within DL frameworks can cascade into critical issues in higher-level applications, jeopardizing reliability and security. While numerous techniques have been proposed to detect bugs in DL frameworks, research exploring common API patterns across frameworks and the potential risks they entail remains limited. Notably, many DL frameworks expose similar APIs with overlapping input parameters and functionalities, rendering them vulnerable to shared bugs, where a flaw in one API may extend to analogous APIs in other frameworks. To address this challenge, we propose MirrorFuzz, an automated API fuzzing solution to discover shared bugs in DL frameworks. MirrorFuzz operates in three stages: First, MirrorFuzz collects historical bug data for each API within a DL framework to identify potentially buggy APIs. Second, it matches each buggy API in a specific framework with similar APIs within and across other DL frameworks. Third, it employs large language models (LLMs) to synthesize code for the API under test, leveraging the historical bug data of similar APIs to trigger analogous bugs across APIs. We implement MirrorFuzz and evaluate it on four popular DL frameworks (TensorFlow, PyTorch, OneFlow, and Jittor). Extensive evaluation demonstrates that MirrorFuzz improves code coverage by 39.92% and 98.20% compared to state-of-the-art methods on TensorFlow and PyTorch, respectively. Moreover, MirrorFuzz discovers 315 bugs, 262 of which are newly found, and 80 bugs are fixed, with 52 of these bugs assigned CNVD IDs. Shiwen Ou, Yuwei Li 0002, Chengkun Wei, Tingke Wen, Qiangpu Chen, Yu Chen 0053, Haizhi Tang, Zulie Pan |
IEEE Trans. Software Eng. | 9 |
| 2025 | SeqFuzz: Efficient Kernel Directed Fuzzing via Effective Component Inference
Yuwei Li 0002, Tingke Wen, Huimin Ma 0004, Zulie Pan |
Inscrypt (3) | 9 |
| 2025 | Insvdf: Interface-State-Aware Virtual Device FuzzingabstractHypervisor is the core technology of virtualization for emulating independent hardware resources for each virtual machine. Virtual devices serve as the main interface of the hypervisor, making the security of virtual devices crucial, as any vulnerabilities can impact the entire virtualization environment and pose a threat to the host machine's security. Direct Memory Access (DMA) is the interface of virtual devices, enabling communication with the host machine. Recently, many efforts have focused on fuzzing against DMA to discover the hypervisor's vulnerabilities. However, the lack of sensitivity to the DMA state causes these efforts to be hindered in efficiency during fuzzing. Specifically, there are two main issues: the uncertain interaction moment and the unclear interaction depth. In this paper, we introduce InSVDF, a DMA interface stateaware fuzzing engine. InSVDF first models the intra-interface state of the DMA interface and incorporates an asynchronyaware state snapshot mechanism along with a depth-aware seed preservation mechanism. To validate our approach, we compare InSVDF with a state-of-the-art fuzzer. The results demonstrate that InSVDF significantly enhances vulnerability discovery speed, with improvements of up to 24.2 x in the best case. Furthermore, InSVDF has identified 2 new vulnerabilities, one of which has been assigned a CVE ID. Zexiang Zhang, Yiming Tao, Zulie Pan, Cheng Tu, Min Zhang 0054, Yang Li 0215, Yi Shen 0012, Chunming Wu 0001 |
ICSE | 5 |
| 2025 | DMut: Optimize Mutation Strategy in Directed Greybox Fuzzing by Multi-Population Genetic AlgorithmabstractDirected greybox fuzzing has become a crucial technique for discovering vulnerabilities in software. The seed mutation plays an important role in fuzzing by generating new inputs that explore diverse program states and find the target vulnerability. While seed mutation is critical to the effectiveness of fuzzing, most existing mutation strategies are designed for coverage-based fuzzing and lack the guidance required in directed scenarios. This limits the quality of generated testcases and reduces fuzzing efficiency in directed greybox fuzzing.In this paper, we propose DMut, a novel seed mutation strategy based on a multi-population genetic algorithm, designed to address these limitations. DMut models the seed mutation process using a genetic algorithm, optimizing the seed mutation probability distribution and iteratively evolving it to generate higher-quality testcases. The approach incorporates a well-designed fitness function and selection strategy that aligns with directed fuzzing scenarios to guide the evolution of the mutation strategy. Through comprehensive experiments on real-world CVEs, we demonstrate that DMut significantly improves the effectiveness of directed fuzzing. Compared to the widely adopted directed greybox fuzzing tool, AFLGo, DMut reduces the time to expose the target vulnerability by 41% on average. Additionally, DMut improves path exploration efficiency, covering more unique execution paths and speeding up the exploration process. In summary, the experimental results show that DMut provides a robust, efficient method for improving directed fuzzing performance, offering a significant advancement over existing approaches. Tingke Wen, Yuwei Li 0002, Huimin Ma 0004, Yang Li 0215, Zulie Pan |
SMC | 7 |
| 2025 | Fusing Multimodal Binary Code Representations for Enhanced Similarity DetectionabstractAs software reuse has become increasingly prevalent in the modern era, binary code similarity detection plays a critical role in program analysis. Numerous machine learning methods have been introduced to this field, with the primary challenge lying in the effective representation of binary code. While existing approaches leverage multimodal features for embedding, they still adopt relatively simple fusion methods to combine different modalities. Therefore, their performance can be limited by inadequate modality interaction during the feature fusion step and an over-reliance on a single modality in the final embedding step. To address these issues, we propose FuseBinRepr, a method that fuses text modality and graph modality representation techniques using a fusion model architecture and three specialized learning tasks to enhance binary code similarity detection. We design a fusion architecture integrating both text and graph embeddings via cross-attention and self-attention mechanisms. For model training, we developed three tasks: text-graph alignment (TGA), graph masking recovery (GMR), and contrastive learning (CL), to capture and align high-level semantics across modalities. Through evaluation, FuseBinRepr demonstrates improvements in Mean Reciprocal Rank (MRR) and Recall compared to state-of-the-art baseline methods, achieving increases of up to 40.8 % and 42 %, respectively. Ablation studies confirm the robustness of our model design, as well as the effectiveness of pretraining tasks. In the downstream software vulnerability detection task, FuseBinRepr achieves the best MRR and Recall performance in ranking CVE binary functions. Taiyan Wang, Yu Chen 0053, Zulie Pan, Min Zhang 0054 |
SRDS | 5 |
| 2025 | A survey of binary code representation technologyabstractBinary analysis, as an important foundational technology, provides support for numerous applications in the fields of software engineering and security research. With the continuous expansion of software scale and the complex evolution of software architecture, binary analysis technology is facing new challenges. To break through existing bottlenecks, researchers have applied artificial intelligence (AI) technology to the understanding and analysis of binary code. The core lies in characterizing binary code, i.e., how to use intelligent methods to generate representation vectors containing semantic information for binary code, and apply them to multiple downstream tasks of binary analysis. In this paper, we provide a comprehensive survey of recent advances in binary code representation technology, and introduce the workflow of existing research in two parts, i.e., binary code feature selection methods and binary code feature embedding methods. The feature selection section includes mainly two parts: definition and classification of features, and feature construction. First, the abstract definition and classification of features are systematically explained, and second, the process of constructing specific representations of features is introduced in detail. In the feature embedding section, based on the different intelligent semantic understanding models used, the embedding methods are classified into four categories based on the usage of text-embedding models and graph-embedding models. Finally, we summarize the overall development of existing research and provide prospects for some potential research directions related to binary code representation technology. Taiyan Wang, Qingsong Xie, Zulie Pan, Min Zhang 0054 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2025 | Whiskey: Large-Scale Identification of Mobile Mini-App Session Key Leakage With LLMsabstractMini-apps, which run on super-apps, have attracted a large number of users due to their lightweight nature and the convenience of supporting the authorized use of super-app user information. Super-apps employ encryption to protect the transmission of sensitive identity information authorized by users to the mini-app, using the session key as the key. However, we have identified a risk of session key leakage, which could be exploited to maliciously manipulate sensitive user identity information, thereby posing a significant threat to user data security. To reveal this damage, we explore potential business scenarios of session key leakage in detail. Nevertheless, the diversity in design among various mini-apps makes automated testing of these business scenarios at a large scale challenging. This diversity is reflected in the inconsistent naming of identical types of controls and the disparate execution orders of controls within the same business scenarios across different mini-apps. To overcome these challenges, we propose Whiskey, which can adaptively and intelligently optimize dynamic testing strategies for mini-apps with diverse designs using large language models to detect session key leakage at scale. We evaluated Whiskey on 157,063 WeChat mini-apps and 10,000 TikTok mini-apps, and found that 15,712 of WeChat mini-apps and 678 of TikTok mini-apps had session key leakage vulnerabilities. Further analysis showed that this leakage could lead to account takeover and promotion abuse attacks. We responsibly reported the detection results to Tencent and the mini-app vendors. At the time of submission, 17 reported issues had been assigned CNVD IDs. Yu Chen 0053, Yuanchao Chen, Taiyan Wang, Shouling Ji, Hong Shan, Zulie Pan |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2025 | Understanding the Security Risks of Websites Using Cloud Storage for Direct User File UploadsabstractWith the rising demand for website data storage, leveraging cloud storage services for vast user file storage has become prevalent. Nowadays, a new file upload scenario has been introduced, allowing web users to upload files directly to the cloud storage service. This new scenario offers convenience but involves more roles (i.e., web users, web servers, and cloud storage services) and their interactions, bringing new security threats. In this paper, we perform the first systematic security study in this scenario. With in-depth analysis, we identify six new types of vulnerabilities and conduct large-scale real-world measurements on the top 500 Alexa Rank websites. Among these websites, 182 (36.4%) use cloud storage services, illustrating the widespread use of the cloud. Then, we perform a detailed analysis of 28 popular websites that allow user upload. Surprisingly, they all have at least one of the six vulnerabilities. Totally, we discover 79 new vulnerabilities and responsibly report them to the websites. Many popular websites respond positively, including Google, Reddit, and CSDN. We discuss the root causes of these vulnerabilities and propose possible mitigation methods. In summary, our work offers significant value in understanding the security risks of cloud storage services for websites and facilitating future research. Yuanchao Chen, Yuwei Li 0002, Yuliang Lu, Zulie Pan, Shouling Ji, Yu Chen 0053, Yang Li 0103, Yi Shen 0012 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Enhancing Black-box Compiler Option Fuzzing with LLM through Command FeedbackabstractSince the compiler acts as a core component in software building, it is essential to ensure its availability and reliability through software testing and security analysis. Most research has focused on compiler robustness when compiling various test cases, while the reliability of compiler options lacks attention, especially since each option can activate a specific compiler function. Although some researchers have made efforts in testing it, the insufficient utilization of compiler command feedback messages leads to the poor efficiency, which hinders more diverse and in-depth testing.In this paper, we propose a novel solution to enhance black-box compiler option fuzzing by utilizing command feedback, such as error messages, standard output and compiled files, to guide the error fixing and option pruning via prompting large language models for suggestions. We have implemented the prototype and evaluated it on 4 versions of LLVM. Experiments show that our method significantly improves the detection of crashes, reduces false negatives, and even increase the success rate of compilation when compared to the baseline. To date, our method has identified hundreds of unique bugs, and 9 of them are previously unknown. Among these, 8 have been assigned CVE numbers, and 1 has been fixed following our report. Taiyan Wang, Yu Chen 0053, Zulie Pan, Min Zhang 0054, Huimin Ma 0004, Jinghua Zheng |
ISSRE | 5 |
| 2024 | An Empirical Study on the Distance Metric in Guiding Directed Grey-box FuzzingabstractDirected grey-box fuzzing (DGF) aims to discover vulnerabilities in specific code areas efficiently. Distance metric, which is used to measure the quality of seed in DGF, is a crucial factor in affecting the fuzzing performance. Despite distance metrics being widely applied in existing DGF frameworks, it remains opaque about how different distance metrics guide the fuzzing process and affect the fuzzing result in practice. In this paper, we conduct the first empirical study to explore how different distance metrics perform in guiding DGFs. Specifically, we systematically discuss different distance metrics in the aspect of calculation method and granularity. Then, we implement different distance metrics based on AFLGo. On this basis, we conduct comprehensive experiments to evaluate the performance of these distance metrics on the benchmarks widely used in existing DGF-related work. The experimental results demonstrate the following insights. First, the difference among different distance metrics with varying methods of calculation and granularities is not significant. Second, the distance metrics may not be effective in describing the difficulty of triggering the target vulnerability. In addition, by scrutinizing the quality of testcases, our research highlights the inherent limitation of existing mutation strategies in generating high-quality testcases, calling for designing effective mutation strategies for directed fuzzing. We open-source the implementation code and experiment dataset to facilitate future research in DGF. Tingke Wen, Yuwei Li 0002, Huimin Ma 0004, Zulie Pan |
ISSRE | 5 |
| 2024 | SyzLego: Enhancing Kernel Directed Greybox Fuzzing via Dependency Inference and Scheduling
Chengxiang Liao, Juxing Chen, Zulie Pan |
ISC (1) | 6 |
| 2024 | G-Fuzz: A Directed Fuzzing Framework for gVisorabstractgVisor is a Google-published application-level kernel for containers. As gVisor is lightweight and has sound isolation, it has been widely used in many IT enterprises [1],[2],[3]. When a new vulnerability of the upstream gVisor is found, it is important for the downstream developers to test the corresponding code to maintain the security. To achieve this aim, directed fuzzing is promising. Nevertheless, there are many challenges in applying existing directed fuzzing methods for gVisor. The core reason is that existing directed fuzzers are mainly for general C/C++ applications, while gVisor is an OS kernel written in the Go language. To address the above challenges, we propose G-Fuzz, a directed fuzzing framework for gVisor. There are three core methods in G-Fuzz, including lightweight and fine-grained distance calculation, target related syscall inference and utilization, and exploration and exploitation dynamic switch. Note that the methods of G-Fuzz are general and can be transferred to other OS kernels. We conduct extensive experiments to evaluate the performance of G-Fuzz. Compared to Syzkaller, the state-of-the-art kernel fuzzer, G-Fuzz outperforms it significantly. Furthermore, we have rigorously evaluated the importance for each core method of G-Fuzz. G-Fuzz has been deployed in industry and has detected multiple serious vulnerabilities. Yuwei Li 0002, Shouling Ji, Xuhong Zhang 0002, Guanglu Yan, Alex X. Liu, Chunming Wu 0001, Zulie Pan |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2024 | URadar: Discovering Unrestricted File Upload Vulnerabilities via Adaptive Dynamic TestingabstractUnrestricted file upload (UFU) vulnerabilities, especially unrestricted executable file upload (UEFU) vulnerabilities, pose severe security risks to web servers. For instance, attackers can leverage such vulnerabilities to execute arbitrary code to gain the control of a whole web server. Therefore, it is significant to develop effective and efficient methods to detect UFU and UEFU vulnerabilities. Towards this, most state-of-the-art methods are designed based on dynamic testing. Nevertheless, they still entail two critical limitations. 1) They heavily rely on manual efforts, which are error-prone and have poor adaptability. 2) They seldom leverage effective information to guide the testing, resulting in generating a large number of invalid test cases. Such limitations severely hinder the performance of UFU vulnerability detection. In this paper, we propose URadar, an adaptive dynamic testing-based method for detecting UFU and UEFU vulnerabilities. There are three core designs in URadar, including file upload interface identification, file type restriction inference, and invalid mutation combination filtration, which can effectively solve the two limitations of existing methods. To evaluate the performance of URadar, we conduct extensive experiments and compare URadar with state-of-the-art methods (e.g., FUSE, RIPS). In testing 18 web applications, URadar discovers 26 UEFU vulnerabilities, where 8 are new, and 6 have been assigned new CVE/CNNVD IDs. By contrast, FUSE and RIPS find 14 and 2 UEFU vulnerabilities, respectively. To discover the same number of UFU vulnerabilities, FUSE needs to send 73,261 request packets with a time cost of 2,791.1s on average, 23.43 and 20.53 times of the requirements for URadar. The above results demonstrate that URadar significantly outperforms the state-of-the-art methods. In addition, we have open-sourced URadar to facilitate future research on UFU vulnerability detection. Yuanchao Chen, Yuwei Li 0002, Zulie Pan, Yuliang Lu, Juxing Chen, Shouling Ji |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Optir-SBERT: Cross-Architecture Binary Code Similarity Detection Based on Optimized LLVM IR
Yintong Yan, Taiyan Wang, Zulie Pan |
ICDF2C (2) | 5 |
| 2023 | AlphaEXP: An Expert System for Identifying Security-Sensitive Kernel Objects
Kaixiang Chen, Chao Zhang 0008, Zulie Pan, Qianyu Li 0001, Siliang Qin, Shenglin Xu, Min Zhang 0054, Yang Li 0215 |
USENIX Security Symposium | 4 |
| 2023 | Tunter: Assessing Exploitability of Vulnerabilities with Taint-Guided Exploitable States Exploration
Kaixiang Chen, Zulie Pan, Yuwei Li 0002, Qianyu Li 0001, Yang Li 0215, Min Zhang 0054, Chao Zhang 0008 |
Comput. Secur. | 3 |
| 2022 | Secure User Authentication Leveraging Keystroke Dynamics via Wi-Fi SensingabstractUser authentication plays a critical role in access control of a man-machine system, where the knowledge factor, such as a personal identification number, constitutes the most widely used authentication element. However, knowledge factors are usually vulnerable to the spoofing attack. Recently, the inheritance factor, such as fingerprints, emerges as an efficient alternative resilient to malicious users, but it normally requires special equipment. To this end, in this article, we propose WiPass, a device-free authentication system only leveraging the pervasive Wi-Fi infrastructure to explore keystroke dynamics (manner and rhythm of keystrokes) captured by the channel state information to recognize legitimate users while rejecting spoofers. However, it remains an open challenge to characterize the behavioral features hidden in the human subtle motions, such as keystrokes. Therefore, we build a signal enhancement model using Ricean distribution to amplify user keystroke dynamics and a hybrid learning model for user authentication, which consists of two parts, i.e., convolutional neural network based feature extraction and support vector machine based classification. The former relies on visualizing the channel responses into time-series images to learn the behavioral features of keystrokes in energy and spectrum domains, whereas the latter exploits such behavioral features for user authentication. We prototype WiPass on the low-cost off-the-shelf Wi-Fi devices and verify its performance. Empirical results show that WiPass achieves on average 92.1% authentication accuracy, 5.9% false accept rate, and 6.3% false reject rate in three real environments. Yu Gu 0003, Yantong Wang, Meng Wang 0037, Zulie Pan, Zhi Liu 0002, Mianxiong Dong |
IEEE Trans. Ind. Informatics | 4 |
| 2021 | Multiparty verification in image secret sharingabstractMultiparties in image secret sharing (ISS) need to verify (detect and recognize) each other, which is seldom considered and realized in traditional methods. In this paper, we introduce the definition of multiparty verification. It includes two stages, i.e., a detection stage and a recognition stage, with evaluation methods that are also discussed. A multiparty verification scheme without pixel expansion is developed, which is suitable for both dealer attendance and nonattendance. The classic hash function, public key cryptography and visual cryptography are technically fused in the developed scheme. In the shadow distribution phase, each participant can verify the received shadow using his private key. In the restoration phase, for the case of dealer attendance, he can verify each shadow received using his secret key; for the case of dealer nonattendance, participants can verify each other before exchanging their shadows. We conduct analyses and illustrations to validate the developed scheme. Xuehu Yan, Zulie Pan, Xiaofeng Zhong, Guozheng Yang |
Inf. Sci. | 3 |
| 2021 | Webshell Detection Based on Executable Data Characteristics of PHP CodeabstractA webshell is a malicious backdoor that allows remote access and control to a web server by executing arbitrary commands. The wide use of obfuscation and encryption technologies has greatly increased the difficulty of webshell detection. To this end, we propose a novel webshell detection model leveraging the grammatical features extracted from the PHP code. The key idea is to combine the executable data characteristics of the PHP code with static text features for webshell classification. To verify the proposed model, we construct a cleaned data set of webshell consisting of 2,917 samples from 17 webshell collection projects and conduct extensive experiments. We have designed three sets of controlled experiments, the results of which show that the accuracy of the three algorithms has reached more than 99.40%, the highest reached 99.66%, the recall rate has been increased by at least 1.8%, the most increased by 6.75%, and the F1 value has increased by 2.02% on average. It not only confirms the efficiency of the grammatical features in webshell detection but also shows that our system significantly outperforms several state‐of‐the‐art rivals in terms of detection accuracy and recall rate. Zulie Pan, Yuanchao Chen, Yu Chen 0053, Yi Shen 0012, Xuanzhen Guo |
Wirel. Commun. Mob. Comput. | 1 |
| 2020 | Binary File's Visualization and Entropy Features Analysis Combined with Multiple Deep Learning Networks for Malware ClassificationabstractIn recent years, the research on malware variant classification has attracted much more attention. However, there are still many challenges, including the low accuracy of classification of samples of similar malware families, high time, and resource consumption. This paper proposes a new method of malware classification based on multiple visual features of malware and deep learning algorithms. In prior research, visualization techniques and entropy demonstrated exemplary performance in many areas. This paper extracts numerous visual features from the raw bytes and entropy sequence of the malware, which makes it more sensitive to malware samples of similar families and endows it the ability to classify malware variants more accurately. To evaluate the proposed method, this paper conducted a series of experiments on two malware datasets with a total of more than 20,000 samples provided by the Malware Research Lab and Microsoft Research. Through experiments, the method showed its superiority compared with some leading malware visual classification methods, achieving good performance on the accuracy with at least 1% improvement. The accuracy of the method even could reach 99.73% and 99.54%, respectively, on the two datasets. Shuguang Huang, Cheng Huang 0003, Fan Shi 0003, Min Zhang 0054, Zulie Pan |
Secur. Commun. Networks | 6 |