VLDB 2026 Research / reviewers in the wild / expert
Weiyu Dong
dblp:193/2570
· DBLP profile ↗
15ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0001-6137-7782ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 8 · 7 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rator: detecting fine-grained semantic code clones using tree encoding based on node degrees of freedomabstractAbstract Code clone detection has garnered significant attention across various fields, including code refactoring, plagiarism detection, and software maintenance. Numerous methods have been proposed for detecting code clones; however, while text-based and token-based approaches are scalable, they often fail to consider code semantics and are unable to effectively handle semantic code clones. Although tree-based methods perform well in semantic code clone detection, they are limited by the complex structure of trees, making it challenging to apply to large-scale clone detection. Moreover, these methods struggle to achieve fine-grained semantic code clone detection, lacking the ability to pinpoint specific code blocks within semantic clones. In this paper, we propose Rator , a tree-based code clone detector that combines scalability and fine-grained analysis capabilities while effectively detecting semantic clones. Specifically, we design a tree encoding method based on node degrees of freedom, which can transform complex tree structures into simple vector representations while preserving the structural details of the tree. In this way, we can encode all the subtrees of the abstract syntax tree into separate sets of vectors and derive similar features by calculating the similarity between these vectors. The derived similar features serve dual purposes: firstly, they are employed to train a machine learning-based code clone detector, and then, by analyzing the subtree types corresponding to the feature values, specific clone code blocks can be precisely located, thus achieving fine-grained code clone detection. Experimental results show that Rator outperforms nine state-of-the-art code clone detectors with F1 scores of 0.99 and 0.91 on BigCloneBench and Google Code Jam datasets, respectively. As for scalability, Rator is about 93 times faster than ASTNN , another state-of-the-art tree-based semantic clone detector. Regarding fine-grained detection, Rator correctly identifies the concrete clone block with a Top-3 ranked list. Furthermore, the accuracy of fine-grained detection on the Google Code Jam dataset is up to 100% with a Top-2 ranked list. Rui Lou, Huanwei Wang, Weiyu Dong |
Cybersecur. | 6 |
| 2026 | SABLM-VD: Vulnerability detection with a semantic-aware binary language model
Qinghao Li, Tieming Liu, Wei Liu 0164, Yonghe Tang, Weiyu Dong |
Inf. Softw. Technol. | 6 |
| 2025 | Shortest Printable Shellcode Encoding Algorithm Based on Dynamic Bitwidth Selection
Guoan Liu, Weiyu Dong, Jiaan Liu, Tieming Liu |
ACISP (3) | 3 |
| 2025 | SyzForge: An Automated System Call Specification Generation Process for Efficient Kernel Fuzzing
ZhiZhuo Tang, Weiyu Dong, Tieming Liu |
DIMVA (1) | 3 |
| 2025 | Devmp: A Virtual Instruction Extraction Method for Commercial Code Virtualization ObfuscatorsabstractIn code virtualization deobfuscation, extracting virtual instructions is a crucial first step for reverse-engineering programs protected by virtual machine obfuscation.This process is essential for uncovering concealed malicious code, yet existing methods face significant limitations, such as the inability to resolve virtual branch jumps and support multi-version of specified obfuscators, severely hindering their effectiveness.To address these challenges, we introduce a novel method for virtual instruction extraction based on dynamic binary instrumentation and symbolic execution.We implement this method in Devmp, a prototype system designed to extract virtual instructions and facilitate virtualization deobfuscation.Devmp dynamically generates instruction traces through binary instrumentation and performs offline analysis to partition handler sets based on virtual machine structures and jump rules.Then it employs symbolic execution to derive state expressions for semantic analysis of handlers and extracts virtual instructions with complete semantics.We evaluate Devmp on eight test programs protected by two versions of VMProtect.Experimental results demonstrate that Devmp outperforms state-of-the-art tools like VMP Analysis Plugin and NoVmpy, achieving a 28.49% increase in virtual instruction recognition rate by optimized virtual branch processing and accurately analyzing all extracted virtual instructions through enhanced cross-version applicability.These results indicate that Devmp not only improves the accuracy and completeness of virtual instruction extraction but also provides a robust and versatile solution for analyzing programs obfuscated by commercial code virtualization obfuscators. Shenqianqian Zhang, Weiyu Dong, Jian Lin 0007 |
Internetware | 2 |
| 2025 | Deep neural network modeling attacks on arbiter-PUF-based designsabstractAbstract Physical Unclonable Functions (PUFs) are novel circuit structures that provide hardware security solutions in application areas such as chip design and IoT, due to characteristics of their lightweight, key-free and tamper-resistant. PUFs are not immune to threats like machine learning modeling attacks and side channel modeling attacks. Strong PUFs are susceptible to classical machine learning attacks, however, machine learning’s effectiveness in attacking complex structured strong PUFs is limited, and its efficiency is relatively low. Side-channel modeling attacks, on the other hand, incur high implementation costs. Hence, employing deep learning for modeling attacks becomes an effective and cost-efficient choice when attacking complex structured PUFs. In this paper, we introduce a method that employs deep neural network to assess the modeling resilience of combination logic operation-based PUFs with APUFs as components for the first time. We employed a 4-layer DNN model to investigate the security resilience of PUF models involving any combination of OR AND and XOR logical operations. We explored the security regular patterns of modeling resilience. We have demonstrated for the first time that bias in PUF responses can reduce or destroy the security of PUFs. OR or AND logic operations do not provide any security benefit in PUF design, while XOR operations enhance the security of PUFs. Huanwei Wang, Weining Hao, Yonghe Tang, Weiyu Dong, Wei Liu 0164 |
Cybersecur. | 5 |
| 2025 | MLAF-VD: A vulnerability detection model based on multi-level abstract features
Qinghao Li, Wei Liu 0164, Yisen Wang 0002, Weiyu Dong |
J. Inf. Secur. Appl. | 4 |
| 2024 | FA-Fuzz: A Novel Scheduling Scheme Using Firefly Algorithm for Mutation-Based FuzzingabstractMutation-based fuzzing has been widely used in both academia and industry. Recently, researchers observe that the mutation scheduling scheme affects the efficiency of fuzzing. Accordingly, they propose PSO algorithm or machine learning-based technique to optimize the scheduling process. However, these methods fail to consider the fact that the optimal operator distribution of different seeds is different, even for the same program. In this paper, we propose a novel general scheduling scheme, named FA-fuzz, to find the optimal selecting probability distribution of mutation operators, which is based on the observations that the effective mutation operators are different for different seeds. Specifically, our method is based on the firefly algorithm. The positions of fireflies are mapped to the selection probability distribution of different mutation operators. The brightness of fireflies is expressed as the efficiency of discovering unique testcases. We implement prototype systems on multiple state-of-art fuzzers, and perform evaluations on two datasets. Our proposed method improves both the number of unique paths and unique bugs on real-world datasets. In addition, we discover 30 zero-day vulnerabilities in eight real-world programs, which demonstrate the effectiveness of FA-fuzz. Zicong Gao, Weiyu Dong, Yajin Zhou, Liehui Jiang |
IEEE Trans. Software Eng. | 3 |
| 2023 | Dynamic Resampling Based Boosting Random Forest for Network Anomaly Traffic Detection
Huajuan Ren, Weiyu Dong, Yonghe Tang |
IEA/AIE (2) | 3 |
| 2023 | Improvements to code2vec: Generating path vectors using RNNabstractSource code analysis has many application scenarios, such as code plagiarism detection and software vulnerability search. Source code analysis can benefit from machine learning , but it typically requires a standard vector representation and cannot be directly applied to the source code. Thus, we are required to embed source code into vector representation while maintaining the semantics of the code as much as possible. Code2vec proposes a code embedding method that converts source code into code vector through Abstract Syntax Tree(AST). However, we found that code2vec uses a hashing algorithm to generate the identifier for the path in the path context, which leads to the loss of node information in the path and also causes the model training parameters to be very large. Therefore, we present a new path representation which utilizes RNN to generate vectors for paths. We also proposed alternative model designs and evaluated their impact on the model in the experiments. The results we obtained in a challenging source code classification task suggest that, compared to code2vec, the RNN-based paths representation can produce a better embedding model with fewer training parameters. Xuekai Sun, Weiyu Dong, Tieming Liu |
Comput. Secur. | 3 |
| 2023 | BHMDC: A byte and hex n-gram based malware detection and classification method
Yonghe Tang, Xuyan Qi, Jing Jing 0004, Weiyu Dong |
Comput. Secur. | 5 |
| 2023 | DUEN: Dynamic ensemble handling class imbalance in network intrusion detection
Huajuan Ren, Yonghe Tang, Weiyu Dong, Liehui Jiang |
Expert Syst. Appl. | 3 |
| 2022 | CaDeCFF: Compiler-Agnostic Deobfuscator of Control Flow FlatteningabstractWith the increasing influence of malware and various attacks, malware detection methods have been continuously proposed. However, in order to evade malware detection, code obfuscation which makes programs harder to understand, is widely used by malware writers. Control Flow Flattening (CFF) is a common control-flow obfuscation method. However, Control Flow Flattening deobfuscation tools have a low success rate for compilation-optimized binaries because the structural features on which the tools depend have changed. Weiyu Dong, Jian Lin 0007 |
Internetware | 1 |
| 2022 | Fw-fuzz: A code coverage-guided fuzzing framework for network protocols on firmwareabstractSummary Fuzzing is an effective approach to detect software vulnerabilities utilizing changeable generated inputs. However, fuzzing the network protocol on the firmware of IoT devices is limited by inefficiency of test case generation, cross‐architecture instrumentation, and fault detection. In this article, we propose the Fw‐fuzz, a coverage‐guided and crossplatform framework for fuzzing network services running in the context of firmware on embedded architectures, which can generate more valuable test cases by introspecting program runtime information and using a genetic algorithm model. Specifically, we propose novel dynamic instrumentation in Fw‐fuzz to collect the running state of the firmware program. Then Fw‐fuzz adopts a genetic algorithm model to guide the generation of inputs with high code coverage. We fully implement the prototype system of Fw‐fuzz and conduct evaluations on network service programs of various architectures in MIPS, ARM, and PPC. By comparing with the protocol fuzzers Boofuzz and Peach in metrics of edge coverage, our prototype system achieves an average growth of 33.7% and 38.4%, respectively. We further verify six known vulnerabilities and discover 5 0‐day vulnerabilities with the Fw‐fuzz, which prove the validity and utility of our framework. The overhead of our system expressed as an additional 5% of memory growth. Zicong Gao, Weiyu Dong, Yisen Wang 0002 |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | An Effective Authentication for Client Application Using ARM TrustZone
Weiyu Dong, Liehui Jiang, Shuiqiao Yang |
ISPEC | 4 |