VLDB 2026 Research / reviewers in the wild / expert
Wei Wang 0428
dblp:35/7092-428
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0003-0373-8863ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CTISum: A new benchmark dataset for Cyber Threat Intelligence summarization
Wei Peng 0008, Junmei Ding, Wei Wang 0428, Lei Cui 0003, Zhiyu Hao, Xiao-chun Yun |
Comput. Secur. | 3 |
| 2025 | Bottom Aggregating, Top Separating: An Aggregator and Separator Network for Encrypted Traffic UnderstandingabstractEncrypted traffic classification refers to the task of identifying the application, service or malware associated with network traffic that is encrypted. Previous methods mainly have two weaknesses. Firstly, from the perspective of word-level (namely, byte-level) semantics, current methods use pre-training language models like BERT, learned general natural language knowledge, to directly process byte-based traffic data. However, understanding traffic data is different from understanding words in natural language, using BERT directly on traffic data could disrupt internal word sense information so as to affect the performance of classification. Secondly, from the perspective of packet-level semantics, current methods mostly implicitly classify traffic using abstractive semantic features learned at the top layer, without further explicitly separating the features into different space of categories, leading to poor feature discriminability. In this paper, we propose a simple but effective Aggregator and Separator Network (ASNet) for encrypted traffic understanding, which consists of two core modules. Specifically, a parameter-free word sense aggregator enables BERT to rapidly adapt to understanding traffic data and keeping the complete word sense without introducing additional model parameters. And a category-constrained semantics separator with task-aware prompts (as the stimulus) is introduced to explicitly conduct feature learning independently in semantic spaces of different categories. Experiments on five datasets across seven tasks demonstrate that our proposed model achieves the current state-of-the-art results without pre-training in both the public benchmark and real-world collected traffic dataset. Statistical analyses and visualization experiments also validate the interpretability of the core modules. Furthermore, what is important is that ASNet does not need pre-training, which dramatically reduces the cost of computing power and time. The model code and dataset will be released inhttps://github.com/pengwei-iie/ASNET. Wei Peng 0008, Lei Cui 0003, Wei Wang 0428, Xiaoyu Cui, Zhiyu Hao, Xiao-chun Yun |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | APIBeh: Learning Behavior Inclination of APIs for Malware ClassificationabstractMalware classification involves categorizing mal-ware samples based on their characteristics. While deep learning techniques applied to malware execution traces, mainly API calls, have shown potential in this field, they still perform poorly. This is primarily because they treat all APIs equally and train classifiers directly on native APIs, which inadequately capture the under-lying family-related semantics. In this paper, we first investigate the behaviors of multiple malware families and observe that different families exhibit divergent behaviors, with each family consistently favoring certain behaviors over time. Motivated by this, we propose APIBeh, a new embedding method designed to enhance malware classification. APIBeh first utilizes Benignity Degree Algorithm to identify and exclude insignificant, likely benign APIs from sequences. Then, it introduces the concept of Behavior Inclination, which quantifies the association between an API and malicious behaviors, facilitating high-level behavior encoding for each API. This Behavior Inclination embedding is then concatenated with raw embedding to represent an API, and fed into a DL model for classifier training. Experimental results show that APIBeh outperforms existing embedding methods in classification performance, e.g., 3.18% boost in weighted f1-score over a recent study using word2vec. In addition, it offers robustness to concept drift and adversarial attacks. Lei Cui 0003, Yiran Zhu, Junnan Yin, Zhiyu Hao, Wei Wang 0428, Peng Liu 0044, Xiao-chun Yun |
ISSRE | 5 |
| 2022 | ClusterRR: a record and replay framework for virtual machine clusterabstractThe Record and Replay (RnR) technology provides the ability to reproduce past execution of systems deterministically. It has many prominent applications, including fault tolerance, security analysis, and failure diagnosis. In system virtualization, previous RnR researches mainly focus on individual VM, including coherent replaying of multi-core systems, reducing performance penalty and storage overhead. However, with the emerging of distributed systems deployed in virtual machine clusters (VMC), the existing RnR technology of individual VM can not meet the requirements of analyzers and developers. The critical challenge for VMC RnR is to maintain the consistency of global state. In this paper, we propose ClusterRR, a RnR framework for VMC. To solve the inconsistency problem, we propose coordination protocols to schedule the record and replay process of VMs. Meanwhile, we employ a Hybrid RnR approach to reduce the performance penalty and storage costs caused by recording network events. Moreover, we implement ClusterRR on QEMU/KVM platform and utilize a network packets retransmission framework to guarantee the reproducibility of VMC replay. Last, we conduct a series of experiments to measure its efficiency and overhead. The results show that ClusterRR would efficiently replay the execution of the whole VMC at instruction-level granularity. Wei Wang 0428, Zhiyu Hao, Lei Cui 0003 |
VEE | 1 |
| 2022 | iConSnap: An Incremental Continuous Snapshots System for Virtual MachinesabstractThe reliability of data and services hosted on a virtual machine (VM) is a top concern in cloud environments. The Continuous Snapshots can reduce the data loss in case of failures and thus is prevailing for protecting long-running systems. However, existing methods suffer from long VM downtime, long snapshot interval and significant performance loss. In this article, we present iConSnap, a system designed to take fine-grained continuous snapshots of virtual machines without compromising VM performance. First, iConSnap adopts the copy-on-write (COW) mechanism to save the memory pages on-demand, and thus decreases the VM downtime to about 200 milliseconds. Second, we extend the idea of COW and propose a lazily incremental approach to save the delta data between two successive snapshots only once, thereby reducing the snapshot duration and snapshot data a lot. Third, we propose a scheduling mechanism to mitigate the VM performance penalty issue. Last, we introduce a method combined of compression and time-aware multi-granularity reclamation strategy to reduce the storage costs without losing performance and availability. We implement iConSnap on QEMU/KVM and evaluate it through a set of experiments. The experimental results show that iConSnap outperforms existing approaches in terms of VM downtime, snapshot duration, storage costs and VM performance. Zhiyu Hao, Wei Wang 0428, Lei Cui 0003, Xiao-chun Yun, Zhenquan Ding |
IEEE Trans. Serv. Comput. | 2 |
| 2020 | pRnR: A Parallel Record-Replay Framework for Virtual MachinesabstractThe record and replay(RnR) technology of virtual machine(VM) provides the ability to reproduce the past execution of a VM deterministically. It has many promising applications in the cloud environment, including fault tolerance, security analysis, and failure diagnosis. Existing studies in this area pay more effort in optimizing the record method, such as reducing performance penalty and storage costs. However, considering that many practical applications follow the record once, replay many mode, the optimization for the replay is more critical, especially for efficiency. In this paper, we propose pRnR, a novel parallel RnR framework, to support efficient replay. By combining the native RnR framework with an improved continuous snapshots mechanism, pRnR divides the full execution into many independent and complete slices, each of which supports arbitrary replay. In addition, it supports two replay modes to improve replay efficiency, i.e., multi-slice parallel replay and multi-dimension parallel replay. Moreover, we apply our pRnR framework to syscall-based diagnosis to demonstrate its usability. The experimental results show that pRnR is more efficient than existing RnR frameworks. Wei Wang 0428, Lei Cui 0003, Zhiyu Hao, Haiqiang Fei, Chonghua Wang, Yaqiong Peng |
ICCD | 1 |