Diwei Chen

dblp:316/0428 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0001-6080-9969ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Synthetic Malware at Scale: Malicious Code Generation With Code Transplanting
abstract
Malicious code detection is one of the most essential tasks in safeguarding against security breaches, data compromise, and related threats. While machine learning has emerged as a predominant method for pattern detection, the training process is intricate due to the severe scarcity of malicious code samples. Consequently, machine learning detectors often encounter malicious patterns in limited and isolated scenarios, hindering their ability to generalize effectively across diverse threat landscapes. In this paper, we introduce MalCoder, a novel method for synthesizing malicious code samples. MalCoder enlarges the quantity and diversity of malicious instances by transplanting a set of malicious prototypes into a vast pool of benign code, thereby crafting a diverse array of malicious instances tailored to various application scenarios. For each malware prototype, MalCoder treats it as an incomplete code fragment and crafts its preceding and subsequent contexts through right-to-left and left-to-right code completion respectively. By leveraging GPTs with various sampling strategies, we can instantiate a large number of code samples bearing the malware prototype. Subsequently, MalCoder masks the original prototypes within the transplanted samples and fine-tunes an LLM code generator to reconstruct the original prototype. This process enables the model to seamlessly transplant malicious code fragments into benign code. During inference, MalCoder can automatically insert malicious fragments into benign samples at random positions, transforming benign code into malicious code. We apply MalCoder to a large pool of benign code in CodeSearchNet and craft over 50,000 malicious samples stemming from 39 malicious prototypes. Both qualitative and quantitative analyses show that the generated samples maintain key characteristics of malicious code while blending seamlessly with benign code, which helps in creating realistic and varied training data. Additionally, by using the generated samples as augmented training data, we witness a remarkable surge in malicious code detection capabilities. Specifically, the F1-score experiences a significant increase compared to utilizing only the original prototype samples.
Guangzhan Wang, Diwei Chen, Xiaodong Gu 0002, Yuting Chen 0001, Beijun Shen
IEEE Trans. Software Eng.2
2022 Answering Software Deployment Questions via Neural Machine Reading at Scale
abstract
As software systems continue to grow in complexity and scale, deploying and delivering them becomes increasingly difficult. In this work, we develop DeployQA, a novel QA bot that automatically answers software deployment questions over user manuals and Stack Overflow posts. DeployQA is built upon RoBERTa. To bridge the gap between natural language and the domain of software deployment, we propose three adaptations in terms of vocabulary, pre-training, and fine-tuning, respectively. We evaluate our approach on our constructed DeQuAD dataset. The results show that DeployQA remarkably outperforms baseline methods by leveraging the three domain adaptation strategies.
Guanjie Qiu, Diwei Chen, Yitian Chai, Xiaodong Gu 0002, Beijun Shen
ASE2
2021 ConLAR: Learning to Allocate Resources to Docker Containers under Time-Varying Workloads
abstract
Cloud platforms are increasingly using containers for lightweight virtualization. However, the mainstream operating systems are currently limited in their capabilities in customizing containers’ resource management. There remains two main challenges in resource allocations. First, the application workloads can be time-varying, leading to a problem of resource over- or under-allocations. Second, it becomes difficult to minimize the resource provisioning cost while guaranteeing the SLO (service level objective). To address these challenges, we propose ConLAR, a learning-to-allocate approach that predicts and allocates resources to Docker containers under time-varying workloads. ConLAR efficiently reduces over-provisioning cost with the SLO guarantees by taking an Observing-Predicting-Allocating-Executing paradigm: given a container, it observes the running of the online application and its environment, leverages the LSTM (long short term memory) model to predict its future workload, adaptively learns to construct resource allocation strategies with two objectives through RL (reinforcement learning), and executes them to scale container resources dynamically. We have evaluated ConLAR on two real-world workloads of ClarkNet and GoogleClusterData. The results clearly show the effectiveness of ConLAR. In particular, ConLAR achieves a resource over-provisioning cost of less than 16.5% and an SLO violations rate of 8.9%; it also shows good flexibility to learn different adaption policies.
Diwei Chen, Beijun Shen, Yuting Chen 0001
QRS1