Wenjia Song

dblp:288/0339 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0002-9597-3587ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 How Can ChatGPT Support Human Security Testers to Help Mitigate Supply Chain Attacks?
abstract
Developers often build software on top of third-party libraries (Libs) to improve programmer productivity and software quality. The libraries may contain vulnerabilities exploitable by hackers to attack the applications (Apps) built on top of them. Such attacks are known as software supply chain attacks, the documented number of which has increased 742% in 2022. Researchers and developers created tools to mitigate such attacks, by scanning the library dependencies of Apps, identifying the usage of vulnerable library versions, and suggesting secure alternatives to vulnerable dependencies. However, recent studies show that many developers do not trust the reports by these tools; they need code or evidence to demonstrate how library vulnerabilities lead to security exploits, in order to assess vulnerability severity and modification necessity. Unfortunately, manually crafting demos of application-specific attacks is challenging and timeconsuming, and there is insufficient tool support to automate that procedure.To help developers enhance software security, in this study, we systematically explored the usage of a large language model (LLM)–ChatGPT-4.0–to generate security tests, which unit tests demonstrate how vulnerable library dependencies facilitate the supply chain attacks to givenApps. In our exploration, we defined prompt templates to take in the various vulnerability-relevant information we manually collected, and generated prompts from those templates to query ChatGPT for security test generation. We found that ChatGPT-generated tests demonstrated 24 evidence or proof of vulnerability for 49 Apps. To assess the consistency of test generation, we also evaluated another five state-of-the-art LLMs. All the models generated security tests for at least 17 cases that successfully demonstrate the vulnerabilities. We filed six reports for the newly revealed vulnerabilities in Apps, and got four Common Vulnerability Entries (CVEs) assigned. Our use of ChatGPT outperformed two state-of-the-art security test generators (TRANSFER and SIEGE), by generating a lot more tests and achieving more attacks. Our research will shed light on new research in security test generation.
Ying Zhang 0066, Wenjia Song, Zhengjie Ji, Danfeng Yao, Na Meng 0001
IEEE Trans. Software Eng.2
2024 Madeline: Continuous and Low-cost Monitoring with Graph-free Representations to Combat Cyber Threats
abstract
Advanced persistent threats (APTs) have caused significant financial losses for enterprises, making the development of effective detection systems a critical priority. While existing provenance graph-based APT defenses demonstrate high accuracy, the high complexity and cost of graph operations (e.g., construction, iteration) require extensive processing time and computational resources, making them impractical for real-time detection. To address this challenge, we introduce Madeline, a graph-free, lightweight APT detection system that leverages historical system statistics. Through a multi-step state score calculation for a set of behavioral attributes, MADE-LINE meticulously captures subtle, gradual system changes indicative of stealthy APT activities. Using an LSTM autoencoder, Madeline performs anomaly detection effectively without the need for prior knowledge or manual labeling of attacks. Additionally, Madeline supports continuous monitoring, enhancing the assessment of ongoing risks. This feature helps reduce the need for extensive investigative resources by prioritizing genuine high-risk periods. Our experiments show that Madeline achieves comparable detection accuracy with the state-of-the-art APT detection, with 0.996 recall and 0.011 false positive rate on average. Madeline also significantly reduces the computational overhead, with over 1000x reduction in processing time and 5x in memory utilization.
Wenjia Song, Hailun Ding, Na Meng 0001, Danfeng Yao
ACSAC1
2024 Measurement of Embedding Choices on Cryptographic API Completion Tasks
abstract
In this article, we conduct a measurement study to comprehensively compare the accuracy impacts of multiple embedding options in cryptographic API completion tasks. Embedding is the process of automatically learning vector representations of program elements. Our measurement focuses on design choices of three important aspects, program analysis preprocessing , token-level embedding , and sequence-level embedding . Our findings show that program analysis is necessary even under advanced embedding. The results show 36.20% accuracy improvement, on average, when program analysis preprocessing is applied to transfer bytecode sequences into API dependence paths. With program analysis and the token-level embedding training, the embedding dep2vec improves the task accuracy from 55.80% to 92.04%. Moreover, only a slight accuracy advantage (0.55%, on average) is observed by training the expensive sequence-level embedding compared with the token-level embedding. Our experiments also suggest the differences made by the data. In the cross-app learning setup and a data scarcity scenario, sequence-level embedding is more necessary and results in a more obvious accuracy improvement (5.10%).
Ya Xiao 0002, Wenjia Song, Salman Ahmed 0001, Xinyang Ge, Bimal Viswanath, Na Meng 0001, Danfeng Yao
ACM Trans. Softw. Eng. Methodol.2
2023 Specializing Neural Networks for Cryptographic Code Completion Applications
abstract
Similarities between natural languages and programming languages have prompted researchers to apply neural network models to software problems, such as code generation and repair. However, program-specific characteristics pose unique prediction challenges that require the design of new and specialized neural network solutions. In this work, we identify new prediction challenges in application programming interface (API) completion tasks and find that existing solutions are unable to capture complex program dependencies in program semantics and structures. We design a new neural network model Multi-HyLSTM to overcome the newly identified challenges and comprehend complex dependencies between API calls. Our neural network is empowered with a specialized dataflow analysis to extract multiple global API dependence paths for neural network predictions. We evaluate Multi-HyLSTM on 64,478 Android Apps and predict 774,460 Java cryptographic API calls that are usually challenging for developers to use correctly. Our Multi-HyLSTM achieves an excellent top-1 API completion accuracy at 98.99%. Moreover, we show the effectiveness of our design choices through an ablation study and have released our dataset.
Ya Xiao 0002, Wenjia Song, Jingyuan Qi, Bimal Viswanath, Patrick D. McDaniel, Danfeng Yao
IEEE Trans. Software Eng.2