Guo-Wei Wong

dblp:339/7712 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-1748-7818ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SAGA: Synthetic Audit Log Generation for APT Campaigns
abstract
With the increasing sophistication of Advanced Persistent Threats (APTs), the demand for effective detection and mitigation strategies and methods has escalated. Program execution leaves traces in the system audit log, which can be analyzed to detect malicious activities. However, collecting and analyzing large volumes of audit logs over extended periods is challenging, further compounded by insufficient labeling that hinders their usability. Addressing these challenges, this paper introduces SAGA (Synthetic Audit log Generation for APT campaigns), a novel approach for generating find-grained labeled synthetic audit logs that mimic real-world system logs while embedding stealthy APT attacks. SAGA generates configurable audit logs for arbitrary duration, blending benign logs from normal operations with malicious logs based on the definitions the MITRE ATT&CK framework. Malicious audit logs follow an APT lifecycle, incorporating various attack techniques at each stage. These synthetic logs can serve as benchmark datasets for training machine learning models and assessing diverse APT detection methods. To demonstrate the usefulness of synthetic audit logs, we ran established baselines of event-based technique hunting and APT campaign detection using various synthetic audit logs. In addition, we show that a deep learning model trained on synthetic audit logs can detect previously unseen techniques within audit logs.
Yi-Ting Huang, Ying-Ren Guo, Yu-Sheng Yang, Guo-Wei Wong, Yu-Zih Jheng, Yeali S. Sun, Jessemyn Modini, Timothy Lynar, Meng Chang Chen
IEEE Trans. Dependable Secur. Comput.4
2026 Resilient Dynamic Analysis for Windows Malware Technique Discovery against Behavior Obfuscation
abstract
In this article, we focus on the robustness of behavior-based malware analysis models, justified by the need to address the high mutation rates of malware executables that debilitate conventional signature-based approaches and even behavior-based AI solutions. In response to these challenges, we propose MAMBA + , an obfuscation-resistant dynamic analysis approach tailored for uncovering malware behavior. We have assembled a comprehensive collection of behavioral obfuscation attacks designed to undermine behavior-based models. The central concept behind MAMBA + involves treating obfuscated calls as perturbed data and introducing a novel loss function to effectively balance ground-truth predictions and the handling of these perturbations. To facilitate this approach, MAMBA + designs adapted embedding mechanisms to transform traces of API calls into high-dimensional vectors for attention calculations. Through a comprehensive empirical study with seven obfuscations and three unseen attacks, we reveal important qualitative properties of MAMBA + , and quantitatively demonstrate its superiority in performance and robustness to all compared methods.
Yi-Ting Huang, Lisa Liu, Ying-Ren Guo, Guo-Wei Wong, Timothy Lynar, Meng Chang Chen
ACM Trans. Priv. Secur.4
2025 Poster: When Logs Misbehave: Retrieving Known APTs from Noisy Graphs
abstract
The task of retrieving known Advanced Persistent Threat (APT) campaigns from system activity graphs, where nodes represent MITRE ATT&CK techniques and edges encode temporal or resource-level relationships, requires reasoning over structures. In operational settings, these target graphs are often noisy due to incomplete detection, technique misclassification, and benign-induced structural artifacts. To address this issue, we formulate the task as approximate subgraph matching between a known APT query graph and a noisy, partially observed technique graph. In this poster, we introduce a preliminary embedding-based retrieval method, aiming to promote it as a robust and practical framework for retrieving known APTs in real-world environments.
Guo-Wei Wong, Yi-Ting Huang, Ying-Ren Guo, Shou-De Lin, Wang-Chien Lee, Meng Chang Chen
CCS1
2025 Poster: LogCraft: Crafting CVE-Aware Synthetic Worlds (Logs)
Kai-Xian Wong, Chan-Jien Tan, Yi-Ting Huang, Ying-Ren Guo, Yu-Zih Jheng, Guo-Wei Wong, Meng Chang Chen
CCS6
2025 A Cascade Approach for APT Campaign Attribution in System Event Logs: Technique Hunting and Subgraph Matching
abstract
As Advanced Persistent Threats (APTs) grow increasingly sophisticated, the demand for effective detection methods has intensified. This study addresses the challenge of identifying APT campaign attacks through system event logs. A cascading approach, name SFM, combines Technique hunting and APT campaign attribution. The approach assumes that real-world system event logs contain a vast majority of normal events interspersed with few suspiciously malicious ones and that the logs are annotated with Techniques of MITRE ATT&CK framework for attack pattern recognition. After identifying Techniques from the log, we attribute APT campaign attacks by aligning detected Techniques with known attack sequences to determine the most likely APT campaign. Evaluations on five synthetic real-world APT campaigns indicate that the proposed approach demonstrates reliable performance.
Yi-Ting Huang, Ying-Ren Guo, Guo-Wei Wong, Meng Chang Chen
ICC3
2024 Attention-Based API Locating for Malware Techniques
abstract
This paper presents APILI, an innovative approach to behavior-based malware analysis that utilizes deep learning to locate the API calls corresponding to discovered malware techniques in dynamic execution traces. APILI defines multiple attentions between API calls, resources, and techniques, incorporating MITRE ATT&CK framework, adversary tactics, techniques and procedures, through a neural network. We employ fine-tuned BERT for arguments/resources embedding, SVD for technique representation, and several design enhancements, including layer structure and noise addition, to improve the locating performance. To the best of our knowledge, this is the first attempt to locate low-level API calls that correspond to high-level malicious behaviors (that is, techniques). Our evaluation demonstrates that APILI outperforms other traditional and machine learning techniques in both technique discovery and API locating. These results indicate the promising performance of APILI, thus allowing it to reduce the analysis workload.
Guo-Wei Wong, Yi-Ting Huang, Ying-Ren Guo, Yeali S. Sun, Meng Chang Chen
IEEE Trans. Inf. Forensics Secur.1
2024 Sparse Grid Imputation Using Unpaired Imprecise Auxiliary Data: Theory and Application to PM2.5 Estimation
abstract
Sparse grid imputation (SGI) is a challenging problem, as its goal is to infer the values of the entire grid from a limited number of cells with values. Traditionally, the problem is solved using regression methods such as KNN and kriging, whereas in the real world, there is often extra information—usually imprecise—that can aid inference and yield better performance. In the SGI problem, in addition to the limited number of fixed grid cells with precise target domain values, there are contextual data and imprecise observations over the whole grid. To solve this problem, we propose a distribution estimation theory for the whole grid and realize the theory via the composition architecture of the Target-Embedding and the Contextual CycleGAN trained with contextual information and imprecise observations. Contextual CycleGAN is structured as two generator–discriminator pairs and uses different types of contextual loss to guide the training. We consider the real-world problem of fine-grained PM2.5 inference with realistic settings: a few (less than 1%) grid cells with precise PM2.5 data and all grid cells with contextual information concerning weather and imprecise observations from satellites and microsensors. The task is to infer reasonable values for all grid cells. As there is no ground truth for empty cells, out-of-sample mean squared error and Jensen–Shannon divergence measurements are used in the empirical study. The results show that Contextual CycleGAN supports the proposed theory and outperforms the methods used for comparison.
Guo-Wei Wong, Meng Chang Chen
ACM Trans. Knowl. Discov. Data2
2022 Building Cybersecurity Ontology for Understanding and Reasoning Adversary Tactics and Techniques
abstract
Cyber threats have become more prevalent than ever. Cyber Threat Intelligence (CTI) reports and MITRE ATTCK® framework play an imperative role in helping experts and organizations assess current and potential attacks, such as Advanced Persistent Threats (APT). However, the task of extracting valuable information from unstructured texts remains an ongoing challenge. In this work, we present a framework for understanding and reasoning adversary tactics and techniques. We construct an ontology structure and propose an automatic information extraction method that is capable of integrating the parsed information from CTI reports into each instance. The ontology is represented in the Web Ontology Language (OWL) accessible with the SPARQL query language. Our evaluation shows that the proposed information extraction method outperforms other state-of-the-art neural network-based methods in terms of precision. Furthermore, our framework can effectively infer adversary information, which efficiently supports security analysts recognize tactics and techniques.
Chiao-Cheng Huang, Pei-Yu Huang, Ying-Ren Kuo, Guo-Wei Wong, Yi-Ting Huang, Yeali S. Sun, Meng Chang Chen
IEEE Big Data4