Steven H. H. Ding

dblp:161/1183 · DBLP profile ↗
← Back
37ranked-venue papers
4as first author
31since 2021 · last 2026
0000-0003-4513-200XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 8 since 2021Security and privacy · 10 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 9 · 9 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Security risk assessment of android automotive OS software supply chain using firmware reverse engineering
abstract
As Android Automotive OS (AAOS) becomes the in-vehicle platform of choice for infotainment and domain-controller functions in modern passenger cars, its software supply chain has emerged as a critical security frontier. AAOS spans both infotainment and vehicle-control domains within the automotive electronics architecture by supporting media streaming, over-the-air updates, navigation, and sensor fusion. Its open-source foundations and reliance on third-party libraries introduce risks, from outdated components to malicious modules, that can undermine vehicle functionality and passenger safety. In recognition of these threats, ISO/SAE 21434 and UNECE WP.29 R155 mandate structured security assessments for vehicular systems to prevent software-chain vulnerabilities from compromising safety. In this study, we apply a shift-right security analysis via firmware reverse engineering to AAOS images from four leading OEMs. We unpack each firmware image, extract software bills of materials (SBOMs), map Common Vulnerabilities and Exposures (CVE) to components, and characterize system-level attack surfaces across infotainment and control subsystems. Proof-of-concept exploits were developed for high-risk vulnerabilities. One critical CVE was successfully triggered, while others were mitigated by missing dependencies or built-in protections. Our work delivers a reproducible firmware-analysis workflow for automotive supply-chain risk assessment, a comparative survey of third-party and proprietary component management, and the evidence of inconsistent security postures in AAOS-based vehicular electronics. These vulnerabilities underscore the need for harmonized SBOM practices and targeted hardening in next-generation in-vehicle systems.
Hanbo Yu, Faiyaz Khan, Steven H. H. Ding, Natalia Stakhanova, Benjamin C. M. Fung
Comput. Secur.3
2025 Understanding Abandonment and Slowdown Dynamics in the Maven Ecosystem
abstract
The sustainability of libraries is critical for modern software development, yet many libraries face abandonment, posing significant risks to dependent projects. This study explores the prevalence and patterns of library abandonment in the Maven ecosystem. We investigate abandonment trends over the past decade, revealing that approximately one in four libraries fail to survive beyond their creation year. We also analyze the release activities of libraries, focusing on their lifespan and release speed, and analyze the evolution of these metrics within the lifespan of libraries. We find that while slow release speed and relatively long periods of inactivity are often precursors to abandonment, some abandoned libraries exhibit bursts of high frequent release activity late in their life cycle. Our findings contribute to a new understanding of library abandonment dynamics and offer insights for practitioners to identify and mitigate risks in software ecosystems.
Kazi Amit Hasan, Jerin Yasmin, Huizi Hao, Yuan Tian 0008, Safwat Hassan, Steven H. H. Ding
MSR6
2025 ProvSpider: A Robust and Universal Toolkit for Binary Provenance Analysis Using Deep Learning
abstract
Binary provenance analysis recovers essential information, such as architecture, structure, and toolchain, from executables lacking reliable metadata. This is crucial for reverse engineering. However, provenance recovery from binaries is highly challenging, due to three key factors: (1) binaries span diverse CPU architectures; (2) Raw byte sequences are often extremely long without clear boundaries; and (3) Compilation alters control flow, register usage, and memory layout, obscuring the original code structure and complicating analysis. To address these challenges, we propose a novel and robust analysis toolset, namely ProvSpider, to identify segment boundaries, types of segments as well as target CPU architectures, bitness, and endianness based on code-only sections. ProvSpider is built based on a convolutional neural network (CNN) to learn local execution patterns. We embed byte sequences into eight-dimensional vectors to capture bytes’ global dependencies. The gating mechanism after convolutional layers filters out noise and keeps most representative features. At last, the sliding window divides lengthy byte sequences into fixedlength processable chunks. Our model achieves high accuracy in all five analysis tasks, significantly outperforming the state-of-the-art models. By providing a universal and robust approach, ProvSpider lays the foundation for advancing binary provenance analysis, facilitating future improvements in binary analysis and reverse engineering.
Zhiwei Fu, Hanbo Yu, Steven H. H. Ding, Furkan Alaca, Philippe Charland
NCA4
2025 Transforming Generic Coder LLMs to Effective Binary Code Embedding Models for Similarity Detection
abstract
Cybersecurity and software research have crossed paths with modern deep learning research for a few years. The power of large language models (LLMs) in particular has intrigued us to apply them to understanding binary code. In this paper, we investigate some of the many ways LLMs can be applied to binary code similarity detection, as it is a significantly more difficult task compared to source code similarity detection due to the sparsity of information and less meaningful syntax. It also has great practical implications, such as vulnerability and malware detection. We find that pretrained LLMs are mostly capable of detecting similar binary code, even with a zero-shot setting. Our main contributions and findings are to provide several supervised fine-tuning methods that, when combined, significantly surpass zero-shot LLMs and state-of-the-art binary code similarity detection methods. Specifically, we up-train the model through data augmentation, translation-style causal learning, LLM2Vec, and cumulative GTE loss. With a complete ablation study, we show that our training method can transform a generic language model into a powerful binary similarity expert, and is also robust and general enough for cross-optimization, cross-architecture, and cross-obfuscation detection.
Litao Li, Leo Song, Steven H. H. Ding, Benjamin C. M. Fung, Philippe Charland
NeurIPS3
2025 MalGPT: A Generative Explainable Model for Malware Binaries
Mohd Saqib, Benjamin C. M. Fung, Steven H. H. Ding, Philippe Charland
ECML/PKDD (4)3
2025 NeuroYara: Learning to Rank for Yara Rules Generation Through Deep Language Modeling and Discriminative N-Gram Encoding
abstract
Signature-based malware detection methods are recognized for their simplicity, explainability, and efficiency. One of the most commonly used tools is Yara, which provides the syntax for crafting malware signatures. However, while developing high-quality Yara rules requires significant expertise in malware analysis, training such skilled analysts can be both resource-intensive and time-consuming. While a few works have been conducted to automate the generation of signatures, signatures generated by those works typically underperform the manually generated ones. In addition, these automated methods often depend on large static databases of hard-coded byte n-grams to minimize false positives. Instead of storing a large non-inclusive database to score byte n-grams, we propose a novel architecture utilizing two learning to rank neural networks to understand the underlying effectiveness and correlations among n-grams extracted for rule construction. This approach provides better flexibility and coverage of possible n-grams while reducing the required storage size from several GBs to only 10MBs. Combining these two models with a hierarchical density-based clustering method allows us to group multiple n-grams into logical conditions as Yara rules of higher quality. Experimental results show that our framework, NeuroYara, reduces the resources invested by analysts while generating rules with a low false-positive rate outperforming existing tools and manually-generated rules.
Ziad Mansour, Weihan Ou, Steven H. H. Ding, Mohammad Zulkernine, Philippe Charland
IEEE Trans. Dependable Secur. Comput.3
2025 PulseAnomaly: Unsupervised Anomaly Detection on Avionic Platforms With Seasonality and Trend Modeling in Transformer Networks
abstract
For communication within military avionic platforms (e.g., F-15 and F-35), the US Department of Defense established MIL-STD-1553 military standard. It has been released for more than 50 years and is still used in platforms other than military avionics. It was originally produced to be used with military avionics, but in the following decades, it was adopted into all branches of the armed forces, as well as spacecraft and commercial avionics. However, potential attacks against the MIL-STD-1553 may exist due to the demand for internet communication between planes and the lack of security. The current study presentsPulseAnomaly, a novel unsupervised anomaly detection model for the MIL-STD-1553 bus that utilizes time-feature and message sequences. Our model demonstrates better performance compared to baseline models in the test, achieving a higher F1-score and showing excellent AUROC compared to existing methods. Additionally, we have used data from a recently developed open-source MIL-STD-1553 real-time bus simulator, which features a more diverse range of attacks and data points that more closely resemble real-world scenarios. Evaluation results show that our model outperforms existing unsupervised solutions.
Hanbo Yu, Sudipta Acharya, Steven H. H. Ding, Mohammad Zulkernine
IEEE Trans. Dependable Secur. Comput.3
2025 Toward a Robust Detection of PowerShell Malware against Code Mixing and Obfuscation by Using Sentence Transformer and Similarity Learning
abstract
Embedded PowerShell commands or scripts are among the most popular malware payloads. For malware that prioritizes stealthiness, such as fileless malware, PowerShell’s access to Windows API functions without additional libraries makes it useful for evading detection. Detecting malicious PowerShell scripts and commands is an open challenge for proactive endpoint protection due to three major issues: (1) The malicious commands are usually hidden in a long script beyond the processing limit of typical machine learning models. (2) They are usually mixed with bulky benign scripts. (3) Script obfuscation can easily conceal their potential matching signatures. In this article, we introduce a novel model addressing these challenges. It incorporates similarity learning, sentence transformer, sliding window method, and stochastic gradient descent (SGD) classifier. Our key insight is that malicious PowerShell code, particularly when obfuscated, exhibits semantic and statistical deviations from benign administrative usage, and these deviations can be captured by contrastive sentence embeddings without the need for de-obfuscation or handcrafted features. We operate this insight through a Siamese similarity learning framework that improves robustness against Out-of-Vocabulary tokens due to unseen code obfuscation methods. The sliding window method enables the model to handle long scripts, and the SGD classifier evaluates segment-level maliciousness. Our model achieves accuracies of 99.01%, 97.59%, 98.70%, and 99.73% across multiple obfuscated and mixed script benchmarks, outperforming existing baselines by over 30% in all cases. This work demonstrates a scalable and effective strategy for robust PowerShell malware detection in real-world scenarios.
Zhiwei Fu, Leo Song, Steven H. H. Ding, Furkan Alaca, Sudipta Acharya
ACM Trans. Priv. Secur.3
2025 Obfuscated Clone Search in JavaScript based on Reinforcement Subsequence Learning
abstract
Finding similar code is important for software engineering, defense of intellectual property, and security, and one of the increasingly common ways adversaries use to defeat the detection of similar code is through obfuscations such as code transformation and scattering the code they wish to hide among long sequences. Moving code far enough apart poses a specific challenge for solutions with localized features (e.g., n-grams), or attention mechanisms as the code parts are distributed beyond the local context window. We introduce a neural network solution pattern called “Cybertron” that addresses this problem by utilizing reinforcement learning to train a code abstraction and summarization function; this converts arbitrarily long code into fixed-length real vectors in a way that is optimized for similarity search. The key to the design is the smart selection of important elements of the code and abstraction to preserve semantic function while minimizing syntactic feature information. We evaluated the approach on a three-challenge benchmark of obfuscated JavaScript, a scripting language that is commonly obfuscated and for which code-mixing is a rising challenge. The evaluation shows our approach identifies obfuscated code within even large scripts with an AUC of 78%, which outperforms current state-of-the-art sequence models by 7–35%.
Leo Song, Steven H. H. Ding, Yuan Tian 0008, Li Tao Li, Weihan Ou, Philippe Charland, Andrew Walenstein
ACM Trans. Softw. Eng. Methodol.2
2025 Mecha: A Neural-Symbolic Open-Set Homogeneous Decision Fusion Approach for Zero-Day Malware Similarity Detection
abstract
With increasing numbers of novel malware each year, tools are required for efficient and accurate variant matching under the same family, for the purpose of effective proactive threat detection, retro-hunting, and attack campaign tracking. All of the state-of-the-art Deep Learning (DL) approaches assume that the incoming samples originate from known families and incorrectly identify novel families. Additionally, most of the existing solutions that leverage the Siamese Neural Network architecture either rely on pair-wise comparisons or computationally expensive preprocessing steps that are not scalable to a real-world malware triage volume requirement. We propose a different route, Mecha, a Neural-Symbolic Machine Learning (ML) system for malware variant matching and zero-day family detection. Mecha is comprised of an embedding network trained in two different scenarios for byte string embedding and an open-set approximate nearest neighbour algorithm for variant matching and zero-day detection. Our embedding network uses triplet loss for embedding generation and reinforcement-based Expectation Maximization (EM) learning for full deployment optimization. We conduct multiple in-sample and out-of-sample experiments to demonstrate the model's generalizability toward novel variants and families. We also show that Mecha can detect samples outside the known set of malware samples with an accuracy greater than 0.990.
Christopher Molloy, Jeremy Banks, Steven H. H. Ding, Furkan Alaca, Philippe Charland, Andrew Walenstein
IEEE Trans. Software Eng.3
2024 Dynamic Neural Control Flow Execution: an Agent-Based Deep Equilibrium Approach for Binary Vulnerability Detection
abstract
Software vulnerabilities are a challenge in cybersecurity. Manual security patches are often difficult and slow to be deployed, while new vulnerabilities are created. Binary code vulnerability detection is less studied and more complex compared to source code, and this has important practical implications. Deep learning has become an efficient and powerful tool in the security domain, where it provides end-to-end and accurate prediction. Modern deep learning approaches learn the program semantics through sequence and graph neural networks, using various intermediate representation of programs, such as abstract syntax trees (AST) or control flow graphs (CFG). Due to the complex nature of program execution, the output of an execution depends on the many program states and inputs. Also, a CFG generated from static analysis can be an overestimation of the true program flow. Moreover, the size of programs often does not allow a graph neural network with fixed layers to aggregate global information. To address these issues, we propose DeepEXE, an agent-based implicit neural network that mimics the execution path of a program. We use reinforcement learning to enhance the branching decision at every program state transition and create a dynamic environment to learn the dependency between a vulnerability and certain program states. An implicitly defined neural network enables nearly infinite state transitions until convergence, which captures the structural information at a higher level. The experiments are conducted on two semi-synthetic and two real-world datasets. We show that DeepEXE is an accurate and efficient method and outperforms the state-of-the-art vulnerability detection methods.
Li Tao Li, Steven H. H. Ding, Andrew Walenstein, Philippe Charland, Benjamin C. M. Fung
CIKM2
2024 Ch4os: Discretized Generative Adversarial Network for Functionality-Preserving Evasive Modification on Malware
Christopher Molloy, Furkan Alaca, Steven H. H. Ding
ICANN (9)3
2024 An empirical study on developers' shared conversations with ChatGPT in GitHub pull requests and issues
Huizi Hao, Kazi Amit Hasan, Hong Qin 0014, Marcos Macedo, Yuan Tian 0008, Steven H. H. Ding, Ahmed E. Hassan
Empir. Softw. Eng.6
2024 VeriBin: A Malware Authorship Verification Approach for APT Tracking through Explainable and Functionality-Debiasing Adversarial Representation Learning
abstract
Malware attacks are posing a significant threat to national security, cooperate network, and public endpoint security. Identifying the Advanced Persistent Threat (APT) groups behind the attacks and grouping their activities into attack campaigns help security investigators trace their activities thus providing better security protections against future attacks. Existing Cyber Threat Intelligent (CTI) components mainly focus on malware family identification and behavior characterization, which cannot solve the APT tracking problem: while APT tracking needs one to link malware binaries of multiple families to a single threat actor, these behavior or function-based techniques are tightened up to a specific attack technique and would fail on connecting different families. Binary Authorship Attribution (AA) solutions could discriminate against threat actors based on their stylometric traits. However, AA solutions assume that the author of a binary is within a fixed candidate author set. However, real-world malware binaries may be created by a new unknown threat actor. To address this research gap, we propose VeriBin for the Binary Authorship Verification (BAV) problem. VeriBin is a novel adversarial neural network that extracts functionality-agnostic style representations from assembly code for the AV task. The extracted style representations can be visualized and are explainable with VeriBin’s multi-head attention mechanism. We benchmark VeriBin with state-of-the-art coding style representations on a standard dataset and a recent malware-APT dataset. Given two anonymous binaries of out-of-sample authors, VeriBin can accurately determine whether they belong to the same author or not. VeriBin is resilient to compiler optimizations and robust against malware family variants.
Weihan Ou, Steven H. H. Ding, Mohammad Zulkernine, Li Tao Li, Sarah Labrosse
ACM Trans. Priv. Secur.2
2023 Variational Autoencoder with Temporal Condition for Effective Shape-based Calcium Imaging Neuron Registration
abstract
Thanks to recent advances in optical imaging techniques, calcium imaging can now record the activities of thousands of neurons simultaneously, through several sessions and over long periods. Neuron registration assumes a vital role in this process, enabling the monitoring of neurons across multiple movies and extended timeframes by aligning their spatial patterns. Previous approaches often relied on clusters or probabilistic models based on simple distance metrics like overlapping pixels or center-to-center distances, neglecting crucial neuron shape and spatial relationship details. In this paper, we introduce a novel technique for cell registration. Our investigation revealed that a neuron’s shape is influenced by its temporal behaviors, leading us to suggest a temporal-conditional variational autoencoder (tcVAE) for precise shape modeling. A comprehensive evaluation demonstrates that incorporating shape-related details can significantly enhance the quality of neuron registration.
Cyrus Y. H. Fung, Sudipta Acharya, Tak Pan Wong, Steven H. H. Ding
BIBM4
2023 MaGnn: Binary-Source Code Matching by Modality-Sharing Graph Convolution for Binary Provenance Analysis
abstract
The number and variety of binaries running on electrical devices, public clouds, and on-premise infrastructure have been increasing rapidly. Recent successful supply chain attacks indicate that even for binaries known to be developed by trustful developers, they can still contain malicious functionalities and copy-and-pasted vulnerabilities that pose security risks to operational systems and end users. By analyzing the origin of a target code, code provenance analysis helps to relieve such problem by revealing information about the origin of a binary sample such as the author or the included software bill-of-materials. Since in most cases source symbol information is removed during the compilation process, given a binary code sample, matching it to its corresponding source code could improve the accuracy and efficiency of the provenance analysis. Existing binary-source code matching methods focus on comparing manually selected code literals (e.g. the number of if/else statements). However, these methods suffer from the issue of generalizability and require significant manual efforts.Different from the previous methods, we propose a machine learning-based binary-source code matching system, MaGnn, which measures the consistency of an input binary-source code pair by automatically extracting high-dimensional feature representations of the input and calculating the functionality similarity. With the Siamese architecture that shares a unified encoder across two modalities, McGnn is able to calculate the similarity of the input binary-source code pair with the automatically-extracted functionality representations. With the graph convolution neural network as the representation encoder, MaGnn is able to learn and encode the functionality information of the input pairs from their graph features into high-dimensional representation vectors. We benchmark MaGnn with a state-of-the-art binary-source code matching method and two machine-learning models on six out-of-sample datasets collected from five real-world libraries. Our experiment results show that MaGnn outperforms the baselines on most out-of-sample datasets.
Weihan Ou, Steven H. H. Ding
COMPSAC2
2023 Milo: Attacking Deep Pre-trained Model for Programming Languages Tasks with Anti-analysis Code Obfuscation
Leo Song, Steven H. H. Ding
COMPSAC2
2023 GenTAL: Generative Denoising Skip-gram Transformer for Unsupervised Binary Code Similarity Detection
abstract
Binary code similarity detection serves a critical role in cybersecurity. It alleviates the huge manual effort required in the reverse engineering process for malware analysis and vulnerability detection, where the original source code is often not available. Most of the existing solutions focus on a manual feature engineering process and customized code matching algorithms that are inefficient and inaccurate. Recent deep learning-based solutions embed the semantics of binary code into a latent space through supervised contrastive learning. However, one cannot cover all the possible forms in the training set to learn the variance of the same semantics. In this paper, we propose an unsupervised model aiming to learn the intrinsic representation of assembly code semantics. Specifically, we propose a Transformer-based auto-encoder like language model for the low-level assembly code grammar to capture the abstract semantic representation. By coupling a Transformer encoder and a skip-gram style loss design, it can learn a compact representation that is robust against different compilation options. We conduct experiments on four different block-level code similarity tasks. It shows that our method is more robust compared to the state-of-the-art solutions.
Li Tao Li, Steven H. H. Ding, Philippe Charland
IJCNN2
2023 Understanding the Time to First Response in GitHub Pull Requests
abstract
The pull-based development is widely adopted in modern open-source software (OSS) projects, where developers propose changes to the codebase by submitting a pull request (PR). However, due to many reasons, PRs in OSS projects frequently experience delays across their lifespan, including prolonged waiting times for the first response. Such delays may significantly impact the efficiency and productivity of the development process, as well as the retention of new contributors as long-term contributors.In this paper, we conduct an exploratory study on the time-to-first-response for PRs by analyzing 111,094 closed PRs from ten popular OSS projects on GitHub. We find that bots frequently generate the first response in a PR, and significant differences exist in the timing of bot-generated versus human-generated first responses. We then perform an empirical study to examine the characteristics of bot- and human-generated first responses, including their relationship with the PR’s lifetime. Our results suggest that the presence of bots is an important factor contributing to the time-to-first-response in the pull-based development paradigm, and hence should be separately analyzed from human responses. We also report the characteristics of PRs that are more likely to experience long waiting for the first human-generated response. Our findings have practical implications for newcomers to understand the factors contributing to delays in their PRs.
Kazi Amit Hasan, Marcos Macedo, Yuan Tian 0008, Bram Adams, Steven H. H. Ding
MSR5
2023 TDRLM: Stylometric learning for authorship verification by Topic-Debiasing
Weihan Ou, Sudipta Acharya, Steven H. H. Ding, Ryan D'Gama, Hanbo Yu
Expert Syst. Appl.4
2023 AIM: An Android Interpretable Malware detector based on application class modeling
Farnood Faghihi, Mohammad Zulkernine, Steven H. H. Ding
J. Inf. Secur. Appl.3
2023 VulANalyzeR: Explainable Binary Vulnerability Detection with Multi-task Learning and Attentional Graph Convolution
abstract
Software vulnerabilities have been posing tremendous reliability threats to the general public as well as critical infrastructures, and there have been many studies aiming to detect and mitigate software defects at the binary level. Most of the standard practices leverage both static and dynamic analysis, which have several drawbacks like heavy manual workload and high complexity. Existing deep learning-based solutions not only suffer to capture the complex relationships among different variables from raw binary code but also lack the explainability required for humans to verify, evaluate, and patch the detected bugs. We propose VulANalyzeR, a deep learning-based model, for automated binary vulnerability detection, Common Weakness Enumeration-type classification, and root cause analysis to enhance safety and security. VulANalyzeR features sequential and topological learning through recurrent units and graph convolution to simulate how a program is executed. The attention mechanism is integrated throughout the model, which shows how different instructions and the corresponding states contribute to the final classification. It also classifies the specific vulnerability type through multi-task learning as this not only provides further explanation but also allows faster patching for zero-day vulnerabilities. We show that VulANalyzeR achieves better performance for vulnerability detection over the state-of-the-art baselines. Additionally, a Common Vulnerability Exposure dataset is used to evaluate real complex vulnerabilities. We conduct case studies to show that VulANalyzeR is able to accurately identify the instructions and basic blocks that cause the vulnerability even without given any prior knowledge related to the locations during the training phase.
Litao Li, Steven H. H. Ding, Yuan Tian 0008, Benjamin C. M. Fung, Philippe Charland, Weihan Ou, Leo Song, Congwei Chen
ACM Trans. Priv. Secur.2
2023 SCS-Gan: Learning Functionality-Agnostic Stylometric Representations for Source Code Authorship Verification
abstract
In recent years, the number of anonymous script-based fileless malware attacks, software copyright disputes, and code plagiarism issues has increased rapidly. In the literature, automated Code Authorship Analysis (CAA) techniques have been proposed to reduce the manual effort in identifying those attacks and issues. Most CAA techniques aim to solve the task of Authorship Attribution (AA), i.e., identifying the actual author of a source code fragment from a given set of candidate authors. However, in many real-world scenarios, investigators do not have a predefined set of authors containing the actual author at the time of investigation, i.e., contradicting AA's assumption. Additionally, existing AA techniques ignore the influence of code functionality when identifying the authorship, which leads to biased matching simply based on code functionality. Different from AA, the task of (extreme) Authorship Verification (AV) is to decide if two texts were written by the same person or not. AV techniques do not need a predefined author set and thus could be applied in more code authorship-related applications than AA. To our knowledge, there is no previous work attempting to solve the AV problem for the source code. To fill the gap, we propose a novel adversarial neural network, namely SCS-Gan, that can learn a stylometric representation of code for automated AV. With the multi-head attention mechanism, SCS-Gan focuses on the code parts that are most informative regarding personal styles and generates functionality-agnostic stylometric representations through adversarial training. We benchmark SCS-Gan and two state-of-the-art code representation models on four out-of-sample datasets collected from a real-world programming competition. Our experiment results show that SCS-Gan outperforms the baselines on all four out-of-sample datasets.
Weihan Ou, Steven H. H. Ding, Yuan Tian 0008, Leo Song
IEEE Trans. Software Eng.2
2022 Adversarial Variational Modality Reconstruction and Regularization for Zero-Day Malware Variants Similarity Detection
abstract
Matching malware variants in the same malware family has always been a significant challenge for Cyber Threat Intelligence (CTI). For zero-day malware that does not belong to an existing family, a timely matching of its variants is essential for effective threat tracing and prompt response to the cyber incident. However, malware variants are of diverse forms that make them difficult to match. Additionally, the information extracted from a given malware sample is inaccurate, especially on zero-day malware. Existing malware solutions only focus on detecting known malware or find if two samples are similar without creating any reusable representation of the samples. In this paper, we propose the first practical and efficient solution for zero-day malware variant matching with reconstruction. By combining multi-modality learning and a Siamese-based structure, our model can navigate across different modalities and match zero-day variants. To address the missing or noisy modality issue, we propose a Conditional Variable Autoencoder with a Generative Adversarial Network for heightened resolution. We trained the model on 100,000 malware triplet pairs. Our experiments on real-world noisy samples show that the model out-performs the state-of-the-art and can accurately match not only zero-day malware, but also out-of-sample benign binaries of the same category.
Christopher Molloy, Jeremy Banks, Steven H. H. Ding, Philippe Charland, Andrew Walenstein, Litao Li
ICDM3
2022 TASR: Adversarial learning of topic-agnostic stylometric representations for informed crisis response through social media
abstract
The impact of crisis events can be devastating in a multitude of ways, many of which are unpredictable due to the suddenness in which they occur. The evolution of social media (for example Twitter) has given directly affected individuals or those with valuable information a platform to effectively share their stories to the masses. As a result, these platforms have become vast repositories of helpful information for emergency organizations. However, different crisis events often contain event-specific keywords, which results in the difficult extraction of useful information with a single model. In this paper, we put forward TASR, which stands for Topic-Agnostic Stylometric Representations, a novice deep learning architecture that uses stylometric and adversarial learning to remove topical bias to better manage the unknown surrounding unseen events. As an alternative to domain adaptive approaches requiring data from the unseen event, it reduces the work for those responding to the onset of a crisis. Overall, we conduct a comprehensive study of the situational properties of TASR, the benefits of its architecture including its topic-agnostic and explainable properties, and how it improves upon comparable models in past research. From two experiments, on average, TASR is able to outperform state-of-the-art methods such as transfer learning and domain adoption by 11% in AUC. The ablation study illustrates how different architecture choices of TASR impact the results and that TASR has been optimized for this task. Finally, we conduct a case study to show that explainable results from our model can be used to help guide human analysts through crisis information extraction.
Litao Li, Rylen Sampson, Steven H. H. Ding, Leo Song
Inf. Process. Manag.3
2022 CamoDroid: An Android application analysis environment resilient against sandbox evasion
Farnood Faghihi, Mohammad Zulkernine, Steven H. H. Ding
J. Syst. Archit.3
2022 OD1NF1ST: True Skip Intrusion Detection and Avionics Network Cyber-attack Simulation
abstract
MIL-STD-1553 is a communication bus that has been used by many military avionics platforms, such as the F-15 and F-35 fighter jets, for almost 50 years. Recently, it has become clear that the lack of security on MIL-STD-1553 and the requirement for internet communication between planes has revealed numerous potential attack vectors for malicious parties. Prevention of these attacks by modernizing the MIL-STD-1553 is not practical due to the military applications and existing far-reaching installations of the bus. We present a software system that can simulate bus transmissions to create easy, replicable, and large datasets of MIL-STD-1553 communications. We also propose an intrusion detection system (IDS) that can identify anomalies and the precise type of attack using recurrent neural networks with a reinforcement learning true-skip data selection algorithm. Our IDS outperforms existing algorithms designed for MIL-STD-1553 in binary anomaly detection tasks while also performing attack classification and minimizing computational resource cost. Our simulator can generate more data with higher fidelity than existing methods and integrate attack scenarios with greater detail. Furthermore, the simulator and IDS can be combined to form a web-based attack-defense game.
Michael Wrana, Marwa Elsayed, Karim Lounis, Ziad Mansour, Steven H. H. Ding, Mohammad Zulkernine
ACM Trans. Cyber Phys. Syst.5
2022 AdaptIDS: Adaptive Intrusion Detection for Mission-Critical Aerospace Vehicles
abstract
Aerospace and defense industries are particularly vulnerable to cyber threats given their sensitive nature, significantly extending the consequences of security breaches to the national level. Aerospace vehicles are augmented by cooperative control, intelligent, connected, and autonomous systems. The risk against such systems is further amplified due to commonly relying on the MIL-STD-1553 communication bus developed with a high focus on reliability and fault tolerance, albeit with security as a second priority. MIL-STD-1553 (a.k.a., STANAG 3838 by NATO) is a standard that describes a serial data communication bus primarily used in aerospace vehicles for military and civilian applications, including avionics, aircraft, and spacecraft data handling. In the absence of core security measures such as authentication, authorization, and encryption, the bus connecting sensitive functions, including autopilot, GPS, fuel valve switches, and other avionics equipment, is easily vulnerable to a wide range of attacks. This paper proposes, AdaptIDS, a novel adaptive intrusion detection system as a security analytics framework for the MIL-STD-1553 communication bus. AdaptIDS mainly adopts data science principles and leverages advanced deep learning techniques (i.e., the stacking ensemble) to boost its generalization capabilities for detecting unseen patterns of attacks in the dynamic changing environment of aerospace vehicles. Extensive experiments are conducted using two datasets generated from an open-source simulation system, reflecting dynamic real-life scenarios. The evaluation results demonstrate that our solution outperforms existing solutions with high detection effectiveness of 0.99 F1-measure and computational time efficiency.
Marwa Elsayed, Michael Wrana, Ziad Mansour, Karim Lounis, Steven H. H. Ding, Mohammad Zulkernine
IEEE Trans. Intell. Transp. Syst.5
2021 A Novel and Dedicated Machine Learning Model for Malware Classification
Miles Q. Li, Benjamin C. M. Fung, Philippe Charland, Steven H. H. Ding
ICSOFT4
2021 ER-AE: Differentially Private Text Generation for Authorship Anonymization
abstract
Haohan Bo, Steven H. H. Ding, Benjamin C. M. Fung, Farkhund Iqbal. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Haohan Bo, Steven H. H. Ding, Benjamin C. M. Fung, Farkhund Iqbal
NAACL-HLT2
2021 I-MAD: Interpretable malware detector using Galaxy Transformer
Miles Q. Li, Benjamin C. M. Fung, Philippe Charland, Steven H. H. Ding
Comput. Secur.4
2020 Detecting breaking news rumors of emerging topics in social media
Sarah A. Alkhodair, Steven H. H. Ding, Benjamin C. M. Fung, Junqiang Liu
Inf. Process. Manag.2
2019 Asm2Vec: Boosting Static Representation Robustness for Binary Clone Search against Code Obfuscation and Compiler Optimization
abstract
Reverse engineering is a manually intensive but necessary technique for understanding the inner workings of new malware, finding vulnerabilities in existing systems, and detecting patent infringements in released software. An assembly clone search engine facilitates the work of reverse engineers by identifying those duplicated or known parts. However, it is challenging to design a robust clone search engine, since there exist various compiler optimization options and code obfuscation techniques that make logically similar assembly functions appear to be very different. A practical clone search engine relies on a robust vector representation of assembly code. However, the existing clone search approaches, which rely on a manual feature engineering process to form a feature vector for an assembly function, fail to consider the relationships between features and identify those unique patterns that can statistically distinguish assembly functions. To address this problem, we propose to jointly learn the lexical semantic relationships and the vector representation of assembly functions based on assembly code. We have developed an assembly code representation learning model \emph{Asm2Vec}. It only needs assembly code as input and does not require any prior knowledge such as the correct mapping between assembly functions. It can find and incorporate rich semantic relationships among tokens appearing in assembly code. We conduct extensive experiments and benchmark the learning model with state-of-the-art static and dynamic clone search approaches. We show that the learned representation is more robust and significantly outperforms existing methods against changes introduced by obfuscation and optimizations.
Steven H. H. Ding, Benjamin C. M. Fung, Philippe Charland
IEEE Symposium on Security and Privacy1
2019 Arabic Authorship Attribution: An Extensive Study on Twitter Posts
abstract
Law enforcement faces problems in tracing the true identity of offenders in cybercrime investigations. Most offenders mask their true identity, impersonate people of high authority, or use identity deception and obfuscation tactics to avoid detection and traceability. To address the problem of anonymity, authorship analysis is used to identify individuals by their writing styles without knowing their actual identities. Most authorship studies are dedicated to English due to its widespread use over the Internet, but recent cyber-attacks such as the distribution of Stuxnet indicate that Internet crimes are not limited to a certain community, language, culture, ideology, or ethnicity. To effectively investigate cybercrime and to address the problem of anonymity in online communication, there is a pressing need to study authorship analysis of languages such as Arabic, Chinese, Turkish, and so on. Arabic, the focus of this study, is the fourth most widely used language on the Internet. This study investigates authorship of Arabic discourse/text, especially tiny text, Twitter posts. We benchmark the performance of a profile-based approach that uses n -grams as features and compare it with state-of-the-art instance-based classification techniques. Then we adapt an event-visualization tool that is developed for English to accommodate both Arabic and English languages and visualize the result of the attribution evidence. In addition, we investigate the relative effect of the training set, the length of tweets, and the number of authors on authorship classification accuracy. Finally, we show that diacritics have an insignificant effect on the attribution process and part-of-speech tags are less effective than character-level and word-level n -grams.
Malik H. Altakrori, Farkhund Iqbal, Benjamin C. M. Fung, Steven H. H. Ding, Abdallah Tubaishat
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2019 Learning Stylometric Representations for Authorship Analysis
abstract
Authorship analysis (AA) is the study of unveiling the hidden properties of authors from textual data. It extracts an author's identity and sociolinguistic characteristics based on the reflected writing styles in the text. The process is essential for various areas, such as cybercrime investigation, psycholinguistics, political socialization, etc. However, most of the previous techniques critically depend on the manual feature engineering process. Consequently, the choice of feature set has been shown to be scenario- or dataset-dependent. In this paper, to mimic the human sentence composition process using a neural network approach, we propose to incorporate different categories of linguistic features into distributed representation of words in order to learn simultaneously the writing style representations based on unlabeled texts for AA. In particular, the proposed models allow topical, lexical, syntactical, and character-level feature vectors of each document to be extracted as stylometrics. We evaluate the performance of our approach on the problems of authorship characterization, authorship identification and authorship verification with the Twitter, blog, review, novel, and essay datasets. The experiments suggest that our proposed text representation outperforms the static stylometrics, dynamic n -grams, latent Dirichlet allocation, latent semantic analysis, distributed memory model of paragraph vectors, distributed bag of words version of paragraph vector, word2vec representations, and other baselines.
Steven H. H. Ding, Benjamin C. M. Fung, Farkhund Iqbal, William Kwok-Wai Cheung
IEEE Trans. Cybern.1
2016 Kam1n0: MapReduce-based Assembly Clone Search for Reverse Engineering
abstract
Assembly code analysis is one of the critical processes for detecting and proving software plagiarism and software patent infringements when the source code is unavailable. It is also a common practice to discover exploits and vulnerabilities in existing software. However, it is a manually intensive and time-consuming process even for experienced reverse engineers. An effective and efficient assembly code clone search engine can greatly reduce the effort of this process, since it can identify the cloned parts that have been previously analyzed. The assembly code clone search problem belongs to the field of software engineering. However, it strongly depends on practical nearest neighbor search techniques in data mining and databases. By closely collaborating with reverse engineers and Defence Research and Development Canada (DRDC), we study the concerns and challenges that make existing assembly code clone approaches not practically applicable from the perspective of data mining. We propose a new variant of LSH scheme and incorporate it with graph matching to address these challenges. We implement an integrated assembly clone search engine called Kam1n0. It is the first clone search engine that can efficiently identify the given query assembly function's subgraph clones from a large assembly code repository. Kam1n0 is built upon the Apache Spark computation framework and Cassandra-like key-value distributed storage. A deployed demo system is publicly available. Extensive experimental results suggest that Kam1n0 is accurate, efficient, and scalable for handling large volume of assembly code.
Steven H. H. Ding, Benjamin C. M. Fung, Philippe Charland
KDD1
2015 A Visualizable Evidence-Driven Approach for Authorship Attribution
abstract
The Internet provides an ideal anonymous channel for concealing computer-mediated malicious activities, as the network-based origins of critical electronic textual evidence (e.g., emails, blogs, forum posts, chat logs, etc.) can be easily repudiated. Authorship attribution is the study of identifying the actual author of the given anonymous documents based on the text itself, and for decades, many linguistic stylometry and computational techniques have been extensively studied for this purpose. However, most of the previous research emphasizes promoting the authorship attribution accuracy, and few works have been done for the purpose of constructing and visualizing the evidential traits. In addition, these sophisticated techniques are difficult for cyber investigators or linguistic experts to interpret. In this article, based on the End-to-End Digital Investigation (EEDI) framework, we propose a visualizable evidence-driven approach, namely VEA, which aims at facilitating the work of cyber investigation. Our comprehensive controlled experiment and the stratified experiment on the real-life Enron email dataset demonstrate that our approach can achieve even higher accuracy than traditional methods; meanwhile, its output can be easily visualized and interpreted as evidential traits. In addition to identifying the most plausible author of a given text, our approach also estimates the confidence for the predicted result based on a given identification context and presents visualizable linguistic evidence for each candidate.
Steven H. H. Ding, Benjamin C. M. Fung, Mourad Debbabi
ACM Trans. Inf. Syst. Secur.1