VLDB 2026 Research / reviewers in the wild / expert
Andrew Walenstein
dblp:79/5970
· DBLP profile ↗
15ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0003-1103-2465ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Obfuscated Clone Search in JavaScript based on Reinforcement Subsequence LearningabstractFinding similar code is important for software engineering, defense of intellectual property, and security, and one of the increasingly common ways adversaries use to defeat the detection of similar code is through obfuscations such as code transformation and scattering the code they wish to hide among long sequences. Moving code far enough apart poses a specific challenge for solutions with localized features (e.g., n-grams), or attention mechanisms as the code parts are distributed beyond the local context window. We introduce a neural network solution pattern called “Cybertron” that addresses this problem by utilizing reinforcement learning to train a code abstraction and summarization function; this converts arbitrarily long code into fixed-length real vectors in a way that is optimized for similarity search. The key to the design is the smart selection of important elements of the code and abstraction to preserve semantic function while minimizing syntactic feature information. We evaluated the approach on a three-challenge benchmark of obfuscated JavaScript, a scripting language that is commonly obfuscated and for which code-mixing is a rising challenge. The evaluation shows our approach identifies obfuscated code within even large scripts with an AUC of 78%, which outperforms current state-of-the-art sequence models by 7–35%. Leo Song, Steven H. H. Ding, Yuan Tian 0008, Li Tao Li, Weihan Ou, Philippe Charland, Andrew Walenstein |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2025 | Mecha: A Neural-Symbolic Open-Set Homogeneous Decision Fusion Approach for Zero-Day Malware Similarity DetectionabstractWith increasing numbers of novel malware each year, tools are required for efficient and accurate variant matching under the same family, for the purpose of effective proactive threat detection, retro-hunting, and attack campaign tracking. All of the state-of-the-art Deep Learning (DL) approaches assume that the incoming samples originate from known families and incorrectly identify novel families. Additionally, most of the existing solutions that leverage the Siamese Neural Network architecture either rely on pair-wise comparisons or computationally expensive preprocessing steps that are not scalable to a real-world malware triage volume requirement. We propose a different route, Mecha, a Neural-Symbolic Machine Learning (ML) system for malware variant matching and zero-day family detection. Mecha is comprised of an embedding network trained in two different scenarios for byte string embedding and an open-set approximate nearest neighbour algorithm for variant matching and zero-day detection. Our embedding network uses triplet loss for embedding generation and reinforcement-based Expectation Maximization (EM) learning for full deployment optimization. We conduct multiple in-sample and out-of-sample experiments to demonstrate the model's generalizability toward novel variants and families. We also show that Mecha can detect samples outside the known set of malware samples with an accuracy greater than 0.990. Christopher Molloy, Jeremy Banks, Steven H. H. Ding, Furkan Alaca, Philippe Charland, Andrew Walenstein |
IEEE Trans. Software Eng. | 6 |
| 2024 | Dynamic Neural Control Flow Execution: an Agent-Based Deep Equilibrium Approach for Binary Vulnerability DetectionabstractSoftware vulnerabilities are a challenge in cybersecurity. Manual security patches are often difficult and slow to be deployed, while new vulnerabilities are created. Binary code vulnerability detection is less studied and more complex compared to source code, and this has important practical implications. Deep learning has become an efficient and powerful tool in the security domain, where it provides end-to-end and accurate prediction. Modern deep learning approaches learn the program semantics through sequence and graph neural networks, using various intermediate representation of programs, such as abstract syntax trees (AST) or control flow graphs (CFG). Due to the complex nature of program execution, the output of an execution depends on the many program states and inputs. Also, a CFG generated from static analysis can be an overestimation of the true program flow. Moreover, the size of programs often does not allow a graph neural network with fixed layers to aggregate global information. To address these issues, we propose DeepEXE, an agent-based implicit neural network that mimics the execution path of a program. We use reinforcement learning to enhance the branching decision at every program state transition and create a dynamic environment to learn the dependency between a vulnerability and certain program states. An implicitly defined neural network enables nearly infinite state transitions until convergence, which captures the structural information at a higher level. The experiments are conducted on two semi-synthetic and two real-world datasets. We show that DeepEXE is an accurate and efficient method and outperforms the state-of-the-art vulnerability detection methods. Li Tao Li, Steven H. H. Ding, Andrew Walenstein, Philippe Charland, Benjamin C. M. Fung |
CIKM | 3 |
| 2024 | GAGE: Genetic Algorithm-Based Graph Explainer for Malware AnalysisabstractMalware analysts often prefer reverse engineering using Call Graphs, Control Flow Graphs (CFGs), and Data Flow Graphs (DFGs), which involves the utilization of black-box Deep Learning (DL) models. The proposed research introduces a structured pipeline for reverse engineering-based analysis, offering promising results compared to state-of-the-art methods and providing high-level interpretability for malicious code blocks in subgraphs. We propose the Canonical Executable Graph (CEG) as a new representation of Portable Executable (PE) files, uniquely incorporating syntactical and semantic information into its node embeddings. At the same time, edge features capture structural aspects of PE files. This is the first work to present a PE file representation encompassing syntactical, semantic, and structural characteristics, whereas previous efforts typically focused solely on syntactic or structural properties. Furthermore, recognizing the limitations of existing graph explanation methods within Explainable Artificial Intelligence (XAI) for malware analysis, primarily due to the specificity of malicious files, we introduce Genetic Algorithm-based Graph Explainer (GAGE). GAGE operates on the CEG, striving to identify a precise subgraph relevant to predicted malware families. Through experiments and comparisons, our proposed pipeline exhibits substantial improvements in model robustness scores and discriminative power compared to the previous benchmarks. Furthermore, we have successfully used GAGE in practical applications on real-world data, producing meaningful insights and interpretability. This research offers a robust solution to enhance cybersecurity by delivering a transparent and accurate understanding of malware behaviour. Moreover, the proposed algorithm is specialized in handling graph-based data, effectively dissecting complex content and isolating influential nodes. Mohd Saqib, Benjamin C. M. Fung, Philippe Charland, Andrew Walenstein |
ICDE | 4 |
| 2024 | AsmDocGen: Generating Functional Natural Language Descriptions for Assembly Codeabstract"This study explores the field of software reverse engineering through the lens of code summarization, which involves generating informative and concise summaries of code functionality. A significant aspect of this research is the application of assembly code summarization in malware analysis, highlighting its critical role in understanding and mitigating potential security threats. Although there have been recent efforts to develop code summarization techniques for high-level programming languages, to the best of our knowledge, this study is the first attempt to generate comments for assembly code. For this purpose, we first built a carefully curated dataset of assembly function-comment pairs. We then focused on automatic assembly code summarization using transfer learning with pre-trained natural language processing (NLP) models, including BERT, DistilBERT, RoBERTa, and CodeBERT. The results of our experiments show a notable advantage of Code- BERT: despite its initial training on high-level programming languages alone, it excels in learning assembly language, outperforming other pre-trained NLP models."@eng Jesia Quader Yuki, Mohammadhossein Amouei, Benjamin C. M. Fung, Philippe Charland, Andrew Walenstein |
ICSOFT | 5 |
| 2024 | VulEXplaineR: XAI for Vulnerability Detection on Assembly Code
Samaneh Mahdavifar, Mohd Saqib, Benjamin C. M. Fung, Philippe Charland, Andrew Walenstein |
ECML/PKDD (9) | 5 |
| 2023 | An Efficient Resilient MPC Scheme via Constraint Tightening Against Cyberattacks: Application to Vehicle Cruise Control
Milad Farsi, Shuhao Bian, Nasser L. Azad, Xiaobing Shi, Andrew Walenstein |
ICINCO (1) | 5 |
| 2023 | Finding associations between natural and computer languages: A case-study of bilingual LDA applied to the bleeping computer forum posts
Kundi Yao, Gustavo Ansaldi Oliva, Ahmed E. Hassan, Muhammad Asaduzzaman, Andrew J. Malton, Andrew Walenstein |
J. Syst. Softw. | 6 |
| 2022 | Adversarial Variational Modality Reconstruction and Regularization for Zero-Day Malware Variants Similarity DetectionabstractMatching malware variants in the same malware family has always been a significant challenge for Cyber Threat Intelligence (CTI). For zero-day malware that does not belong to an existing family, a timely matching of its variants is essential for effective threat tracing and prompt response to the cyber incident. However, malware variants are of diverse forms that make them difficult to match. Additionally, the information extracted from a given malware sample is inaccurate, especially on zero-day malware. Existing malware solutions only focus on detecting known malware or find if two samples are similar without creating any reusable representation of the samples. In this paper, we propose the first practical and efficient solution for zero-day malware variant matching with reconstruction. By combining multi-modality learning and a Siamese-based structure, our model can navigate across different modalities and match zero-day variants. To address the missing or noisy modality issue, we propose a Conditional Variable Autoencoder with a Generative Adversarial Network for heightened resolution. We trained the model on 100,000 malware triplet pairs. Our experiments on real-world noisy samples show that the model out-performs the state-of-the-art and can accurately match not only zero-day malware, but also out-of-sample benign binaries of the same category. Christopher Molloy, Jeremy Banks, Steven H. H. Ding, Philippe Charland, Andrew Walenstein, Litao Li |
ICDM | 5 |
| 2021 | Autonomic Security Management for IoT Smart SpacesabstractEmbedded sensors and smart devices have turned the environments around us into smart spaces that could automatically evolve, depending on the needs of users, and adapt to the new conditions. While smart spaces are beneficial and desired in many aspects, they could be compromised and expose privacy, security, or render the whole environment a hostile space in which regular tasks cannot be accomplished anymore. In fact, ensuring the security of smart spaces is a very challenging task due to the heterogeneity of devices, vast attack surface, and device resource limitations. The key objective of this study is to minimize the manual work in enforcing the security of smart spaces by leveraging the autonomic computing paradigm in the management of IoT environments. More specifically, we strive to build an autonomic manager that can monitor the smart space continuously, analyze the context, plan and execute countermeasures to maintain the desired level of security, and reduce liability and risks of security breaches. We follow the microservice architecture pattern and propose a generic ontology named Secure Smart Space Ontology (SSSO) for describing dynamic contextual information in security-enhanced smart spaces. Based on SSSO, we build an autonomic security manager with four layers that continuously monitors the managed spaces, analyzes contextual information and events, and automatically plans and implements adaptive security policies. As the evaluation, focusing on a current BlackBerry customer problem, we deployed the proposed autonomic security manager to maintain the security of a smart conference room with 32 devices and 66 services. The high performance of the proposed solution was also evaluated on a large-scale deployment with over 1.8 million triples. Changyuan Lin, Hamzeh Khazaei, Andrew Walenstein, Andrew J. Malton |
ACM Trans. Internet Things | 3 |
| 2020 | Enterprise Security with Adaptive Ensemble Learning on Cooperation and Interaction PatternsabstractSocial networking research has primarily focused on public social networking services and applications, while rich social interactions in an enterprise setting and their related context has received less attention. In this paper, we focus on using the enterprise social context to augment traditional authentication tools. This is motivated by the emergence of smart mobile devices which introduce ease of remote access to work from almost anywhere and anytime, adding spatio-temporal dimension to the social context. However, it remains a challenge to efficiently manage access-controlled events by using different contextual properties. This paper analyzes specific actions under specific access-control rules to extract context-aware machine learning predictions. Such analysis includes the introduction of three contextual metrics: document shareability, valuation, and user cooperation. Furthermore, these socially-dependent metrics are combined with our Smart Enterprise Access Control (SEAC) technique to achieve authenticity precision of 99% while improving the corresponding efficiency trade-off associated with high and strict security. Kyle Quintal, Burak Kantarci, Melike Erol-Kantarci, Andrew J. Malton, Andrew Walenstein |
CCNC | 5 |
| 2018 | A Survey of Anomaly Detection for Connected Vehicle Cybersecurity and SafetyabstractAnomaly detection techniques have been applied to the challenging problem of ensuring both cybersecurity and safety of connected vehicles. We propose a taxonomy of prior research in this domain. Our proposed taxonomy has 3 overarching dimensions subsuming 9 categories and 38 subcategories. Key observations emerging from the survey are: Real-world datasets are seldom used, but instead, most results are derived from simulations; V2V/V2I communications and in vehicle communication are not considered together; proposed techniques are seldom evaluated against a baseline; safety of the vehicles does not attract as much attention as cybersecurity. Gopi Krishnan Rajbahadur, Andrew J. Malton, Andrew Walenstein, Ahmed E. Hassan |
Intelligent Vehicles Symposium | 3 |
| 2011 | Guest editor's introduction to the special section on source code analysis and manipulation
Sibylle Schupp, Andrew Walenstein |
Softw. Qual. J. | 2 |
| 2002 | Evaluating Theories for Managing Imperfect Knowledge in Human-Centric Database Reengineering EnvironmentsabstractModernizing heavily evolved and poorly documented information systems is a central software engineering problem in our current IT industry. It is often necessary to reverse engineer the design documentation of such legacy systems. Several interactive CASE tools have been developed to support this human-intensive process. However, practical experience indicates that their applicability is limited because they do not adequately handle imperfect knowledge about legacy systems. In this paper, we investigate the applicability of several major theories of imperfect knowledge management in the area of soft computing and approximate reasoning. The theories are evaluated with respect to how well they meet requirements for generating effective human-centred reverse engineering environments. The requirements were elicited with help from practical case studies in the area of database reverse engineering. A particular theory called "possibilistic logic" was found to best meet these requirements most comprehensively. This evaluation highlights important challenges to the designers of knowledge management techniques, and should help reverse engineering tool implementers select appropriate technologies. Jens H. Weber, Andrew Walenstein |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 1998 | Developing the Designer's Toolkit with Software Comprehension ModelsabstractCognitive models of software comprehension are potential sources of theoretical knowledge for tool designers. Although their use in the analysis of existing tools is fairly well established, the literature has shown only limited use of such models for directly developing design ideas. This paper suggests a way of utilizing existing cognitive models of software comprehension to generate design goals and suggest design strategies early in the development cycle. A crucial part of our method is a scheme for explaining the value of tool features by describing the mechanisms that are presumed to underly the expected improvements in task performance. Andrew Walenstein |
ASE | 1 |