Sima Arasteh

dblp:228/0555 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0000-5950-8245ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 TRM: An Efficient Hypervisor-Based Framework For Malware Analysis and Memory Reconstruction
abstract
Modern rootkits leverage kernel privileges to hide from analysis tools, while obfuscation techniques render them resistant to static analysis. Reverse engineering malware requires observing memory usage and reconstructing data structures. Existing tools rely on instrumentation or emulation, which introduce high overhead, leave detectable artifacts, and cannot reliably analyze kernel-level malware. We present The Reversing Machine (TRM), a hypervisor-based framework for high-performance introspection of evasive malware. It is the first system to support selective memory tracing and structure reconstruction in the hypervisor. TRM repurposes hardware virtualization to efficiently detect user-kernel mode transitions and obtain memory traces independently of a potentially compromised guest kernel. TRM introduces new insights into leveraging hardware virtualization for runtime memory reconstruction and analysis of data structures while remaining invisible to malware. We demonstrate automatic reconstruction of function signatures and data structures, reduce system call latency overhead from 142% to 57% compared to prior work, and accelerate manual reverse engineering by 43% on average, even for complex kernel objects. TRM shows that hypervisor-level memory tracing makes data structure reconstruction practical in hostile environments, bridging the gap between prior feasibility studies and real-world malware analysis.
Mohammad Sina Karvandi, Soroush Meghdadi Zanjani, Sima Arasteh, Saleh Khalaj Monfared, Mohammad K. Fallah, Saeid Gorgin 0001, Jeong-A Lee, Asia Slowinska, Erik van der Kouwe
AsiaCCS3
2026 Data Flows in You: Benchmarking and Improving Static Data-flow Analysis on Binary Executables
abstract
Data-flow analysis is a critical component of security research. Theoretically, accurate data-flow analysis in binary executables is an undecidable problem, due to complexities of binary code. Practically, many binary analysis engines offer some data-flow analysis capability, but we lack understanding of the accuracy of these analyses, and their limitations. We address this problem by introducing a labeled benchmark data set, including 215, 072 microbenchmark test cases, mapping to 277, 072 binary executables, created specifically to evaluate data-flow analysis implementations. Additionally, we augment our benchmark set with dynamically-discovered data flows from 6 real-world executables. Using our benchmark data set, we evaluate three state of the art data-flow analysis implementations, in angr, Ghidra and Miasm and discuss their very low accuracy and reasons behind it. We further propose three model extensions to static data-flow analysis that significantly improve accuracy, achieving almost perfect recall (0.99) and increasing precision from 0.13 to 0.32 for the GCC compiler, and achieving a recall of 0.86 with a precision increase from 0.12 to 0.22 for Clang. Finally, we show that leveraging these model extensions in a vulnerability-discovery context leads to a tangible improvement in vulnerable instruction identification.
Nicolaas Weideman, Sima Arasteh, Mukund Raghothaman, Jelena Mirkovic, Christophe Hauser
AsiaCCS2
2024 BinHunter: A Fine-Grained Graph Representation for Localizing Vulnerabilities in Binary Executables*
abstract
The success of deep learning techniques in diverse fields has prompted research into their application for automatic software vulnerability discovery. The first step in the design of a deep learning based vulnerability detector fundamentally involves selecting an appropriate binary representation. A second challenge arises from the need to automatically localize the vulnerability to specific instructions, so as to allow for better detection and to enable downstream applications such as triage and patching.In this paper, we propose BinHunter, an automated tool for vulnerability discovery in binary programs. BinHunter leverages a new graph representation derived from slices of the combined control and data dependency graphs of a binary executable, and can learn code properties by propagating information through the graph edges. This representation enables graph convolutional network (GCN) learning algorithms to both detect and pinpoint the locations of vulnerabilities in binary programs.We evaluate our approach both using the Juliet test suite and a dataset consisting of historical CVEs from the Debian packages. In both evaluations, we observe that BinHunter is significantly more effective than the baselines: On the Juliet test programs, our model has 6.77%, 26.53%, 24.65% and 41.59% higher true positive rates and 19%, 47.64%, 31.47% and 39.82% lower false positive rates than our baselines respectively (Bin2vec [1], Asm2vec [10], Genius [12] and Jtrans [46]). Furthermore, our model is able to detect 17 of 21 bugs from the Debian dataset, Bin2vec detects 2 bugs, and the remaining three baselines are unable to detect any vulnerabilities at all.
Sima Arasteh, Jelena Mirkovic, Mukund Raghothaman, Christophe Hauser
ACSAC1
2021 Bin2vec: learning representations of binary executable programs for security tasks
abstract
Abstract Tackling binary program analysis problems has traditionally implied manually defining rules and heuristics, a tedious and time consuming task for human analysts. In order to improve automation and scalability, we propose an alternative direction based on distributed representations of binary programs with applicability to a number of downstream tasks. We introduce Bin2vec, a new approach leveraging Graph Convolutional Networks (GCN) along with computational program graphs in order to learn a high dimensional representation of binary executable programs. We demonstrate the versatility of this approach by using our representations to solve two semantically different binary analysis tasks – functional algorithm classification and vulnerability discovery. We compare the proposed approach to our own strong baseline as well as published results, and demonstrate improvement over state-of-the-art methods for both tasks. We evaluated Bin2vec on 49191 binaries for the functional algorithm classification task, and on 30 different CWE-IDs including at least 100 CVE entries each for the vulnerability discovery task. We set a new state-of-the-art result by reducing the classification error by 40% compared to the source-code based inst2vec approach, while working on binary code. For almost every vulnerability class in our dataset, our prediction accuracy is over 80% (and over 90% in multiple classes).
Shushan Arakelyan, Sima Arasteh, Christophe Hauser, Erik Kline, Aram Galstyan
Cybersecur.2
2018 Security analysis of a key based color image watermarking vs. a non-key based technique in telemedicine applications
Pegah Nikbakht Bideh, Shahram Etemadi Borujeni, Sima Arasteh
Multim. Tools Appl.4