Monika Santra

dblp:205/8912 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0001-6219-6545ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 BinType: Type based Indirect Call Target Refinement on Binary Programs
abstract
Constructing precise and sound control flow graphs (CFGs) is critical for enforcing control flow integrity (CFI) defense against control-flow hijacking related exploits. One major challenge of constructing such CFGs is to infer the targets of indirect calls, which suffers extra difficulty on commercial off-the-shelf (COTS) binaries due to the absence of source-level information. One classic direction is to use points-to analysis, but it often suffers from scalability issues. Thus, signature-matching based approaches are employed to mitigate the scalability issues. However, existing binary-level signatures are coarse-grained and the paired matching policy must be conservative to pursue soundness, resulting in CFG precision decrease. In this paper, we present BinType, a new signature-matching approach that relies on type inference to improve the signature granularity. Methodology-wise, BinType identifies storage locations of high-confidence types, generates type equivalence relations between storage locations, and propagates types to callsite arguments and function parameters by following the type equivalence relations. Our evaluations show that BinType achieves >20% higher precision than the previous arity-based technique. Moreover, BinType shows comparable precision against the state-of-the-art points-to analysis based approach, while significantly improving the efficiency.
Sun Hyoung Kim, Dongrui Zeng, Monika Santra, Gang Tan
CODASPY3
2025 AI-Augmented Static Analysis: Bridging Heuristics and Completeness for Practical Reverse Engineering
abstract
Reverse engineering poses significant challenges for several reasons, including the presence of interleaved code and data, the absence of names, types, and stack frames, aggressive compiler optimizations, and a range of obfuscation techniques. Traditional static analysis methods and existing tools have sought to mitigate the impact of these missing critical elements through various heuristic-based strategies. However, recent advancements in artificial intelligence (AI) have shown promise in tackling these challenges by uncovering complex hidden patterns from incomplete or low-level representations, particularly in predicting high-level semantic constructs that may be lost during compilation. Despite these advancements, AI-only solutions frequently struggle to deliver the completeness and reliability necessary for security-critical binary analysis. To address this shortcoming, we aim to establish an innovative synergy between AI and static analysis—leveraging AI to replace brittle heuristics for enhanced generalization while using static analysis to reinforce AI with a best-effort approach to completeness, thereby meeting the rigorous demands of security applications. In this thesis, we will concentrate on three essential tasks in reverse engineering that are notably underserved in both academic research and existing tools: instruction boundary identification, function boundary identification, and Control Flow Graph (CFG) construction, more specifically for indirect call targets. Our goal is to develop a novel integration of AI and static analysis, creating an end-to-end disassembly framework.
Monika Santra
CCS1
2025 Disa: Accurate Learning-based Static Disassembly with Attentions
abstract
For reverse engineering related security domains, such as vulnerability detection, malware analysis, and binary hardening, disassembly is crucial yet challenging. The fundamental challenge of disassembly is to identify instruction and function boundaries. Classic approaches rely on file-format assumptions and architecture-specific heuristics to guess the boundaries, resulting in incomplete and incorrect disassembly, especially when the binary is obfuscated. Recent advancements of disassembly have demonstrated that deep learning can improve both the accuracy and efficiency of disassembly. In this paper, we propose Disa, a new learning-based disassembly approach that uses the information of superset instructions over the multi-head self-attention to learn the instructions' correlations, thus being able to infer function entry-points and instruction boundaries. Disa can further identify instructions relevant to memory block boundaries to facilitate an advanced block-memory model based value-set analysis for an accurate control flow graph (CFG) generation. Our experiments show that Disa outperforms prior deep-learning disassembly approaches in function entry-point identification, especially achieving 9.1% and 13.2% F1-score improvement on binaries respectively obfuscated by the disassembly desynchronization technique and popular source-level obfuscator. By achieving an 18.5% improvement in the memory block precision, Disa generates more accurate CFGs with a 4.4% reduction in Average Indirect Call Targets (AICT) compared with the state-of-the-art heuristic-based approach.
Monika Santra, Cong Sun 0001, Dongrui Zeng, Gang Tan
CCS2
2025 D24D: Dynamic Deep 4-Dimensional Analysis for Malware Detection
abstract
In the era of ubiquitous computing devices, malware is the primary weapon of cyber attacks, and malware-related security breaches remain a significant security concern. Nowadays, adversaries require fewer resources to exploit a system with the help of contemporary malicious payloads and AI tools than in the old days. Despite many advances in malware defense research, adversaries continually employ sophisticated tools and techniques to evade existing defense mechanisms and create chaos. Moreover, it is challenging to recognize these malicious binaries with shallow features such as section names, entropies, virtual sizes, and strings, which are not robust. The proposed work mainly focuses on identifying robust features that can help to detect more sophisticated (i) seen and (ii) never-seen-before malware effectively. Unlike the existing research works,$D^{2}4D$concentrates on four types of analysis: Registry key, API function, network, and memory analysis. Above all,$D^{2}4D$identifies the binaries that perform fast-flux attacks, DGA-based attacks, homoglyphs attacks, and other attack types. The evaluation results indicate that the$D^{2}4D$achieves an accuracy of 99.67%, with a 0.10% False Positive Rate for seen binaries and more than 91% accuracy for never-seen-before binaries. Beyond that,$D^{2}4D$outperforms 33 existing anti-malware. The extracted features prove robust in identifying seen and never-seen-before binaries based on the experimental analysis, comparison with the state-of-the-art models, and ablation study.
Rama Krishna Koppanati, Monika Santra, Sateesh Kumar Peddoju
IEEE Trans. Inf. Forensics Secur.2