Dongrui Zeng

dblp:216/4152 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0003-0032-2571ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 13 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BinType: Type based Indirect Call Target Refinement on Binary Programs
abstract
Constructing precise and sound control flow graphs (CFGs) is critical for enforcing control flow integrity (CFI) defense against control-flow hijacking related exploits. One major challenge of constructing such CFGs is to infer the targets of indirect calls, which suffers extra difficulty on commercial off-the-shelf (COTS) binaries due to the absence of source-level information. One classic direction is to use points-to analysis, but it often suffers from scalability issues. Thus, signature-matching based approaches are employed to mitigate the scalability issues. However, existing binary-level signatures are coarse-grained and the paired matching policy must be conservative to pursue soundness, resulting in CFG precision decrease. In this paper, we present BinType, a new signature-matching approach that relies on type inference to improve the signature granularity. Methodology-wise, BinType identifies storage locations of high-confidence types, generates type equivalence relations between storage locations, and propagates types to callsite arguments and function parameters by following the type equivalence relations. Our evaluations show that BinType achieves >20% higher precision than the previous arity-based technique. Moreover, BinType shows comparable precision against the state-of-the-art points-to analysis based approach, while significantly improving the efficiency.
Sun Hyoung Kim, Dongrui Zeng, Monika Santra, Gang Tan
CODASPY2
2026 Fine-grained information-flow-driven program partitioning
Xue Rao, Cong Sun 0001, Dongrui Zeng
Comput. Secur.3
2025 Disa: Accurate Learning-based Static Disassembly with Attentions
abstract
For reverse engineering related security domains, such as vulnerability detection, malware analysis, and binary hardening, disassembly is crucial yet challenging. The fundamental challenge of disassembly is to identify instruction and function boundaries. Classic approaches rely on file-format assumptions and architecture-specific heuristics to guess the boundaries, resulting in incomplete and incorrect disassembly, especially when the binary is obfuscated. Recent advancements of disassembly have demonstrated that deep learning can improve both the accuracy and efficiency of disassembly. In this paper, we propose Disa, a new learning-based disassembly approach that uses the information of superset instructions over the multi-head self-attention to learn the instructions' correlations, thus being able to infer function entry-points and instruction boundaries. Disa can further identify instructions relevant to memory block boundaries to facilitate an advanced block-memory model based value-set analysis for an accurate control flow graph (CFG) generation. Our experiments show that Disa outperforms prior deep-learning disassembly approaches in function entry-point identification, especially achieving 9.1% and 13.2% F1-score improvement on binaries respectively obfuscated by the disassembly desynchronization technique and popular source-level obfuscator. By achieving an 18.5% improvement in the memory block precision, Disa generates more accurate CFGs with a 4.4% reduction in Average Indirect Call Targets (AICT) compared with the state-of-the-art heuristic-based approach.
Monika Santra, Cong Sun 0001, Dongrui Zeng, Gang Tan
CCS5
2025 Sliver: A Scalable Slicing-Based Verification for Information Flow Security
abstract
Static information flow analysis has been studied for a long time. It is usually considered more precise than dynamic taint analysis and more flexible and indispensable when running individual modules or the entire program is difficult. The state-of-the-art static information flow analyses are scalable on analyzing Java programs or mobile apps, and several type systems have enforced information flow security on different languages. However, static information-flow analyses have rarely scaled up to real-world C programs. This work presents Sliver, a slicing-based approach to verify information flow security on real-world C programs. The principle of Sliver is to convert the information-flow-involved parts of the original program into behavior-equivalent slices and use bounded model checking to enforce the end-to-end noninterference property or detect security violations on the slices after self-composition. We develop automated path-signature-guided slicing and adaptive self-composition approaches to ensure Sliver's efficacy and scalability. We also develop a consistency testing technique and metrics to estimate the correctness of slices generated by Sliver. The evaluations demonstrate Sliver's effectiveness, scalability, and the correctness of the generated slices.
Xue Rao, Cong Sun 0001, Dongrui Zeng, Yongzhe Huang, Gang Tan
IEEE Trans. Dependable Secur. Comput.3
2023 LibScan: Towards More Precise Third-Party Library Identification for Android Applications
Cong Sun 0001, Dongrui Zeng, Gang Tan, Siqi Ma 0001
USENIX Security Symposium3
2023 CryptoEval: Evaluating the risk of cryptographic misuses in Android apps with data-flow analysis
abstract
Abstract The misunderstanding and incorrect configurations of cryptographic primitives have exposed severe security vulnerabilities to attackers. Due to the pervasiveness and diversity of cryptographic misuses, a comprehensive and accurate understanding of how cryptographic misuses can undermine the security of an Android app is critical to the subsequent mitigation strategies but also challenging. Although various approaches have been proposed to detect cryptographic misuse in Android apps, studies have yet to focus on estimating the security risks of cryptographic misuse. To address this problem, the authors present an extensible framework for deciding the threat level of cryptographic misuse in Android apps. Firstly, the authors propose a general and unified specification for representing cryptographic misuses to make our framework extensible and develop adapters to unify the detection results of the state‐of‐the‐art cryptographic misuse detectors, resulting in an adapter‐based detection tool chain for a more comprehensive list of cryptographic misuses. Secondly, the authors employ a misuse‐originating data‐flow analysis to connect each cryptographic misuse to a set of data‐flow sinks in an app, based on which the authors propose a quantitative data‐flow‐driven metric for assessing the overall risk of the app introduced by cryptographic misuses. To make the per‐app assessment more useful for app vetting at the app‐store level, the authors apply unsupervised learning to predict and classify the top risky threats to guide more efficient subsequent mitigation. In the experiments on an instantiated implementation of the framework, the authors evaluate the accuracy of our detection and the effect of data‐flow‐driven risk assessment of our framework. Our empirical study on over 40,000 apps, and the analysis of popular apps reveal important security observations on the real threats of cryptographic misuse in Android apps.
Cong Sun 0001, Xinpeng Xu, Dongrui Zeng, Gang Tan, Siqi Ma 0001
IET Inf. Secur.4
2023 DeepCatra: Learning flow- and graph-based behaviours for Android malware detection
abstract
Abstract As Android malware grows and evolves, deep learning has been introduced into malware detection, resulting in great effectiveness. Recent work is considering hybrid models and multi‐view learning. However, they use only simple features, limiting the accuracy of these approaches in practice. This study proposes DeepCatra, a multi‐view learning approach for Android malware detection, whose model consists of a bidirectional LSTM (BiLSTM) and a graph neural network (GNN) as subnets. The two subnets rely on features extracted from statically computed call traces leading to critical APIs derived from public vulnerabilities. For each Android app, DeepCatra first constructs its call graph and computes call traces reaching critical APIs. Then, temporal opcode features used by the BiLSTM subnet are extracted from the call traces, while flow graph features used by the GNN subnet are constructed from all call traces and inter‐component communications. We evaluate the effectiveness of DeepCatra by comparing it with several state‐of‐the‐art detection approaches. Experimental results on over 18,000 real‐world apps and prevalent malware show that DeepCatra achieves considerable improvement, for example, 2.7%–14.6% on the F1 measure, which demonstrates the feasibility of DeepCatra in practice.
Dongrui Zeng, Cong Sun 0001
IET Inf. Secur.4
2023 ABSLearn: a GNN-based framework for aliasing and buffer-size information retrieval
Ke Liang 0006, Jim Tan, Dongrui Zeng, Yongzhe Huang, Gang Tan
Pattern Anal. Appl.3
2023 μDep: Mutation-Based Dependency Generation for Precise Taint Analysis on Android Native Code
abstract
The existence of native code in Android apps plays an important role in triggering inconspicuous propagation of secrets and circumventing malware detection. However, the state-of-the-art information-flow analysis tools for Android apps all have limited capabilities of analyzing native code. Due to the complexity of binary-level static analysis, most static analyzers choose to build conservative models for a selected portion of native code. Though the recent inter-language analysis improves the capability of tracking information flow in native code, it is still far from attaining similar effectiveness of the state-of-the-art information-flow analyzers that focus on non-native Java methods. To overcome the above constraints, we propose a new analysis framework,$\mu$Dep, to detect sensitive information flows of the Android apps containing native code. In this framework, we combine a control-flow based static binary analysis with a mutation-based dynamic analysis to model the tainting behaviors of native code in the apps. Based on the result of the analyses,$\mu$Dep conducts a stub generation for the related native functions to facilitate the state-of-the-art analyzer DroidSafe with fine-grained tainting behavior summaries of native code. The experimental results show that our framework is competitive on the accuracy, and effective in analyzing the information flows in real-world apps and malware compared with the state-of-the-art inter-language static analysis.
Cong Sun 0001, Yuwan Ma, Dongrui Zeng, Gang Tan, Siqi Ma 0001
IEEE Trans. Dependable Secur. Comput.3
2022 BinPointer: towards precise, sound, and scalable binary-level pointer analysis
abstract
Binary-level pointer analysis is critical to binary-level applications such as reverse engineering and binary debloating. In this paper, we propose BinPointer, a new binary-level interprocedural pointer analysis that relies on an offset-sensitive value-tracking analysis to achieve high precision. We also propose a soundness and precision evaluation methodology based on runtime memory accesses triggered by reference input data. Our experimental results demonstrate that BinPointer has higher precision over prior work, while maintaining acceptable scalability. The soundness of BinPointer is also validated through runtime data.
Sun Hyoung Kim, Dongrui Zeng, Cong Sun 0001, Gang Tan
CC2
2021 ReCFA: Resilient Control-Flow Attestation
abstract
Recent IoT applications gradually adapt more complicated end systems with commodity software. Ensuring the runtime integrity of these software is a challenging task for the remote controller or cloud services. Popular enforcement is the runtime remote attestation which requires the end system (prover) to generate evidence for its runtime behavior and a remote trusted verifier to attest the evidence. Control-flow attestation is a kind of runtime attestation that provides diagnoses towards the remote control-flow hijacking at the prover. Most of these attestation approaches focus on small or embedded software. The recent advance to attesting complicated software depends on the source code and CFG traversing to measure the checkpoint-separated subpaths, which may be unavailable for commodity software and cause possible context missing between consecutive subpaths in the measurements.
Xinzhi Liu, Cong Sun 0001, Dongrui Zeng, Gang Tan, Xiao Kan, Siqi Ma 0001
ACSAC4
2021 Refining Indirect Call Targets at the Binary Level
Sun Hyoung Kim, Cong Sun 0001, Dongrui Zeng, Gang Tan
NDSS3
2021 MazeRunner: Evaluating the Attack Surface of Control-Flow Integrity Policies
abstract
Control-Flow Integrity (CFI) enforces a control-flow graph (CFG) to limit attackers' ability to manipulate runtime control flow. CFI variations, enforcing different CFGs, achieve different degrees of attack surface reduction. To compare the security strength of different CFI policies, measuring the remaining attack surface is critical but challenging. Therefore, we propose MazeRunner, a framework that quantitatively estimates the attack surface of a CFI-hardened program. Methodology-wise, it takes a program's CFG, an attack model, and a security-violation policy as input to discover risky program points by an attack-aware data dependency tracking algorithm. Risky program points and the CFG are used to compute a metric for the remaining attack surface. We evaluate MazeRunner with 3 CFG types, 3 attack models, and 4 security-violation policies against 13 realistic benchmarks, and demonstrate that the new metric achieves higher precision than traditional metrics while maintaining completeness.
Dongrui Zeng, Ben Niu 0007, Gang Tan
TrustCom1
2019 Program-mandering: Quantitative Privilege Separation
abstract
Privilege separation is an effective technique to improve software security. However, past partitioning systems do not allow programmers to make quantitative tradeoffs between security and performance. In this paper, we describe our toolchain called PM. It can automatically find the optimal boundary in program partitioning. This is achieved by solving an integer-programming model that optimizes for a user-chosen metric while satisfying the remaining security and performance constraints on other metrics. We choose security metrics to reason about how well computed partitions enforce information flow control to: (1) protect the program from low-integrity inputs or (2) prevent leakage of program secrets. As a result, functions in the sensitive module that fall on the optimal partition boundaries automatically identify where declassification is necessary. We used PM to experiment on a set of real-world programs to protect confidentiality and integrity; results show that, with moderate user guidance, PM can find partitions that have better balance between security and performance than partitions found by a previous tool that requires manual declassification.
Shen Liu 0002, Dongrui Zeng, Yongzhe Huang, Frank Capobianco, Stephen McCamant, Trent Jaeger, Gang Tan
CCS2
2018 From Debugging-Information Based Binary-Level Type Inference to CFG Generation
abstract
Binary-level Control-Flow Graph (CFG) construction is essential for applications such as control-flow integrity. There are two main approaches: the binary-analysis approach and the compiler-modification approach. The binary-analysis approach does not require source code, but it constructs low-precision CFGs. The compiler-modification approach requires source code and modifies compilers for CFG generation. We describe the design and implementation of an alternative system for high-precision CFG construction, which still assumes source code but does not modify compilers. Our approach makes use of standard compiler-generated meta-information, including symbol tables, relocation information, and debugging information. A key component in the system is a type-inference engine that infers types of low-level storage locations such as registers from types in debugging information. Inferred types enable a type-signature matching method for high-precision CFG construction.
Dongrui Zeng, Gang Tan
CODASPY1