Dongpeng Xu 0001

dblp:158/2764 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0001-6596-9101ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 16 · 3 first-author · 8 since 2021Systems, architecture and hardware · 8 · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MBA-Sniffer: Rapidly Locating Mixed Boolean-Arithmetic Obfuscation in Binary Code
Zheyun Feng, Dongpeng Xu 0001
DSN2
2025 Inspecting Virtual Machine Diversification Inside Virtualization Obfuscation
abstract
Virtualization obfuscators are commonly employed to safeguard proprietary code or to impede malware analysis. Despite significant efforts to combat these obfuscators over the past decade, code virtualization continues to be an exceedingly effective obfuscation technique. At the core of modern virtualization obfuscators are the virtual machines (VMs), which employ a variety of diversification techniques to complicate their internal structures. Due to its intricate and diverse nature, reverse engineering one VM is a time-consuming task and is not useful in cracking other VMs. Yet, despite the success of these VMs, there has been no systematic study of their diversification techniques, creating a knowledge gap that needs to be addressed to enhance VM deobfuscation. This work aims to bridge the above gap. First, we categorize and unveil the techniques under the hood of VM diversification, from the perspectives of VM interpretation, byte-code organization, and handler permutation/relocation. This systematic knowledge about modern virtualization is a crucial contribution to the field. Second, we develop an automated tool to identify the VM diversification techniques adopted by state-of-the-art virtualization obfuscators. The results demystify how the VM diversification methods are deployed in practice. Third, our research also involves patching current deobfuscation tools using the newly revealed knowledge of VM diversification to overcome their weaknesses. This outcome highlights how the results of our study pave the way for next-generation VM deobfuscation.
Naiqian Zhang, Dongpeng Xu 0001, Jiang Ming 0002, Jun Xu 0024, Qiaoyan Yu
SP2
2025 MemoryTrap: Booby Trapping Memory to Counter Memory Disclosure Attacks with Hardware Support
Chenke Luo, Jiang Ming 0002, Dongpeng Xu 0001, Guojun Peng, Jianming Fu
USENIX ATC3
2024 An In-Depth Analysis of the Code-Reuse Gadgets Introduced by Software Obfuscation
Naiqian Zhang, Zheyun Feng, Dongpeng Xu 0001
ACNS (3)3
2024 Feature-driven Approximate Computing for Wearable Health-Monitoring Systems
abstract
Real-time health monitoring systems generate a large volume of sensing data, requiring tremendous processing time and storage space. Orthogonal to existing approximate computing mechanisms, this work proposes a Feature-Driven Approximation (FDApx) method to address the pressing need for fast data processing and a limited storage budget in wearable health monitoring devices. The proposed FDApx method reverses the features interested in high-level applications to derive approximation thresholds to retain feature-critical information, rather than aimlessly storing and transmitting all raw data. Case studies in an insole sensing system for fall risk assessment show that FDApx can reduce the data size by up to 87% over raw data and up to 85% over 2-bit precision reduction-based approximation. The approximation from FDApx only results in up to a 2% deviation in swing time; in contrast, the approximation based on precision reduction causes a 30% deviation in the same gait feature1.
Nishanth Goud Chennagouni, Mashrafi Alam Kajol, Diliang Chen, Dongpeng Xu 0001, Qiaoyan Yu
ACM Great Lakes Symposium on VLSI4
2023 No Free Lunch: On the Increased Code Reuse Attack Surface of Obfuscated Programs
abstract
Obfuscation has been widely employed to protect software from the malicious reverse analysis. However, its security risks have not previously been studied in detail. For example, most obfuscation methods introduce large blocks of opaque code that are black boxes to normal users. In this paper, we show that, indeed, obfuscation can increase the attack risk. Existing gadget search tools, while able to find more gadgets in obfuscated code, do not succeed in assembling them into more exploits. However, these tools use strict pattern matching, greedy searching strategies, and only very simple gadgets. We develop Gadget-Planner, a more flexible approach to building code-reuse attacks that overcomes previous limitations via symbolic execution and automated planning. In a study across both benchmark and real-world programs, this approach finds many more exploit payloads on obfuscated programs, both in terms of number and diversity.
Naiqian Zhang, Daroc Alden, Dongpeng Xu 0001, Shuai Wang 0011, Trent Jaeger, Wheeler Ruml
DSN3
2022 Generating Effective Software Obfuscation Sequences With Reinforcement Learning
abstract
Obfuscation is a prevalent security technique which transforms syntactic representation of a program to a complicated form, but still keeps program semantics unchanged. So far, developers heavily rely on obfuscation to harden their products and reduce the risk of adversarial reverse engineering. However, despite its spectacular progress, one crucial hurdle is that each of existing obfuscation method is designed specifically for obfuscating one program feature (e.g., identifier name, control flow), so an effective obfuscation scheme usually composes a considerable amount of different obfuscation methods. Therefore, one primary challenge lies in identifying effective combinations of obfuscation methods. In this research, we propose a principled technique for generating an optimal program obfuscation scheme by adopting a reinforcement learning approach. Given a program and a set of obfuscation transformations, a reinforcement learning model is progressively trained to select a sequence of obfuscation transformations, such that applying each transformation in order toward the program yields the optimal obfuscation result, making programs dissimilar while retaining reasonable instrumentation overhead. Our implementation can directly work on raw binary executables without source code, and our evaluation demonstrates that the trained models can effectively obfuscate executable files with low cost.
Huaijin Wang 0001, Shuai Wang 0011, Dongpeng Xu 0001, Xiao Liu 0025
IEEE Trans. Dependable Secur. Comput.3
2021 Towards Optimal Use of Exception Handling Information for Function Detection
abstract
Function entry detection is critical for security of binary code. Conventional methods heavily rely on patterns, inevitably missing true functions and introducing errors. Recently, call frames have been used in exception-handling for function start detection. However, existing methods have two problems. First, they combine call frames with heuristic-based approaches, which often brings error and uncertain benefits. Second, they trust the fidelity of call frames, without handling the errors that are introduced by call frames. In this paper, we first study the coverage and accuracy of existing approaches in detecting function starts using call frames. We found that although recursive disassembly with call frames can maximize coverage, using extra heuristic-based approaches does not improve coverage and actually hurts accuracy. Second, we unveil call-frame errors and develop the first approach to fix them, making their use more reliable.
Chengbin Pang, Ruotong Yu, Dongpeng Xu 0001, Eric Koskinen, Georgios Portokalidis, Jun Xu 0024
DSN3
2021 GraphMR: Graph Neural Network for Mathematical Reasoning
abstract
Mathematical reasoning aims to infer satisfiable solutions based on the given mathematics questions.Previous natural language processing researches have proven the effectiveness of sequence-to-sequence (Seq2Seq) or related variants on mathematics solving.However, few works have been able to explore structural or syntactic information hidden in expressions (e.g., precedence and associativity).This dissertation set out to investigate the usefulness of such untapped information for neural architectures.Firstly, mathematical questions are represented in the format of graphs within syntax analysis.The structured nature of graphs allows them to represent relations of variables or operators while preserving the semantics of the expressions.Having transformed to the new representations, we proposed a graph-to-sequence neural network GraphMR, which can effectively learn the hierarchical information of graphs inputs to solve mathematics and speculate answers.A complete experimental scenario with four classes of mathematical tasks and three Seq2Seq baselines is built to conduct a comprehensive analysis, and results show that GraphMR outperforms others in hidden information learning and mathematics resolving.
Weijie Feng, Dongpeng Xu 0001, Qilong Zheng
EMNLP (1)3
2021 Software Obfuscation with Non-Linear Mixed Boolean-Arithmetic Expressions
Weijie Feng, Qilong Zheng, Jing Li 0047, Dongpeng Xu 0001
ICICS (1)5
2021 Boosting SMT solver performance on mixed-bitwise-arithmetic expressions
abstract
Satisfiability Modulo Theories (SMT) solvers have been widely applied in automated software analysis to reason about the queries that encode the essence of program semantics, relieving the heavy burden of manual analysis. Many SMT solving techniques rely on solving Boolean satisfiability problem (SAT), which is an NP-complete problem, so they use heuristic search strategies to seek possible solutions, especially when no known theorem can efficiently reduce the problem. An emerging challenge, named Mixed-Bitwise-Arithmetic (MBA) obfuscation, impedes SMT solving by constructing identity equations with both bitwise operations (and, or, negate) and arithmetic computation (add, minus, multiply). Common math theorems for bitwise or arithmetic computation are inapplicable to simplifying MBA equations, leading to performance bottlenecks in SMT solving.
Dongpeng Xu 0001, Weijie Feng, Jiang Ming 0002, Qilong Zheng, Jing Li 0047, Qiaoyan Yu
PLDI1
2021 MBA-Blast: Unveiling and Simplifying Mixed Boolean-Arithmetic Obfuscation
Junfu Shen, Jiang Ming 0002, Qilong Zheng, Jing Li 0047, Dongpeng Xu 0001
USENIX Security Symposium6
2021 Security Threat Analyses and Attack Models for Approximate Computing Systems: From Hardware and Micro-architecture Perspectives
abstract
Approximate computing (AC) represents a paradigm shift from conventional precise processing to inexact computation but still satisfying the system requirement on accuracy. The rapid progress on the development of diverse AC techniques allows us to apply approximate computing to many computation-intensive applications. However, the utilization of AC techniques could bring in new unique security threats to computing systems. This work does a survey on existing circuit-, architecture-, and compiler-level approximate mechanisms/algorithms, with special emphasis on potential security vulnerabilities. Qualitative and quantitative analyses are performed to assess the impact of the new security threats on AC systems. Moreover, this work proposes four unique visionary attack models, which systematically cover the attacks that build covert channels, compensate approximation errors, terminate normal error resilience mechanisms, and propagate additional errors. To thwart those attacks, this work further offers the guideline of countermeasure designs. Several case studies are provided to illustrate the implementation of the suggested countermeasures.
Pruthvy Yellu, Landon Buell, Miguel Mark, Michel A. Kinsy, Dongpeng Xu 0001, Qiaoyan Yu
ACM Trans. Design Autom. Electr. Syst.5
2020 Security Threats and Countermeasures for Approximate Arithmetic Computing
abstract
Approximate computing (AC) emerges as a promising approach for energy-accuracy trade-off in compute-intensive applications. However, recent work reveals that AC techniques could lead to new security vulnerabilities, which are presented in a format of visionary view. There is a lack of in-depth research on concrete attack models and estimation of the significance of the attacks on approximate arithmetic computing systems. This work presents several practical attack examples and then proposes two attack models with quantitative analysis. Input integrity check and exclusive logic based attack detection methods are proposed to address the attacks on AC systems. The experimental results show that the attack detection failure rate of our method is below $2.2*10^{-3}$ and the area and power overhead is less than 6.8% and 1.5%, respectively.
Pruthvy Yellu, Mohammad Mezanur Rahman Monjur, Timothy Kammerer, Dongpeng Xu 0001, Qiaoyan Yu
ASP-DAC4
2020 VAHunt: Warding Off New Repackaged Android Malware in App-Virtualization's Clothing
abstract
Repackaging popular benign apps with malicious payload used to be the most common way to spread Android malware. Nevertheless, since 2016, we have observed an alarming new trend to Android ecosystem: a growing number of Android malware samples abuse recent app-virtualization innovation as a new distribution channel. App-virtualization enables a user to run multiple copies of the same app on a single device, and tens of millions of users are enjoying this convenience. However, cybercriminals repackage various malicious APK files as plugins into an app-virtualization platform, which is flexible to launch arbitrary plugins without the hassle of installation. This new style of repackaging gains the ability to bypass anti-malware scanners by hiding the grafted malicious payload in plugins, and it also defies the basic premise embodied by existing repackaged app detection solutions.
Luman Shi, Jiang Ming 0002, Jianming Fu, Guojun Peng, Dongpeng Xu 0001, Xuanchen Pan
CCS5
2020 Blurring Boundaries: A New Way to Secure Approximate Computing Systems
abstract
Approximate computing (AC) techniques have been widely used to improve the performance of computing systems by trading off accuracy. However, recent literature projects that the utilization of approximation could bring in new security threats to computing systems. This work presents two practical attacks on the AC systems for multilayer perceptron (MLP) and Sobel algorithm based image edge detection. The case studies in this work indicate that the approximation mechanism in AC systems can be exploited to conduct stealthy attacks, which suddenly cause significant degradation in accuracy and lead to unpredictable primary outputs. To address the emerging threats on AC systems, this work proposes to blur the boundary between approximate and precise computing submodules in AC systems. This new defense method obscures that boundary with three obfuscation schemes such that adversary could not easily identify the right target to precisely perform hardware tampering attacks. Simulation results show that the proposed method can effectively reduce the attack success rate.
Pruthvy Yellu, Landon Buell, Dongpeng Xu 0001, Qiaoyan Yu
ACM Great Lakes Symposium on VLSI3
2019 Memory access integrity: detecting fine-grained memory access errors in binary code
abstract
As one of the most notorious programming errors, memory access errors still hurt modern software security. Particularly, they are hidden deeply in important software systems written in memory unsafe languages like C/C++. Plenty of work have been proposed to detect bugs leading to memory access errors. However, all existing works lack the ability to handle two challenges. First, they are not able to tackle fine-grained memory access errors, e.g., data overflow inside one data structure. These errors are usually overlooked for a long time since they happen inside one memory block and do not lead to program crash. Second, most existing works rely on source code or debugging information to recover memory boundary information, so they cannot be directly applied to detection of memory access errors in binary code. However, searching memory access errors in binary code is a very common scenario in software vulnerability detection and exploitation. In order to overcome these challenges, we propose Memory Access Integrity (MAI), a dynamic method to detect fine-grained memory access errors in off-the-shelf binary executables. The core idea is to recover fine-grained accessing policy between memory access behaviors and memory ranges, and then detect memory access errors based on the policy. The key insight in our work is that memory accessing patterns reveal information for recovering the boundary of memory objects and the accessing policy. Based on these recovered information, our method maintains a new memory model to simulate the life cycle of memory objects and report errors when any accessing policy is violated. We evaluate our tool on popular CTF datasets and real world softwares. Compared with the state of the art detection tool, the evaluation result demonstrates that our tool can detect fine-grained memory access errors effectively and efficiently. As the practical impact, our tool has detected three 0-day memory access errors in an audio decoder.
Wenjie Li 0006, Dongpeng Xu 0001, Xiaorui Gong, Xiaobo Xiang, Fangming Gu, Qianxiang Zeng
Cybersecur.2
2018 VMHunt: A Verifiable Approach to Partially-Virtualized Binary Code Simplification
abstract
Code virtualization is a highly sophisticated obfuscation technique adopted by malware authors to stay under the radar. However, the increasing complexity of code virtualization also becomes a "double-edged sword" for practical application. Due to its performance limitations and compatibility problems, code virtualization is seldom used on an entire program. Rather, it is mainly used only to safeguard the key parts of code such as security checks and encryption keys. Many techniques have been proposed to reverse engineer the virtualized code, but they share some common limitations. They assume the scope of virtualized code is known in advance and mainly focus on the classic structure of code emulator. Also, few work verifies the correctness of their deobfuscation results. In this paper, with fewer assumptions on the type and scope of code virtualization, we present a verifiable method to address the challenge of partially-virtualized binary code simplification. Our key insight is that code virtualization is a kind of process-level virtual machine (VM), and the context switch patterns when entering and exiting the VM can be used to detect the VM boundaries. Based on the scope of VM boundary, we simplify the virtualized code. We first ignore all the instructions in a given virtualized snippet that do not affect the final result of that snippet. To better revert the data obfuscation effect that encodes a variable through bitwise operations, we then run a new symbolic execution called multiple granularity symbolic execution to further simplify the trace snippet. The generated concise symbolic formulas facilitate the correctness testing of our simplification results. We have implemented our idea as an open source tool, VMHunt, and evaluated it with real-world applications and malware. The encouraging experimental results demonstrate that VMHunt is a significant improvement over the state of the art.
Dongpeng Xu 0001, Jiang Ming 0002, Dinghao Wu
CCS1
2017 Cryptographic Function Detection in Obfuscated Binaries via Bit-Precise Symbolic Loop Mapping
abstract
Cryptographic functions have been commonly abused by malware developers to hide malicious behaviors, disguise destructive payloads, and bypass network-based firewalls. Now-infamous crypto-ransomware even encrypts victim's computer documents until a ransom is paid. Therefore, detecting cryptographic functions in binary code is an appealing approach to complement existing malware defense and forensics. However, pervasive control and data obfuscation schemes make cryptographic function identification a challenging work. Existing detection methods are either brittle to work on obfuscated binaries or ad hoc in that they can only identify specific cryptographic functions. In this paper, we propose a novel technique called bit-precise symbolic loop mapping to identify cryptographic functions in obfuscated binary code. Our trace-based approach captures the semantics of possible cryptographic algorithms with bit-precise symbolic execution in a loop. Then we perform guided fuzzing to efficiently match boolean formulas with known reference implementations. We have developed a prototype called CryptoHunt and evaluated it with a set of obfuscated synthetic examples, well-known cryptographic libraries, and malware. Compared with the existing tools, CryptoHunt is a general approach to detecting commonly used cryptographic functions such as TEA, AES, RC4, MD5, and RSA under different control and data obfuscation scheme combinations.
Dongpeng Xu 0001, Jiang Ming 0002, Dinghao Wu
IEEE Symposium on Security and Privacy1
2017 BinSim: Trace-based Semantic Binary Diffing via System Call Sliced Segment Equivalence Checking
Jiang Ming 0002, Dongpeng Xu 0001, Yufei Jiang, Dinghao Wu
USENIX Security Symposium2
2016 Generalized Dynamic Opaque Predicates: A New Control Flow Obfuscation Method
Dongpeng Xu 0001, Jiang Ming 0002, Dinghao Wu
ISC1
2015 LOOP: Logic-Oriented Opaque Predicate Detection in Obfuscated Binary Code
abstract
Opaque predicates have been widely used to insert superfluous branches for control flow obfuscation. Opaque predicates can be seamlessly applied together with other obfuscation methods such as junk code to turn reverse engineering attempts into arduous work. Previous efforts in detecting opaque predicates are far from mature. They are either ad hoc, designed for a specific problem, or have a considerably high error rate. This paper introduces LOOP, a Logic Oriented Opaque Predicate detection tool for obfuscated binary code. Being different from previous work, we do not rely on any heuristics; instead we construct general logical formulas, which represent the intrinsic characteristics of opaque predicates, by symbolic execution along a trace. We then solve these formulas with a constraint solver. The result accurately answers whether the predicate under examination is opaque or not. In addition, LOOP is obfuscation resilient and able to detect previously unknown opaque predicates. We have developed a prototype of LOOP and evaluated it with a range of common utilities and obfuscated malicious programs. Our experimental results demonstrate the efficacy and generality of LOOP. By integrating LOOP with code normalization for matching metamorphic malware variants, we show that LOOP is an appealing complement to existing malware defenses.
Jiang Ming 0002, Dongpeng Xu 0001, Dinghao Wu
CCS2
2015 Memoized Semantics-Based Binary Diffing with Application to Malware Lineage Inference
Jiang Ming 0002, Dongpeng Xu 0001, Dinghao Wu
SEC2