VLDB 2026 Research / reviewers in the wild / expert
Jiang Ming 0002
dblp:09/6585-2
· DBLP profile ↗
55ranked-venue papers
9as first author
27since 2021 · last 2026
0000-0001-9682-0502ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 42 · 6 first-author · 21 since 2021Software engineering, systems software and programming languages · 8 · 2 first-author · 4 since 2021Computer networks · 3 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PadNet: Defending Neural Networks Against Adversarial ExamplesabstractMachine learning (ML) suffers from a persistent and critical flaw: adversarial examples. Many new forms of adversarial example attacks have been invented and many narrow defenses have been proposed. Unfortunately, no defensive approach can withstand current attacks. We hypothesize that ML model robustness can be improved with approaches that delineate the data-point-sparse latent space between data-dense regions of a model’s classification space as a barrier class. We introduce one such defense, PadNet, that builds a barrier class using a combination of training samples that mix multiple classes together. It leverages this barrier class to separate decision boundaries between benign classes with regions of padding. PadNet then implements a gradient regularization strategy that penalizes gradients in the direction of the barrier class, causing the decision boundary to draw tighter around training samples increasing boundary thickness between classes. We evaluate PadNet against a sampling of the most effective state-of-the-art attacks, demonstrating that it offers significant robustness and reliability compared to current defenses. We also test it against adaptive attacks and find that PadNet remains robust against them. Armon Barton, Matthew Wright 0001, Shaikh Akib Shahriyar, Edgar W. Jatho III, Mohammad Saidur Rahman 0002, Kantha Girish Gangadhara, Jiang Ming 0002 |
ACM Trans. Priv. Secur. | 7 |
| 2025 | Adversarially Robust Assembly Language Model for Packed Executables DetectionabstractDetecting packed executables is a critical component of large-scale malware analysis and antivirus engine workflows, as it identifies samples that warrant computationally intensive dynamic unpacking to reveal concealed malicious behavior. Traditionally, packer detection techniques have relied on empirical features, such as high entropy or specific binary patterns. However, these empirical, feature-based methods are increasingly vulnerable to evasion by adversarial samples or unknown packers (e.g., low-entropy packers). Furthermore, the dependence on expert-crafted features poses challenges in sustaining and evolving these methods over time. Shijia Li, Jiang Ming 0002, Lanqing Liu, Longwei Yang, Chunfu Jia |
CCS | 2 |
| 2025 | Analyzing PDFs like Binaries: Adversarially Robust PDF Malware Analysis via Intermediate Representation and Language ModelabstractMalicious PDF files have emerged as a persistent threat and become a popular attack vector in web-based attacks. While machine learning-based PDF malware classifiers have shown promise, these classifiers are often susceptible to adversarial attacks, undermining their reliability. To address this issue, recent studies have aimed to enhance the robustness of PDF classifiers. Despite these efforts, the feature engineering underlying these studies remains outdated. Consequently, even with the application of cutting-edge machine learning techniques, these approaches fail to fundamentally resolve the issue of feature instability. To tackle this, we propose a novel approach for PDF feature extraction and PDF malware detection. We introduce the PDFObj IR (PDF Object Intermediate Representation), an assembly-like language framework for PDF objects, from which we extract semantic features using a pretrained language model. Additionally, we construct an Object Reference Graph to capture structural features, drawing inspiration from program analysis. This dual approach enables us to analyze and detect PDF malware based on both semantic and structural features. Experimental results demonstrate that our proposed classifier achieves strong adversarial robustness while maintaining an exceptionally low false positive rate of only 0.07% on baseline dataset compared to state-of-the-art PDF malware classifiers. Side Liu, Jiang Ming 0002, Guodong Zhou 0002, Jianming Fu, Guojun Peng |
CCS | 2 |
| 2025 | Retrofitting XoM for Stripped Binaries without Embedded Data Relocation
Chenke Luo, Jiang Ming 0002, Mengfei Xie, Guojun Peng, Jianming Fu |
NDSS | 2 |
| 2025 | Inspecting Virtual Machine Diversification Inside Virtualization ObfuscationabstractVirtualization obfuscators are commonly employed to safeguard proprietary code or to impede malware analysis. Despite significant efforts to combat these obfuscators over the past decade, code virtualization continues to be an exceedingly effective obfuscation technique. At the core of modern virtualization obfuscators are the virtual machines (VMs), which employ a variety of diversification techniques to complicate their internal structures. Due to its intricate and diverse nature, reverse engineering one VM is a time-consuming task and is not useful in cracking other VMs. Yet, despite the success of these VMs, there has been no systematic study of their diversification techniques, creating a knowledge gap that needs to be addressed to enhance VM deobfuscation. This work aims to bridge the above gap. First, we categorize and unveil the techniques under the hood of VM diversification, from the perspectives of VM interpretation, byte-code organization, and handler permutation/relocation. This systematic knowledge about modern virtualization is a crucial contribution to the field. Second, we develop an automated tool to identify the VM diversification techniques adopted by state-of-the-art virtualization obfuscators. The results demystify how the VM diversification methods are deployed in practice. Third, our research also involves patching current deobfuscation tools using the newly revealed knowledge of VM diversification to overcome their weaknesses. This outcome highlights how the results of our study pave the way for next-generation VM deobfuscation. Naiqian Zhang, Dongpeng Xu 0001, Jiang Ming 0002, Jun Xu 0024, Qiaoyan Yu |
SP | 3 |
| 2025 | MemoryTrap: Booby Trapping Memory to Counter Memory Disclosure Attacks with Hardware Support
Chenke Luo, Jiang Ming 0002, Dongpeng Xu 0001, Guojun Peng, Jianming Fu |
USENIX ATC | 2 |
| 2025 | VAPD: An Anomaly Detection Model for PDF Malware Forensics with Adversarial Robustness
Side Liu, Jiang Ming 0002, Jianming Fu, Guojun Peng |
USENIX Security Symposium | 2 |
| 2025 | The hidden complexities of Android TPL detection: An empirical analysis of techniques, challenges, and effectiveness
Lige Zhan, Jiang Ming 0002, Jianming Fu, Guojun Peng, Letian Sha, Lili Lan |
Comput. Secur. | 2 |
| 2025 | PredicTor: A Global, Machine Learning Approach to Tor Path SelectionabstractTor users derive anonymity in part from the size of the Tor user base, but Tor struggles to attract and support more users due to performance limitations. Previous works have proposed modifications to Tor’s path selection algorithm to enhance both performance and security, but many proposals have unintended consequences due to incorporating information related to client location. We instead propose selecting paths using a global view of the network, independent of client location, and we propose doing so with a machine learning classifier to predict the performance of a given path before building a circuit. We show through a variety of simulated and live experimental settings, across different time periods, that this approach can significantly improve performance compared to Tor’s default path selection algorithm and two previously proposed approaches. In addition to evaluating the security of our approach with traditional metrics, we propose a novel anonymity metric that captures information leakage resulting from location-aware path selection, and we show that our path selection approach leaks no more information than the default path selection algorithm. Armon Barton, Timothy Walsh 0002, Mohsen Imani, Jiang Ming 0002, Matthew Wright 0001 |
ACM Trans. Priv. Secur. | 4 |
| 2024 | Towards Intelligent Automobile Cockpit via A New Container Architecture
Jiang Ming 0002 |
NSDI | 3 |
| 2023 | PackGenome: Automatically Generating Robust YARA Rules for Accurate Malware Packer DetectionabstractBinary packing, a widely-used program obfuscation style, compresses or encrypts the original program and then recovers it at runtime. Packed malware samples are pervasive---they conceal arresting code features as unintelligible data to evade detection. To rapidly respond to large-scale packed malware, security analysts search specific binary patterns to identify corresponding packers. The quality of such packer patterns or signatures is vital to malware dissection. However, existing packer signature rules severely rely on human analysts' experience. In addition to expensive manual efforts, these human-written rules (e.g., YARA) also suffer from high false positives: as they are designed to search the pattern of bytes rather than instructions, they are very likely to mismatch with unexpected instructions. Shijia Li, Jiang Ming 0002, Pengda Qiu, Qiyuan Chen 0006, Lanqing Liu, Huaifeng Bao, Qiang Wang 0059, Chunfu Jia |
CCS | 2 |
| 2023 | Intelligent Zigbee Protocol Fuzzing via Constraint-Field Dependency Inference
Mengfei Ren 0001, Haotian Zhang 0006, Xiaolei Ren 0001, Jiang Ming 0002, Yu Lei 0001 |
ESORICS (2) | 4 |
| 2023 | Assessing Risk in High Performance Computing Attacks
Erika A. Leal, Cimone Wright-Hamor, Joseph B. Manzano, Nicholas J. Multari, Kevin J. Barker, David O. Manz, Jiang Ming 0002 |
ICISSP | 7 |
| 2023 | Leveraging Hardware Performance Counters for Efficient Classification of Binary PackersabstractThe detection and classification of packers serve as a fundamental approach in the study of malware unpacking. In our research, we employ Hardware Performance Counters (HPCs) as features for our classification process. Hardware Performance Counters are integrated into a processor’s Performance Monitoring Unit. The advantages of utilizing hardware performance counters include low overhead access and obviating the need for source code. By selecting hardware features relevant to the unpacking process, we train machine learning classifiers in a supervised learning manner. In our study, we investigate the use of hardware performance counters for the classification purpose of binary packers. Our findings indicate that when configured to alleviate nondeterminism, Hardware Performance Counters possess the capability and potential to classify prevalent packers that are readily accessible. Erika A. Leal, Binlin Cheng, TuQuynh Nguyen, Alfredo Gutierrez Garcia, Nathan Cabero, Jiang Ming 0002 |
TrustCom | 6 |
| 2023 | On the Feasibility of Malware Unpacking via Hardware-assisted Loop Profiling
Binlin Cheng, Erika A. Leal, Haotian Zhang 0006, Jiang Ming 0002 |
USENIX Security Symposium | 4 |
| 2023 | Capturing Invalid Input Manipulations for Memory Corruption DiagnosisabstractMemory corruption diagnosis, especially at the binary level where all high-level program abstractions are missing, is a tedious and time-consuming task. Given a crash, memory corruption diagnosis is expected to not only locate the root cause of the vulnerability, but also deliver rich semantics to understand the vulnerability. However, existing techniques can barely satisfy the above requirements. In this article, we present${{\sf MemRay}}$, a dynamic memory corruption diagnosis technique. The insight behind our approach is that most memory corruption is caused by malformed inputs, which further leads the vulnerable program to manipulate inputs by referencing invalid data structures. We design the “data structure reference sequence” to characterize how a program references various data structures to manipulate program inputs. Then, we identify memory corruptions by detecting violations in the input manipulations via data structures. We demonstrate the effectiveness of${{\sf MemRay}}$on a wide range of memory-corruption vulnerabilities. The result shows that${{\sf MemRay}}$precisely locates the root cause of vulnerabilities. Moreover, the “data structure reference” enables${{\sf MemRay}}$to deliver rich semantics and context information to assist vulnerability diagnosis on binary code. Lei Zhao 0012, Keyang Jiang, Yuncong Zhu, Lina Wang 0001, Jiang Ming 0002 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2023 | Reverse Engineering of Obfuscated Lua Bytecode via Interpreter Semantics TestingabstractAs an efficient and multi-platform scripting language, Lua is gaining increasing popularity in the industry. Unfortunately, Lua’s unique advantages also catch cybercriminals’ attention. A growing number of IoT malware authors switch to Lua for malicious payload development and then distribute malware in bytecode form. To impede malware code analysis, malware authors obfuscate standard Lua bytecode into a customized bytecode specification. Only the attached interpreter can execute that particular bytecode file. Rapid recovery of Lua obfuscated bytecode is essential for a swift response to new malware threats. However, existing generic code deobfuscation approaches cannot keep up with the pace of emerging threats. In this paper, we present a novel reverse engineering technique, calledinterpreter semantics testing. Given a customized interpreter used to execute obfuscated Lua bytecode, we construct a set ofLuaGadgetsthat can adapt to the customized interpreter. Each LuaGadget contains a carefully chosen opcode sequence to fulfill an observable calculation—it is designed to test one or two particular opcodes at a time. Next, we mutate unknown opcode values to generate a bunch of test cases and run them using the customized interpreter; we can observe the expected result only when the mutation hits the opcode’s right value. We perform test case prioritization to cost-effectively recover the semantics of all obfuscated opcodes. Our approach makes no assumptions about the interpreter’s structure and is free from analyzing the numerous execution traces of opcode handlers. We have evaluated our tool,LuaHunt, with Lua malware variants and real-world applications. LuaHunt is able to recover the obfuscated bytecode’s semantics within 90 seconds for each test case, and all of our deobfuscation results can pass the correctness testing. The encouraging results demonstrate that LuaHunt is a promising tool to lighten the burden of security analysts. Chenke Luo, Jiang Ming 0002, Jianming Fu, Guojun Peng, Zhetao Li |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | One size does not fit all: security hardening of MIPS embedded systems via static binary debloating for shared librariesabstractEmbedded systems have become prominent targets for cyberattacks. To exploit firmware’s memory corruption vulnerabilities, cybercriminals harvest reusable code gadgets from the large shared library codebase (e.g., uClibc). Unfortunately, unlike their desktop counterparts, embedded systems lack essential computing resources to enforce security hardening techniques. Recently, we have witnessed a surge of software debloating as a new defense mechanism against code-reuse attacks; it erases unused code to significantly diminish the possibilities of constructing reusable gadgets. Because of the single firmware image update style, static library debloating shows promise to fortify embedded systems without compromising performance and forward compatibility. However, static library debloating on stripped binaries (e.g., firmware’s shared libraries) is still an enormous challenge. In this paper, we show that this challenge is not insurmountable for MIPS firmware. We develop a novel system, named uTrimmer, to identify and wipe out unused basic blocks from shared libraries’ binary code, without causing additional runtime overhead or memory consumption. We propose a new method to identify address-taken blocks/functions, which further help us maintain an inter-procedural control flow graph to conservatively include library code that could be potentially used by firmware. By capturing address access patterns for position-independent code, we circumvent the challenge of determining code-pointer targets and safely eliminate unused code. We run uTrimmer to debloat shared libraries for SPEC CPU2017 benchmarks, popular firmware applications (e.g., Apache, BusyBox, and OpenSSL), and a real-world wireless router firmware image. Our experiments show that not only does uTrimmer deliver functional programs, but also it can cut the exposed code surface and eliminate various reusable code gadgets remarkably. uTrimmer’s debloating capability can compete with the static linking results. Haotian Zhang 0006, Mengfei Ren 0001, Yu Lei 0001, Jiang Ming 0002 |
ASPLOS | 4 |
| 2022 | Chosen-Instruction Attack Against Commercial Code Virtualization Obfuscators
Shijia Li, Chunfu Jia, Pengda Qiu, Qiyuan Chen 0006, Jiang Ming 0002, Debin Gao |
NDSS | 5 |
| 2022 | PolyCruise: A Cross-Language Dynamic Information Flow Analysis
Wen Li 0007, Jiang Ming 0002, Xiapu Luo, Haipeng Cai |
USENIX Security Symposium | 2 |
| 2021 | Towards Transparent and Stealthy Android OS Sandboxing via Customizable Container-Based VirtualizationabstractA fast-growing demand from smartphone users is mobile virtualization.This technique supports running separate instances of virtual phone environments on the same device. In this way, users can run multiple copies of the same app simultaneously,and they can also run an untrusted app in an isolated virtual phone without causing damages to other apps. Traditional hypervisor-based virtualization is impractical to resource-constrained mobile devices.Recent app-level virtualization efforts suffer from the weak isolation mechanism. In contrast, container-based virtualization offers an isolated virtual environment with superior performance.However, existing Android containers do not meet the anti-evasion requirement for security applications: their designs are inherently incapable of providing transparency or stealthiness. Wenna Song, Jiang Ming 0002, Xuanchen Pan, Jianming Fu, Guojun Peng |
CCS | 2 |
| 2021 | App's Auto-Login Function Security Testing via Android OS-Level VirtualizationabstractLimited by the small keyboard, most mobile apps support the automatic login feature for better user experience. Therefore, users avoid the inconvenience of retyping their ID and password when an app runs in the foreground again. However, this auto-login function can be exploited to launch the so-called "data-clone attack": once the locally-stored, auto-login depended data are cloned by attackers and placed into their own smartphones, attackers can break through the login-device number limit and log in to the victim's account stealthily. A natural countermeasure is to check the consistency of device-specific attributes. As long as the new device shows different device fingerprints with the previous one, the app will disable the auto-login function and thus prevent data-clone attacks. In this paper, we develop VPDroid, a transparent Android OS-level virtualization platform tailored for security testing. With VPDroid, security analysts can customize different device artifacts, such as CPU model, Android ID, and phone number, in a virtual phone without user-level API hooking. VPDroid's isolation mechanism ensures that user-mode apps in the virtual phone cannot detect device-specific discrepancies. To assess Android apps' susceptibility to the data-clone attack, we use VPDroid to simulate data-clone attacks with 234 most-downloaded apps. Our experiments on five different virtual phone environments show that VPDroid's device attribute customization can deceive all tested apps that perform device-consistency checks, such as Twitter, WeChat, and PayPal. 19 vendors have confirmed our report as a zero-day vulnerability. Our findings paint a cautionary tale: only enforcing a device-consistency check at client side is still vulnerable to an advanced data-clone attack. Wenna Song, Jiang Ming 0002, Han Yan 0013, Jianming Fu, Guojun Peng |
ICSE | 2 |
| 2021 | Unleashing the hidden power of compiler optimization on binary code difference: an empirical studyabstractHunting binary code difference without source code (i.e., binary diffing) has compelling applications in software security. Due to the high variability of binary code, existing solutions have been driven towards measuring semantic similarities from syntactically different code. Since compiler optimization is the most common source contributing to binary code differences in syntax, testing the resilience against the changes caused by different compiler optimization settings has become a standard evaluation step for most binary diffing approaches. For example, 47 top-venue papers in the last 12 years compared different program versions compiled by default optimization levels (e.g., -Ox in GCC and LLVM). Although many of them claim they are immune to compiler transformations, it is yet unclear about their resistance to non-default optimization settings. Especially, we have observed that adversaries explored non-default compiler settings to amplify malware differences. Xiaolei Ren 0001, Michael Ho, Jiang Ming 0002, Yu Lei 0001, Li Li 0029 |
PLDI | 3 |
| 2021 | Boosting SMT solver performance on mixed-bitwise-arithmetic expressionsabstractSatisfiability Modulo Theories (SMT) solvers have been widely applied in automated software analysis to reason about the queries that encode the essence of program semantics, relieving the heavy burden of manual analysis. Many SMT solving techniques rely on solving Boolean satisfiability problem (SAT), which is an NP-complete problem, so they use heuristic search strategies to seek possible solutions, especially when no known theorem can efficiently reduce the problem. An emerging challenge, named Mixed-Bitwise-Arithmetic (MBA) obfuscation, impedes SMT solving by constructing identity equations with both bitwise operations (and, or, negate) and arithmetic computation (add, minus, multiply). Common math theorems for bitwise or arithmetic computation are inapplicable to simplifying MBA equations, leading to performance bottlenecks in SMT solving. Dongpeng Xu 0001, Weijie Feng, Jiang Ming 0002, Qilong Zheng, Jing Li 0047, Qiaoyan Yu |
PLDI | 4 |
| 2021 | Obfuscation-Resilient Executable Payload Extraction From Packed Malware
Binlin Cheng, Jiang Ming 0002, Erika A. Leal, Haotian Zhang 0006, Jianming Fu, Guojun Peng, Jean-Yves Marion |
USENIX Security Symposium | 2 |
| 2021 | MBA-Blast: Unveiling and Simplifying Mixed Boolean-Arithmetic Obfuscation
Junfu Shen, Jiang Ming 0002, Qilong Zheng, Jing Li 0047, Dongpeng Xu 0001 |
USENIX Security Symposium | 3 |
| 2021 | Z-Fuzzer: device-agnostic fuzzing of Zigbee protocol implementationabstractWith the proliferation of the Internet of Things (IoT) devices, Zigbee is widely adopted as a resource-efficient wireless protocol. Recently, severe vulnerabilities in Zigbee protocol implementations have compromised IoT devices from different manufacturers. It becomes imperative to perform security testing on Zigbee protocol implementations. However, it is not a trivial task to apply the existing vulnerability detection techniques such as fuzzing to Zigbee protocol implementations. In particular, it remains a significant obstacle to deal with low-level hardware events. Many existing protocol fuzzing tools lack a proper execution environment for the Zigbee protocol, which communicates via a radio channel instead of the Internet. Mengfei Ren 0001, Xiaolei Ren 0001, Huadong Feng, Jiang Ming 0002, Yu Lei 0001 |
WISEC | 4 |
| 2020 | Device-agnostic Firmware Execution is Possible: A Concolic Execution Approach for Peripheral EmulationabstractWith the rapid proliferation of IoT devices, our cyberspace is nowadays dominated by billions of low-cost computing nodes, which are very heterogeneous to each other. Dynamic analysis, one of the most effective approaches to finding software bugs, has become paralyzed due to the lack of a generic emulator capable of running diverse previously-unseen firmware. In recent years, we have witnessed devastating security breaches targeting low-end microcontroller-based IoT devices. These security concerns have significantly hamstrung further evolution of the IoT technology. In this work, we present Laelaps, a device emulator specifically designed to run diverse software of microcontroller devices. We do not encode into our emulator any specific information about a device. Instead, Laelaps infers the expected behavior of firmware via symbolic-execution-assisted peripheral emulation and generates proper inputs to steer concrete execution on the fly. This unique design feature makes Laelaps capable of running diverse firmware with no a priori knowledge about the target device. To demonstrate the capabilities of Laelaps, we applied dynamic analysis techniques on top of our emulator. We successfully identified both self-injected and real-world vulnerabilities. Chen Cao 0004, Le Guan, Jiang Ming 0002, Peng Liu 0005 |
ACSAC | 3 |
| 2020 | VAHunt: Warding Off New Repackaged Android Malware in App-Virtualization's ClothingabstractRepackaging popular benign apps with malicious payload used to be the most common way to spread Android malware. Nevertheless, since 2016, we have observed an alarming new trend to Android ecosystem: a growing number of Android malware samples abuse recent app-virtualization innovation as a new distribution channel. App-virtualization enables a user to run multiple copies of the same app on a single device, and tens of millions of users are enjoying this convenience. However, cybercriminals repackage various malicious APK files as plugins into an app-virtualization platform, which is flexible to launch arbitrary plugins without the hassle of installation. This new style of repackaging gains the ability to bypass anti-malware scanners by hiding the grafted malicious payload in plugins, and it also defies the basic premise embodied by existing repackaged app detection solutions. Luman Shi, Jiang Ming 0002, Jianming Fu, Guojun Peng, Dongpeng Xu 0001, Xuanchen Pan |
CCS | 2 |
| 2020 | PatchScope: Memory Object Centric Patch DiffingabstractSoftware patching is one of the most significant mechanisms to combat vulnerabilities. To demystify underlying patch details, the techniques of patch differential analysis (a.k.a. patch diffing) are proposed to find differences between patched and unpatched programs' binary code. Considering the sophisticated security patches, patch diffing is expected to not only correctly locate patch changes but also provide sufficient explanation for understanding patch details and the fixed vulnerabilities. Unfortunately, none of the existing patch diffing techniques can meet these requirements. In this study, we first perform a large-scale study on code changes of security patches for better understanding their patterns. We then point out several challenges and design principles for patch diffing. To address the above challenges, we design a dynamic patch diffing technique PatchScope. Our technique is motivated by two key observations: 1) the way that a program processes its input reveals a wealth of semantic information, and 2) most memory corruption patches regulate the handling of malformed inputs via updating the manipulations of input-related data structures. The core of PatchScope is a new semantics-aware program representation, memory object access sequence, which characterizes how a program references data structures to manipulate inputs. The representation can not only deliver succinct patch differences but also offer rich patch context information such as input-patch correlations. Such information can interpret patch differences and further help security analysts understand patch details, locate vulnerability root causes, and even detect buggy patches. Lei Zhao 0012, Yuncong Zhu, Jiang Ming 0002, Haotian Zhang 0006, Heng Yin 0001 |
CCS | 3 |
| 2020 | Layered obfuscation: a taxonomy of software obfuscation techniques for layered securityabstractAbstract Software obfuscation has been developed for over 30 years. A problem always confusing the communities is what security strength the technique can achieve. Nowadays, this problem becomes even harder as the software economy becomes more diversified. Inspired by the classic idea of layered security for risk management, we propose layered obfuscation as a promising way to realize reliable software obfuscation. Our concept is based on the fact that real-world software is usually complicated. Merely applying one or several obfuscation approaches in an ad-hoc way cannot achieve good obscurity. Layered obfuscation, on the other hand, aims to mitigate the risks of reverse software engineering by integrating different obfuscation techniques as a whole solution. In the paper, we conduct a systematic review of existing obfuscation techniques based on the idea of layered obfuscation and develop a novel taxonomy of obfuscation techniques. Following our taxonomy hierarchy, the obfuscation strategies under different branches are orthogonal to each other. In this way, it can assist developers in choosing obfuscation techniques and designing layered obfuscation solutions based on their specific requirements. Hui Xu 0009, Yangfan Zhou 0002, Jiang Ming 0002, Michael R. Lyu |
Cybersecur. | 3 |
| 2019 | Capturing the Persistence of Facial Expression Features for Deepfake Video Detection
Yiru Zhao, Wanfeng Ge, Run Wang 0001, Lei Zhao 0012, Jiang Ming 0002 |
ICICS | 6 |
| 2019 | "Jekyll and Hyde" is Risky: Shared-Everything Threat Mitigation in Dual-Instance AppsabstractRecent developed application-level virtualization brings a groundbreaking innovation to Android ecosystem: a host app is able to load and launch arbitrary guest APK files without the hassle of installation. Powered by this technology, the so-called "dual-instance apps" are becoming increasingly popular as they can run dual copies of the same app on a single device (e.g., login Facebook simultaneously with two different accounts). Given the large demand from smartphone users, it is imperative to understand how secure dual-instance apps are. However, little work investigates their potential security risks. Even worse, new Android malware variants have been accused of skimming the cream off application-level virtualization. They abuse legitimate virtualization engines to launch phishing attacks or even thwart static detection. We first demonstrate that, current dual-instance apps design introduces serious "shared-everything" threats to users, and severe attacks such as permission escalation and privacy leak have become tremendously easier. Unfortunately, we find that most critical apps cannot discriminate between host app and Android system. In addition, traditional fingerprinting features targeting Android sandboxes are futile as well. To inform users that an app is running in an untrusted environment, we study the inherent features of dual-instance app environment and propose six robust fingerprinting features to detect whether an app is being launched by the host app. We test our approach, called DiPrint, with a set of dual-instance apps collected from popular app stores, Android systems, and virtualization-based malware. Our evaluation shows that DiPrint is able to accurately identify dual-instance apps with negligible overhead. Luman Shi, Jianming Fu, Zhengwei Guo, Jiang Ming 0002 |
MobiSys | 4 |
| 2018 | StateDroid: Stateful Detection of Stealthy Attacks in Android Apps via Horn-Clause VerificationabstractProfit-driven cyber-criminals are motivated to prolong Android malware's lifetime by hiding malicious behaviors from raising suspicion. Stealthy malware has become an emerging challenge to Android security as it can remain undetected for quite a long time. However, traditional defense techniques are insufficient in face of this new threat. Our in-depth study on published malware analysis reports and corresponding code analysis leads to three key observations: 1) a stealthy attack goes through multiple states; 2) state transitions are caused by a sequence of attack actions; 3) an attack action typically involves several Android APIs on different objects. Mohsin Junaid, Jiang Ming 0002, David Chenho Kung |
ACSAC | 2 |
| 2018 | Towards Paving the Way for Large-Scale Windows Malware Analysis: Generic Binary Unpacking with Orders-of-Magnitude Performance BoostabstractBinary packing, encoding binary code prior to execution and decoding them at run time, is the most common obfuscation adopted by malware authors to camouflage malicious code. Especially, most packers recover the original code by going through a set of "written-then-executed" layers, which renders determining the end of the unpacking increasingly difficult. Many generic binary unpacking approaches have been proposed to extract packed binaries without the prior knowledge of packers. However, the high runtime overhead and lack of anti-analysis resistance have severely limited their adoptions. Over the past two decades, packed malware is always a veritable challenge to anti-malware landscape. This paper revisits the long-standing binary unpacking problem from a new angle: packers consistently obfuscate the standard use of API calls. Our in-depth study on an enormous variety of Windows malware packers at present leads to a common property: malware's Import Address Table (IAT), which acts as a lookup table for dynamically linked API calls, is typically erased by packers for further obfuscation; and then unpacking routine, like a custom dynamic loader, will reconstruct IAT before original code resumes execution. During a packed malware execution, if an API is invoked through looking up a rebuilt IAT, it indicates that the original payload has been restored. This insight motivates us to design an efficient unpacking approach, called BinUnpack. Compared to the previous methods that suffer from multiple "written-then-executed" unpacking layers, BinUnpack is free from tedious memory access monitoring, and therefore it introduces very small runtime overhead. To defeat a variety of ever-evolving evasion tricks, we design BinUnpack's API monitor module via a novel kernel-level DLL hijacking technique. We have evaluated BinUnpack's efficacy extensively with more than 238K packed malware and multiple Windows utilities. BinUnpack's success rate is significantly better than that of existing tools with several orders of magnitude performance boost. Our study demonstrates that BinUnpack can be applied to speeding up large-scale malware analysis. Binlin Cheng, Jiang Ming 0002, Jianming Fu, Guojun Peng, Ting Chen 0002, Xiaosong Zhang 0001, Jean-Yves Marion |
CCS | 2 |
| 2018 | VMHunt: A Verifiable Approach to Partially-Virtualized Binary Code SimplificationabstractCode virtualization is a highly sophisticated obfuscation technique adopted by malware authors to stay under the radar. However, the increasing complexity of code virtualization also becomes a "double-edged sword" for practical application. Due to its performance limitations and compatibility problems, code virtualization is seldom used on an entire program. Rather, it is mainly used only to safeguard the key parts of code such as security checks and encryption keys. Many techniques have been proposed to reverse engineer the virtualized code, but they share some common limitations. They assume the scope of virtualized code is known in advance and mainly focus on the classic structure of code emulator. Also, few work verifies the correctness of their deobfuscation results. In this paper, with fewer assumptions on the type and scope of code virtualization, we present a verifiable method to address the challenge of partially-virtualized binary code simplification. Our key insight is that code virtualization is a kind of process-level virtual machine (VM), and the context switch patterns when entering and exiting the VM can be used to detect the VM boundaries. Based on the scope of VM boundary, we simplify the virtualized code. We first ignore all the instructions in a given virtualized snippet that do not affect the final result of that snippet. To better revert the data obfuscation effect that encodes a variable through bitwise operations, we then run a new symbolic execution called multiple granularity symbolic execution to further simplify the trace snippet. The generated concise symbolic formulas facilitate the correctness testing of our simplification results. We have implemented our idea as an open source tool, VMHunt, and evaluated it with real-world applications and malware. The encouraging experimental results demonstrate that VMHunt is a significant improvement over the state of the art. Dongpeng Xu 0001, Jiang Ming 0002, Dinghao Wu |
CCS | 2 |
| 2018 | Towards Predicting Efficient and Anonymous Tor Circuits
Armon Barton, Matthew Wright 0001, Jiang Ming 0002, Mohsen Imani |
USENIX Security Symposium | 3 |
| 2018 | Resetting Your Password Is Vulnerable: A Security Study of Common SMS-Based Authentication in IoT DeviceabstractFirmware vulnerability is an important target for IoT attacks, but it is challenging, because firmware may be publicly unavailable or encrypted with an unknown key. We present in this paper an attack on Short Message Service (SMS for short) authentication code which aims at gaining the control of IoT devices without firmware analysis. The key idea is based on the observation that IoT device usually has an official application (app for short) used to control itself. Customer needs to register an account before using this app, phone numbers are usually suggested to be the account name, and most of these apps have a common feature, calledReset Your Password, that uses an SMS authentication code sent to customer phone to authenticate the customer when he forgot his password. We found that an attacker can perform brute‐force attack on this SMS authentication code automatically by overcoming several challenges, then he can steal the account to gain the control of IoT devices. In our research, we have implemented a prototype tool, calledSACIntruder, to enable performing such brute‐force attack test on IoT devices automatically. We evaluated it and successfully found 12 zero‐day vulnerabilities including smart lock, sharing car, smart watch, smart router, etc. We also discussed how to prevent this attack. Dong Wang 0018, Xiaosong Zhang 0001, Jiang Ming 0002, Ting Chen 0002, Chao Wang 0021, Weina Niu |
Wirel. Commun. Mob. Comput. | 3 |
| 2017 | Cryptographic Function Detection in Obfuscated Binaries via Bit-Precise Symbolic Loop MappingabstractCryptographic functions have been commonly abused by malware developers to hide malicious behaviors, disguise destructive payloads, and bypass network-based firewalls. Now-infamous crypto-ransomware even encrypts victim's computer documents until a ransom is paid. Therefore, detecting cryptographic functions in binary code is an appealing approach to complement existing malware defense and forensics. However, pervasive control and data obfuscation schemes make cryptographic function identification a challenging work. Existing detection methods are either brittle to work on obfuscated binaries or ad hoc in that they can only identify specific cryptographic functions. In this paper, we propose a novel technique called bit-precise symbolic loop mapping to identify cryptographic functions in obfuscated binary code. Our trace-based approach captures the semantics of possible cryptographic algorithms with bit-precise symbolic execution in a loop. Then we perform guided fuzzing to efficiently match boolean formulas with known reference implementations. We have developed a prototype called CryptoHunt and evaluated it with a set of obfuscated synthetic examples, well-known cryptographic libraries, and malware. Compared with the existing tools, CryptoHunt is a general approach to detecting commonly used cryptographic functions such as TEA, AES, RC4, MD5, and RSA under different control and data obfuscation scheme combinations. Dongpeng Xu 0001, Jiang Ming 0002, Dinghao Wu |
IEEE Symposium on Security and Privacy | 2 |
| 2017 | BinSim: Trace-based Semantic Binary Diffing via System Call Sliced Segment Equivalence Checking
Jiang Ming 0002, Dongpeng Xu 0001, Yufei Jiang, Dinghao Wu |
USENIX Security Symposium | 1 |
| 2017 | Semantics-Based Obfuscation-Resilient Binary Code Similarity Comparison with Applications to Software and Algorithm Plagiarism DetectionabstractExisting code similarity comparison methods, whether source or binary code based, are mostly not resilient to obfuscations. Identifying similar or identical code fragments among programs is very important in some applications. For example, one application is to detect illegal code reuse. In the code theft cases, emerging obfuscation techniques have made automated detection increasingly difficult. Another application is to identify cryptographic algorithms which are widely employed by modern malware to circumvent detection, hide network communications, and protect payloads among other purposes. Due to diverse coding styles and high programming flexibility, different implementation of the same algorithm may appear very distinct, causing automatic detection to be very hard, let alone code obfuscations are sometimes applied. In this paper, we propose a binary-oriented, obfuscation-resilient binary code similarity comparison method based on a new concept, longest common subsequence of semantically equivalent basic blocks , which combines rigorous program semantics with longest common subsequence based fuzzy matching. We model the semantics of a basic block by a set of symbolic formulas representing the input-output relations of the block. This way, the semantic equivalence (and similarity) of two blocks can be checked by a theorem prover. We then model the semantic similarity of two paths using the longest common subsequence with basic blocks as elements. This novel combination has resulted in strong resiliency to code obfuscation. We have developed a prototype. The experimental results show that our method can be applied to software plagiarism and algorithm detection, and is effective and practical to analyze real-world software. Lannan Luo, Jiang Ming 0002, Dinghao Wu, Peng Liu 0005, Sencun Zhu |
IEEE Trans. Software Eng. | 2 |
| 2016 | Program-object Level Data Flow Analysis with Applications to Data Leakage and Contamination Forensics
Gaoyao Xiao, Jun Wang 0141, Peng Liu 0005, Jiang Ming 0002, Dinghao Wu |
CODASPY | 4 |
| 2016 | Translingual ObfuscationabstractProgram obfuscation is an important software protection technique that prevents attackers from revealing the programming logic and design of the software. We introduce translingual obfuscation, a new software obfuscation scheme which makes programs obscure by "misusing" the unique features of certain programming languages. Translingual obfuscation translates part of a program from its original language to another language which has a different programming paradigm and execution model, thus increasing program complexity and impeding reverse engineering. In this paper, we investigate the feasibility and effectiveness of translingual obfuscation with Prolog, a logic programming language. We implement translingual obfuscation in a tool called BABEL, which can selectively translate C functions into Prolog predicates. By leveraging two important features of the Prolog language, i.e., unification and backtracking, BABEL obfuscates both the data layout and control flow of C programs, making them much more difficult to reverse engineer. Our experiments show that BABEL provides effective and stealthy software obfuscation, while the cost is only modest compared to one of the most popular commercial obfuscators on the market. With BABEL, we verified the feasibility of translingual obfuscation, which we consider to be a promising new direction for software obfuscation. Pei Wang 0007, Shuai Wang 0011, Jiang Ming 0002, Yufei Jiang, Dinghao Wu |
EuroS&P | 3 |
| 2016 | Generalized Dynamic Opaque Predicates: A New Control Flow Obfuscation Method
Dongpeng Xu 0001, Jiang Ming 0002, Dinghao Wu |
ISC | 2 |
| 2016 | StraightTaint: decoupled offline symbolic taint analysisabstractTaint analysis has been widely applied in ex post facto security applications, such as attack provenance investigation, computer forensic analysis, and reverse engineering. Unfortunately, the high runtime overhead imposed by dynamic taint analysis makes it impractical in many scenarios. The key obstacle is the strict coupling of program execution and taint tracking logic code. To alleviate this performance bottleneck, recent work seeks to offload taint analysis from program execution and run it on a spare core or a different CPU. However, since the taint analysis has heavy data and control dependencies on the program execution, the massive data in recording and transformation overshadow the benefit of decoupling. In this paper, we propose a novel technique to allow very lightweight logging, resulting in much lower execution slowdown, while still permitting us to perform full-featured offline taint analysis. We develop StraightTaint, a hybrid taint analysis tool that completely decouples the program execution and taint analysis. StraightTaint relies on very lightweight logging of the execution information to reconstruct a straight-line code, enabling an offline symbolic taint analysis without frequent data communication with the application. While StraightTaint does not log complete runtime or input values, it is able to precisely identify the causal relationships between sources and sinks, for example. Compared with traditional dynamic taint analysis tools, StraightTaint has much lower application runtime overhead. Jiang Ming 0002, Dinghao Wu, Jun Wang 0141, Gaoyao Xiao, Peng Liu 0005 |
ASE | 1 |
| 2016 | BinCFP: Efficient Multi-threaded Binary Code Control Flow ProfilingabstractIn many tasks of reverse engineering and binary code analysis (e.g., hybrid disassembly, resolving indirect jump, and decoupled taint analysis), the knowledge of detailed dynamic control flow can be of great value. However, the high runtime overhead beset the complete collection of dynamic control flow. The previous efforts on efficient path profiling cannot be directly applied to the obfuscated binary code in which an accurate control flow graph is typically absent. To address these challenges, we present BinCFP, an efficient multi-threaded binary code control flow profiling tool by taking advantage of pervasive multi-core platforms. BinCFP relies on dynamic binary instrumentation to work with the unmodified binary code. The key of BinCFP is a multi-threaded fast buffering scheme that supports processing trace buffers asynchronously. To achieve better performance gains, we also apply a set of optimizations to reduce control flow profile size and instrumentation overhead. Our design enables the complete dynamic control flow collection for an obfuscated binary execution. We have implemented BinCFP on top of Pin. The comparative experiments on SPEC2006 and obfuscated common utility programs show BinCFP outperforms the previous work in several ways. In addition, BinCFP's control flow profile sizes are only about 49.2% that of the conventional design. Jiang Ming 0002, Dinghao Wu |
SCAM | 1 |
| 2016 | Deviation-Based Obfuscation-Resilient Program Equivalence Checking With Application to Software Plagiarism DetectionabstractSoftware plagiarism, an act of illegally copying others' code, has become a serious concern for honest software companies and the open source community. Considerable research efforts have been dedicated to searching the evidence of software plagiarism. In this paper, we continue this line of research and propose LoPD, a deviation-based program equivalence checking approach, which is an ideal fit for the whole-program plagiarism detection. Instead of directly comparing the similarity between two programs, LoPD searches for any dissimilarity between two programs by finding an input that will cause these two programs to behave differently, either with different output states or with semantically different execution paths. As long as we can find one dissimilarity, the programs are semantically different; but if we cannot find any dissimilarity, it is more likely a plagiarism case. We leverage dynamic symbolic execution to capture the semantics of execution paths and to find path deviations. Compared to the existing detection approaches, LoPD's formal program semantics-based method is more resilient to automatic obfuscation schemes. Our evaluation results indicate that LoPD is effective in detecting whole-program plagiarism. Furthermore, we demonstrate that LoPD can be applied to partial software plagiarism detection as well. The encouraging experiment results show that LoPD is an appealing complement to existing software plagiarism detection approaches. Jiang Ming 0002, Fangfang Zhang 0005, Dinghao Wu, Peng Liu 0005, Sencun Zhu |
IEEE Trans. Reliab. | 1 |
| 2015 | Replacement Attacks: Automatically Impeding Behavior-Based Malware Specifications
Jiang Ming 0002, Zhi Xin, Pengwei Lan, Dinghao Wu, Peng Liu 0005, Bing Mao 0001 |
ACNS | 1 |
| 2015 | LOOP: Logic-Oriented Opaque Predicate Detection in Obfuscated Binary CodeabstractOpaque predicates have been widely used to insert superfluous branches for control flow obfuscation. Opaque predicates can be seamlessly applied together with other obfuscation methods such as junk code to turn reverse engineering attempts into arduous work. Previous efforts in detecting opaque predicates are far from mature. They are either ad hoc, designed for a specific problem, or have a considerably high error rate. This paper introduces LOOP, a Logic Oriented Opaque Predicate detection tool for obfuscated binary code. Being different from previous work, we do not rely on any heuristics; instead we construct general logical formulas, which represent the intrinsic characteristics of opaque predicates, by symbolic execution along a trace. We then solve these formulas with a constraint solver. The result accurately answers whether the predicate under examination is opaque or not. In addition, LOOP is obfuscation resilient and able to detect previously unknown opaque predicates. We have developed a prototype of LOOP and evaluated it with a range of common utilities and obfuscated malicious programs. Our experimental results demonstrate the efficacy and generality of LOOP. By integrating LOOP with code normalization for matching metamorphic malware variants, we show that LOOP is an appealing complement to existing malware defenses. Jiang Ming 0002, Dongpeng Xu 0001, Dinghao Wu |
CCS | 1 |
| 2015 | Memoized Semantics-Based Binary Diffing with Application to Malware Lineage Inference
Jiang Ming 0002, Dongpeng Xu 0001, Dinghao Wu |
SEC | 1 |
| 2015 | TaintPipe: Pipelined Symbolic Taint Analysis
Jiang Ming 0002, Dinghao Wu, Gaoyao Xiao, Jun Wang 0141, Peng Liu 0005 |
USENIX Security Symposium | 1 |
| 2014 | Semantics-based obfuscation-resilient binary code similarity comparison with applications to software plagiarism detectionabstractExisting code similarity comparison methods, whether source or binary code based, are mostly not resilient to obfuscations. In the case of software plagiarism, emerging obfuscation techniques have made automated detection increasingly difficult. In this paper, we propose a binary-oriented, obfuscation-resilient method based on a new concept, longest common subsequence of semantically equivalent basic blocks, which combines rigorous program semantics with longest common subsequence based fuzzy matching. We model the semantics of a basic block by a set of symbolic formulas representing the input-output relations of the block. This way, the semantics equivalence (and similarity) of two blocks can be checked by a theorem prover. We then model the semantics similarity of two paths using the longest common subsequence with basic blocks as elements. This novel combination has resulted in strong resiliency to code obfuscation. We have developed a prototype and our experimental results show that our method is effective and practical when applied to real-world software. Lannan Luo, Jiang Ming 0002, Dinghao Wu, Peng Liu 0005, Sencun Zhu |
SIGSOFT FSE | 2 |
| 2011 | Linear Obfuscation to Combat Symbolic Execution
Zhi Wang 0014, Jiang Ming 0002, Chunfu Jia, Debin Gao |
ESORICS | 2 |
| 2011 | Towards ground truthing observations in gray-box anomaly detectionabstractAnomaly detection has been attracting interests from researchers due to its advantage of being able to detect zero-day exploits. A gray-box anomaly detector first observes benign executions of a computer program and then extracts reliable rules that govern the normal execution of the program. However, such observations from benign executions are not necessarily true evidences supporting the rules learned. For example, the observation that a file descriptor being equal to a socket descriptor should not be considered supporting a rule governing the two values to be the same. Ground truthing such observations is a difficult problem since it is not practical to analyze the semantics of every instruction in every program to be protected. In this paper, we propose using taint analysis to automatically help the ground truthing. Intuitively, the same taint source of two values provides ground truth of the data dependence. We implement a host-based anomaly detector with our proposed taint tracking and evaluate the accuracy of rules learned. Results show that we not only manage to filter out incorrect rules that would otherwise be learned (with high support and confidence), but manage recover good rules that are previously believed to be unreliable. We also present overheads of our system and time needed for training. Jiang Ming 0002, Debin Gao |
NSS | 1 |
| 2009 | Denial-of-Service Attacks on Host-Based Generic Unpackers
Jiang Ming 0002, Zhi Wang 0014, Debin Gao, Chunfu Jia |
ICICS | 2 |