EDBT 2026 Demo / reviewers in the wild / expert
Zhenkai Liang
dblp:99/4951
· DBLP profile ↗
110ranked-venue papers
5as first author
38since 2021 · last 2026
0000-0001-7138-5030ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 69 · 4 first-author · 22 since 2021Software engineering, systems software and programming languages · 25 · 12 since 2021Systems, architecture and hardware · 8 · 1 first-author · 1 since 2021Computer networks · 7 · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SymFlow: Event-Chain-Aware Symbolic Execution for Serverless Sensitive Data Flow DetectionabstractServerless applications are widely adopted for their scalability, cost-efficiency, and elastic resource management. However, their event-driven nature introduces complex event chains whose trigger-handler relationships are often determined dynamically by conditional logic, asynchronous callbacks, and resource-state dependencies. Existing security analysis tools, such as CloudFlow, mainly rely on static analysis, making it difficult to capture these dynamic event-chain interactions and the semantics of coarse-grained cloud APIs. As a result, they often fail to bridge the gap between architectural reachability and semantic feasibility, leading to both false positives and false negatives. Yuanpeng Wang, Zhineng Zhong, Zhenkai Liang, Ding Li 0001, Yao Guo 0001, Xiangqun Chen |
LCTES | 3 |
| 2026 | Promoguardian: Detecting Promotion Abuse Fraud with Multi-Relation Fused Graph Neural NetworksabstractAs e-commerce platforms develop, fraudulent activities are increasingly emerging, posing significant threats to the security and stability of these platforms. Promotion abuse is one of the fastest-growing types of fraud in recent years and is characterized by users exploiting promotional activities to gain financial benefits from the platform. To investigate this issue, we conduct the first study on promotion abuse fraud in e-commerce platforms MEITUAN. We find that promotion abuse fraud is a group-based fraudulent activity with two types of fraudulent activities: Stocking Up and Cashback Abuse. Unlike traditional fraudulent activities such as fake reviews, promotion abuse fraud typically involves ordinary customers conducting legitimate transactions and these two types of fraudulent activities are often intertwined. To address this issue, we propose leveraging additional information from the spatial and temporal perspectives to detect promotion abuse fraud. In this paper, we introduce PROMOGUARDIAN, a novel multi-relation fused graph neural network that integrates the spatial and temporal information of transaction data into a homogeneous graph to detect promotion abuse fraud. We conduct extensive experiments on real-world data from MEITUAN, and the results demonstrate that our proposed model outperforms state-of-the-art methods in promotion abuse fraud detection, achieving 93.15% precision, detecting 2.1 to 5.0 times more fraudsters, and preventing 1.5 to 8.8 times more financial losses in production environments. Shaofei Li, Ziqi Zhang 0017, Minyao Hua, Shuli Gao, Zhenkai Liang, Yao Guo 0001, Xiangqun Chen, Ding Li 0001 |
SP | 6 |
| 2026 | Fine-Grained Kernel Auditing Using Augmented Syscall Reference Behavior Analysis and Virtualized Selective Tracing
Chuqi Zhang, Spencer Faith, Feras Al-Qassas, Theodorus Februanto, Zhenkai Liang, Adil Ahmad |
SP | 5 |
| 2025 | Evaluating Disassembly Errors With Only Binaries
Lambang Akbar Wijayadi, Yuancheng Jiang, Roland H. C. Yap, Zhenkai Liang, Zhuohao Liu |
AsiaCCS | 4 |
| 2025 | Your Scale Factors are My Weapon: Targeted Bit-Flip Attacks on Vision Transformers via Scale Factor ManipulationabstractVision Transformers (ViTs) have experienced significant progress and are quantized for deployment in resource-constrained applications. Quantized models are vulnerable to targeted bit-flip attacks (BFAs). A targeted BFA prepares a trigger and a corresponding Trojan/backdoor, inserting the latter (with RowHammer bit flipping) into a victim model, to mislead its classification on samples containing the trigger. Existing targeted BFAs on quantized ViTs are limited in that: (1) they require numerous bit-flips, and (2) the separation between flipped bits is below 4 KB, making attacks infeasible with RowHammer in real-world scenarios. We propose a new and practical targeted attack Flip-S against quantized ViTs. The core insight is that in quantized models, a scale factor change ripples through a batch of model weights. Consequently, flipping bits in scale factors, rather than solely in model weights, enables more cost-effective attacks. We design a Scale-Factor-Search (SFS) algorithm to identify critical bits in scale factors for flipping, and adopt a mutual exclusion strategy to guarantee a 4 KB separation between flips. We evaluate Flip-S on CIFAR-10 and ImageNet datasets across five ViT architectures and two quantization levels. Results show that Flip-S achieves attack success rate (ASR) exceeding 90.0% on all models with 50 bits flipped, outperforming baselines with ASR typically below 80.0%. Furthermore, compared to the SOTA, Flip-S reduces the number of required bit-flips by 8×-20× while reaching equal or higher ASR. Our source code is publicly available1. Jialai Wang, Yuxiao Wu, Chao Zhang 0008, Zongpeng Li, Zhenkai Liang |
CVPR | 8 |
| 2025 | Erebor: A Drop-In Sandbox Solution for Private Data Processing in Untrusted Confidential Virtual MachinesabstractConfidential virtual machines (CVMs) are designed to protect data in cloud machines, but they fail in this task in common Software-as-a-Service (SaaS) cloud environments. In such settings, the software stack within a CVM, including service programs and the operating system, that receives and processes data may intentionally disclose it to attackers. We present Erebor, a sandboxing architecture for CVMs that processes client data in secure containers, where restrictions apply to both (a) access by all untrusted outside components and (b) the sandbox's ability to communicate data through memory and software-controlled direct or covert exits. Erebor enables such restrictions through a security monitor design based on intra-kernel privilege isolation for CVM, fully compatible with emerging cloud deployments without requiring host modifications. Under realistic scenarios, such as large language model inference and private information retrieval, Erebor only adds a performance overhead of 4.5%-13.2%, demonstrating its practicality in terms of enabling strong data sandboxing in modern cloud machines. Chuqi Zhang, Rahul Priolkar, Yuancheng Jiang, Yuan Xiao 0001, Mona Vij, Zhenkai Liang, Adil Ahmad |
EuroSys | 6 |
| 2025 | Fork State-Aware Differential Fuzzing for Blockchain Consensus ImplementationsabstractBlockchain networks allow multiple client implementations of the same consensus algorithm by different developers to coexist in the same system. Ensuring correct implementations among these heterogeneous clients is crucial, as even slight semantic discrepancies in their implementations can lead to safety failures. While existing fuzzing frameworks have discovered implementation flaws in blockchain, they suffer from several challenges in testing them with sequences of conflicting blocks, called forks. Existing tools fail to adequately assess the forkhandling processes in blockchain implementations when relying on traditional code coverage feedback, which lacks the granularity needed to navigate the diverse and complex fork-handling scenarios. This paper introduces FORKY, a fork state-aware differential fuzzing framework designed to detect implementation discrepancies within the critical fork-handling process with its novel fork-aware mutation and fork-diversifying feedback mechanisms. We test FORKY on the two most influential blockchain projects: Bitcoin and Ethereum, which are the representatives of the two major blockchain consensus algorithm families, Proof-of-Work (PoW) and Proof-of-Stake (PoS) consensus algorithms. Wonhoi Kim, Hocheol Nam 0001, Muoi Tran, Amin Jalilov, Zhenkai Liang, Sang Kil Cha, Min Suk Kang |
ICSE | 5 |
| 2025 | ZendDiff: Differential Testing of PHP InterpreterabstractThe PHP interpreter, powering over 70% of web-sites on the internet, plays a crucial role in web development. Existing approaches to finding bugs in PHP primarily focus on detecting explicit security issues through crashes or sanitizer-based oracles, but fail to identify logic bugs that can silently lead to incorrect results. We observe that the introduction of Just-In-Time (JIT) compilation mode in PHP presents an opportunity for differential testing, as it provides an alternative implementation of the same language specification. We propose, ZendDiff, an automatic differential testing framework that effectively detects logic bugs in the PHP interpreter by comparing JIT and non-JIT execution results. Our differential testing incorporates three techniques: program state probing for fine-grained execution state comparison, JIT-aware program mutation to sufficiently exercise JIT functionality, and dual verification to handle non-deterministic behaviors in PHP programs. Our experimental results demonstrate that ZendDiffoutperforms the official test suite used in PHP’s continuous integration, achieving higher code coverage and executing more Zend opcodes. Through ablation studies, we validate the effectiveness of these techniques. To date, ZendDiffhas identified 51 previously unknown logic bugs in the PHP interpreter, with 37 already fixed and 3 confirmed by the PHP maintainers. ZendDiffhas been acknowledged by the PHP community and offers a practical tool for automatically discovering logic bugs in the PHP interpreter. Yuancheng Jiang, Qiange Liu, Yeqi Fu, Roland H. C. Yap, Zhenkai Liang |
ASE | 7 |
| 2025 | Propagation-Based Vulnerability Impact Assessment for Software Supply ChainsabstractIdentifying the impact scope and scale is critical for software supply chain vulnerability assessment. However, existing studies face substantial limitations. First, prior studies either work at coarse package-level granularity—producing many false positives—or fail to accomplish whole-ecosystem vulnerability propagation analysis. Second, although vulnerability assessment indicators like CVSS characterize individual vulnerabilities, no metric exists to specifically quantify the dynamic impact of vulnerability propagation across software supply chains. To address these limitations and enable accurate and comprehensive vulnerability impact assessment, we propose a novel approach: (i) a hierarchical worklist-based algorithm for whole-ecosystem and call-graph-level vulnerability propagation analysis and (ii) the Vulnerability Propagation Scoring System (VPSS), a dynamic metric to quantify the scope and evolution of vulnerability impacts in software supply chains. We implement a prototype of our approach in the Java Maven ecosystem and evaluate it on 100 real-world vulnerabilities. Experimental results demonstrate that our approach enables effective ecosystem-wide vulnerability propagation analysis, and provides a practical, quantitative measure of vulnerability impact through VPSS. Bonan Ruan, Jiahao Liu 0005, Chuqi Zhang, Kaihang Ji, Zhenkai Liang |
ASE | 6 |
| 2025 | Improving LLM-based Log Parsing by Learning from Errors in Reasoning TracesabstractRecent advances in reasoning-capable large lan-guage models (LLMs) have led to their application in a wide range of tasks, including log parsing. These LLMs generate intermediate reasoning traces during inference, offering a unique opportunity to analyze and improve their performance. In this work, we investigate how reasoning traces can be leveraged to enhance LLM-based log parsers. We propose TraceDoctor, a framework that analyzes reasoning traces associated with parsing errors to understand the causes of failure. We categorize these error causes into high-level error types and design targeted log variant generation strategies guided by these high-level error types. The generated variants are then used to fine-tune the LLMs. We instantiate five state-of-the-art (SOTA) reasoning-capable LLMs as log parsers and identify 29 distinct high-level error types. Our approach improves their average parsing accuracy by up to 17.3% and 16.3% on parsing accuracy (PA) and group accuracy (GA), respectively. Jialai Wang, Juncheng Lu, Junjie Wang 0001, Chao Zhang 0008, Zhenkai Liang, Ee-Chien Chang |
ASE | 7 |
| 2025 | SCRUTINIZER: Towards Secure Forensics on Compromised TrustZone
Yiming Zhang 0030, Fengwei Zhang, Xiapu Luo, Rui Hou 0001, Xuhua Ding, Zhenkai Liang, Shoumeng Yan, Tao Wei 0002, Zhengyu He |
NDSS | 6 |
| 2025 | UI-CTX: Understanding UI Behaviors with Code Contexts for Mobile Applications
Jiahao Liu 0005, Jun Zeng 0006, Zhenkai Liang |
NDSS | 5 |
| 2025 | ProvGuard: Detecting SDN Control Policy Manipulation via Contextual Semantics of Provenance Graphs
Jun Zeng 0006, Qixiao Lin, Jiahao Liu 0005, Jianwei Zhuge, Zhenkai Liang |
NDSS | 8 |
| 2025 | RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguardsabstractLarge Language Models (LLMs) continue to exhibit vulnerabilities despite deliberate safety alignment efforts, posing significant risks to users and society. To safeguard against the risk of policy-violating content, system-level moderation via external guard models—designed to monitor LLM inputs and outputs and block potentially harmful content—has emerged as a prevalent mitigation strategy. Existing approaches of training guard models rely heavily on extensive human curated datasets and struggle with out-of-distribution threats, such as emerging harmful categories or jailbreak attacks. To address these limitations, we propose RSafe, an adaptive reasoning-based safeguard that conducts guided safety reasoning to provide robust protection within the scope of specified safety policies. RSafe operates
in two stages: (1) guided reasoning, where it analyzes safety risks of input content through policy-guided step-by-step reasoning, and (2) reinforced alignment, where rule-based RL optimizes its reasoning paths to align with accurate safety prediction. This two-stage training paradigm enables RSafe to internalize safety principles to generalize safety protection capability over unseen or adversarial safety violation
scenarios. During inference, RSafe accepts user-specified safety policies to provide enhanced safeguards tailored to specific safety requirements. Experiments demonstrate that RSafe matches state-of-the-art guard models using limited amount of public data in both prompt- and response-level harmfulness detection, while achieving superior out-of-distribution generalization on both emerging harmful category and jailbreak attacks. Furthermore, RSafe provides human-readable explanations for its safety judgments for better interpretability. RSafe offers a robust, adaptive, and interpretable solution for LLM safety moderation, advancing the development of reliable safeguards in dynamic real-world environments. Our code is available at https://anonymous.4open.science/r/RSafe-996D. Jingnan Zheng, Xiangtian Ji, Chenhang Cui, Weixiang Zhao, Gelei Deng, Zhenkai Liang, An Zhang 0003, Tat-Seng Chua |
NeurIPS | 7 |
| 2025 | TAPPecker: TAP Logic Inference and Violation Detection in Heterogeneous Smart Home SystemsabstractIn IoT environments-particularly within smart home systems-Trigger-Action Programming (TAP) serves as the primary mechanism for specifying automation rules. While TAP enables flexible and user-friendly automation, unintended interactions among TAP rules that violate security or privacy policies can lead to serious security consequences. A common assumption in prior research is that automation rules are readily available for analysis-that is, the TAP logic can be directly accessed. However, many real-world IoT platforms, such as Xiaomi and HomeKit, do not expose their internal automation rule sets. Moreover, existing logic extraction techniques primarily focus on the relationships between individual events, failing to capture the complex semantics of multi-condition TAP rules. The growing heterogeneity of smart home systems complicates the challenge of ensuring consistency between automation logic execution and system security objectives across such diverse “black-box” systems. In this paper, we present TAPPecker, an approach that leverages self-adaptation and evolutionary strategies to automatically infer TAP rules from system events in heterogeneous smart home environments. We analyze the inferred rules to detect potential security violations, with a specific focus on temporal aspects. We prototype TAPPecker and develop a hybrid testbed capable of generating realistic smart-home event logs, enabling comprehensive, multi-scenario testing. Our experimental results demonstrate that TAPPecker improves inference accuracy by $40.85 \%$ over existing approaches, while generating more expressive TAP logic and uncovering previously undetected security violations. Notably, our system revealed two time-related policy violations within official TAP rule sets that had not been previously reported. Qixiao Lin, Zhenkai Liang |
RAID | 4 |
| 2025 | Fuzzing the PHP Interpreter via Dataflow Fusion
Yuancheng Jiang, Chuqi Zhang, Bonan Ruan, Jiahao Liu 0005, Manuel Rigger, Roland H. C. Yap, Zhenkai Liang |
USENIX Security Symposium | 7 |
| 2024 | The HitchHiker's Guide to High-Assurance System Observability Protection with Efficient Permission SwitchesabstractProtecting system observability records (logs) from compromised OSs has gained significant traction in recent times, with several note-worthy approaches proposed. Unfortunately, none of the proposed approaches achieve high performance with tiny log protection delays. They also leverage risky environments for protection (e.g., many use general-purpose hypervisors or TrustZone, which have large TCB and attack surfaces). HitchHiker is an attempt to rectify this problem. The system is designed to ensure (a) in-memory protection of batched logs within a short and configurable real-time deadline by efficient hardware permission switching, and (b) an end-to-end high-assurance environment built upon hardware protection primitives with debloating strategies for secure log protection, persistence, and management. Security evaluations and validations show that HitchHiker reduces log protection delay by 93.3--99.3% compared to the state-of-the-art, while reducing TCB by 9.4--26.9X. Performance evaluations show HitchHiker incurs a geometric mean of less than 6% overhead on diverse real-world programs, improving on the state-of-the-art approach by 61.9--77.5%. Chuqi Zhang, Jun Zeng 0006, Yiming Zhang 0030, Adil Ahmad, Fengwei Zhang, Hai Jin 0001, Zhenkai Liang |
CCS | 7 |
| 2024 | Detecting Logic Bugs in Graph Database Management Systems via Injective and Surjective Graph Query TransformationabstractGraph Database Management Systems (GDBMSs) store graphs as data. They are used naturally in applications such as social networks, recommendation systems and program analysis. However, they can be affected by logic bugs, which cause the GDBMSs to compute incorrect results and subsequently affect the applications relying on them. In this work, we propose injective and surjective Graph Query Transformation (GQT) to detect logic bugs in GDBMSs. Given a query Q, we derive a mutated query Q', so that either their result sets are: (i) semantically equivalent; or (ii) variant based on the mutation to be either a subset or superset of each other. When the expected relationship between the results does not hold, a logic bug in the GDBMS is detected. The key insight to mutate Q is that the graph pattern in graph queries enables systematic query transformations derived from injective and surjective mappings of the directed edge sets between Q and Q'. We implemented injective and surjective Graph Query Transformation (GQT) as a tool called GraphGenie and evaluated it on 6 popular and mature GDBMSs. GraphGenie has found 25 unknown bugs, comprising 16 logic bugs, 3 internal errors, and 6 performance issues. Our results demonstrate the practicality and effectiveness of GraphGenie in detecting logic bugs in GDBMSs which has the potential for improving the reliability of applications relying on these GDBMSs. Yuancheng Jiang, Jiahao Liu 0005, Jinsheng Ba, Roland H. C. Yap, Zhenkai Liang, Manuel Rigger |
ICSE | 5 |
| 2024 | VulZoo: A Comprehensive Vulnerability Intelligence DatasetabstractSoftware vulnerabilities pose critical security and risk concerns. Many techniques are proposed to assess and prioritize vulnerabilities. To evaluate their performance, researchers often craft datasets from limited data sources, lacking a global overview of broad vulnerability intelligence. The repetitive data preparation process complicates the evaluation of new solutions. To solve this issue, we propose VulZoo, a comprehensive vulnerability intelligence dataset that covers 17 vulnerability data sources. We also construct connections among these sources, enabling more straightforward configuration and adaptation for different tasks. VulZoo provides utility scripts for automatic data synchronization and cleaning, relationship mining, and statistics generation. We make VulZoo publicly available and maintain it with incremental updates. We believe that VulZoo serves as a valuable input to vulnerability assessment and prioritization studies. The video is at https://youtu.be/EvoxQmUAHtw. The dataset is at https://github.com/NUS-Curiosity/VulZoo. Bonan Ruan, Jiahao Liu 0005, Weibo Zhao, Zhenkai Liang |
ASE | 4 |
| 2024 | MaskDroid: Robust Android Malware Detection with Masked Graph RepresentationsabstractAndroid malware attacks have posed a severe threat to mobile users, necessitating a significant demand for the automated detection system. Among the various tools employed in malware detection, graph representations (e.g., function call graphs) have played a pivotal role in characterizing the behaviors of Android apps. However, though achieving impressive performance in malware detection, current state-of-the-art graph-based malware detectors are vulnerable to adversarial examples. These adversarial examples are meticulously crafted by introducing specific perturbations to normal malicious inputs. To defend against adversarial attacks, existing defensive mechanisms are typically supplementary additions to detectors and exhibit significant limitations, often relying on prior knowledge of adversarial examples and failing to defend against unseen types of attacks effectively. In this paper, we propose MASKDROID, a powerful detector with a strong discriminative ability to identify malware and remarkable robustness against adversarial attacks. Specifically, we introduce a masking mechanism into the Graph Neural Network (GNN) based framework, forcing MASKDROID to recover the whole input graph using a small portion (e.g., 20%) of randomly selected nodes.This strategy enables the model to understand the malicious semantics and learn more stable representations, enhancing its robustness against adversarial attacks. While capturing stable malicious semantics in the form of dependencies inside the graph structures, we further employ a contrastive module to encourage MASKDROID to learn more compact representations for both the benign and malicious classes to boost its discriminative power in detecting malware from benign apps and adversarial examples. Jingnan Zheng, Jiahao Liu 0005, An Zhang 0003, Jun Zeng 0006, Zhenkai Liang, Tat-Seng Chua |
ASE | 6 |
| 2024 | KernJC: Automated Vulnerable Environment Generation for Linux Kernel VulnerabilitiesabstractLinux kernel vulnerability reproduction is a critical task in system security. To reproduce a kernel vulnerability, the vulnerable environment and the Proof of Concept (PoC) program are needed. Most existing research focuses on the generation of PoC, while the construction of environment is overlooked. However, establishing an effective vulnerable environment to trigger a vulnerability is challenging. Firstly, it is hard to guarantee that the selected kernel version for reproduction is vulnerable, as the vulnerability version claims in online databases can occasionally be incorrect. Secondly, many vulnerabilities cannot be reproduced in kernels built with default configurations. Intricate non-default kernel configurations must be set to include and trigger a kernel vulnerability, but less information is available on how to recognize these configurations. Bonan Ruan, Jiahao Liu 0005, Chuqi Zhang, Zhenkai Liang |
RAID | 4 |
| 2024 | CrypTody: Cryptographic Misuse Analysis of IoT Firmware via Data-flow ReasoningabstractCryptographic techniques form the foundation of the security and privacy of computing solutions. However, if cryptographic APIs are not invoked correctly, they can result in significant security problems. In this paper, we abstract the intricate crypto misuse detection problem as a data-flow reasoning task. Towards this end, we propose CrypTody, a novel logic-inference-based framework for detecting crypto misuses via reasoning about data flows on multi-architecture IoT firmware images. It carries out cross-architecture analysis, with detection strategies to reduce false positives and false negatives, such as cross-flow misuse inference. To evaluate the effectiveness of CrypTody, we conducted a large-scale experiment on 1,431 firmware images from 16 vendors. Our evaluation shows that 46% of the firmware images have high-risk misuses and 95% have at least one cryptographic misuse. In total, we find 6,624 potential crypto misuses, with 760 being cross-flow misuses that are not detected by existing solutions. We have responsibly disclosed portions of our findings to the relevant vendors. From the feedback, we note that CrypTody has a low false-positive rate for the confirmed misuses. Some typical cases have been assigned CVEs and fixed by the vendors. Shanqing Guo, Wenrui Diao, Hai-Xin Duan, Zhenkai Liang |
RAID | 7 |
| 2024 | UIHash: Detecting Similar Android UIs through Grid-Based Visual Appearance Representation
Jun Zeng 0006, Qixiao Lin, Shaowen Feng, Zhenkai Liang |
USENIX Security Symposium | 6 |
| 2023 | Learning Graph-based Code Representations for Source-level Functional Similarity DetectionabstractDetecting code functional similarity forms the basis of various software engineering tasks. However, the detection is challenging as functionally similar code fragments can be implemented differently, e.g., with irrelevant syntax. Recent studies incorporate program dependencies as semantics to identify syntactically different yet semantically similar programs, but they often focus only on local neighborhoods (e.g., one-hop dependencies), limiting the expressiveness of program semantics in modeling functionalities. In this paper, we present Tailor that explicitly exploits deep graph-structured code features for functional similarity detection. Given source-level programs, Tailor first represents them into code property graphs (CPGs) - which combine abstract syntax trees, control flow graphs, and data flow graphs - to collectively reason about program syntax and semantics. Then, Tailor learns representations of CPGs by applying a CPG-based neural network (CPGNN) to iteratively propagate information on them. It improves over prior work on code representation learning through a new graph neural network (GNN) tailored to CPG structures instead of the off-the-shelf GNNs used previously. We systematically evaluate Tailor on C and Java programs using two public benchmarks. Experimental results show that Tailor outperforms the state-of-the-art approaches, achieving 99.8% and 99.9% F-scores in code clone detection and 98.3% accuracy in source code classification. Jiahao Liu 0005, Jun Zeng 0006, Xiang Wang 0010, Zhenkai Liang |
ICSE | 4 |
| 2023 | Securing Web Inputs Using Parallel Session Attachments
Ruite Xu, Qixiao Lin, Shikun Wu, Zhenkai Liang |
SecureComm (2) | 6 |
| 2023 | I Know Your Social Network Accounts: A Novel Attack Architecture for Device-Identity AssociationabstractOnline social networks revolutionize the way people interact with each other. When various social network information aggregates over time, a rich online profile of the user is formed. Owing to the various features provided by mobile devices, a user's online social activities are tightly tied to his phone, and are conveniently, sometimes unnecessarily, available to social networks. In this article, we propose a novel attack architecture to show that attackers can infer a user's social network identities behind a mobile device through new dimensions. Specifically, we first developed a correlation between a user's device system states and the social network events, which leverage multiple mechanisms such as the learning-based memory regression model, to infer the possible accounts of the user in the social network app. Then we exploited the social network to social network correlation, via which we correlated information across different social networks, to identify the accounts of the target user. We implemented and evaluated these attacks on three popular social networks, and the results corroborate the effectiveness of our design. Yinhao Xiao, Xiuzhen Cheng, Shengling Wang 0001, Zhenkai Liang |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2022 | RecIPE: Revisiting the Evaluation of Memory Error DefensesabstractMany detection and defense mechanisms have been proposed to prevent and mitigate memory errors. A good understanding of the pros and cons of memory defense mechanisms is essential before practical deployment. However, the environment, compilers, and defenses change over time. Thus, a benchmark providing comprehensive evaluations of a range of compilation option choices and defenses is helpful to give developers guidance. The most well-known test suite for evaluating spatial memory errors and exploits is RIPE. However, we show that it is no longer applicable, and a new benchmark is needed. Furthermore, a benchmark should be extensible to evolve over time easily. We propose RecIPE, a new extensible and comprehensive benchmark for evaluating memory error defenses. For extensibility and customisability, RecIPE consists of two components. The TestGen component generates vulnerable code from configurable templates across various vulnerable attributes allowing easy changes through the templates. Measuring the attack and exploitation is with the DefEval component, serving as an attacker at runtime and can be updated with new exploitations. We present a comprehensive evaluation of compiler defenses and sanitizers on RecIPE with gcc and clang, showing the strengths and weaknesses of the various defenses. The results show pros and cons which may not be well known, including gaps between the in-principle guarantee and practical implementation. Our results also point out directions for further improvement in defenses. Yuancheng Jiang, Roland H. C. Yap, Zhenkai Liang, Hubert Rosier |
AsiaCCS | 3 |
| 2022 | PalanTír: Optimizing Attack Provenance with Hardware-enhanced System ObservabilityabstractSystem auditing is the foundation of attack provenance to investigate root causes and ramifications of cyber-attacks. However, provenance tracking on coarse-grained audit logs suffers from false causalities caused by dependency explosion. Recent approaches address this problem by increasing provenance granularity using execution partitioning or record-and-replay techniques. Unfortunately, they require program instrumentation and/or impose an unaffordable overhead, which is not practical in deployment. Jun Zeng 0006, Chuqi Zhang, Zhenkai Liang |
CCS | 3 |
| 2022 | Extensible Virtual Call Integrity
Yuancheng Jiang, Gregory J. Duck, Roland H. C. Yap, Zhenkai Liang, Pinghai Yuan |
ESORICS (3) | 4 |
| 2022 | AttacKG: Constructing Technique Knowledge Graph from Cyber Threat Intelligence Reports
Zhenyuan Li, Jun Zeng 0006, Yan Chen 0004, Zhenkai Liang |
ESORICS (1) | 4 |
| 2022 | TeLL: log level suggestions via modeling multi-level code block informationabstractDevelopers insert logging statements into source code to monitor system execution, which forms the basis for software debugging and maintenance. For distinguishing diverse runtime information, each software log is assigned with a separate verbosity level (e.g., trace and error). However, choosing an appropriate verbosity level is a challenging and error-prone task due to the lack of specifications for log level usages. Prior solutions aim to suggest log levels based on the code block in which a logging statement resides (i.e., intra-block features). Such suggestions, however, do not consider information from surrounding blocks (i.e., inter-block features), which also plays an important role in revealing logging characteristics. Jiahao Liu 0005, Jun Zeng 0006, Xiang Wang 0010, Kaihang Ji, Zhenkai Liang |
ISSTA | 5 |
| 2022 | SHADEWATCHER: Recommendation-guided Cyber Threat Analysis using System Audit RecordsabstractSystem auditing provides a low-level view into cyber threats by monitoring system entity interactions. In response to advanced cyber-attacks, one prevalent solution is to apply data provenance analysis on audit records to search for anomalies (anomalous behaviors) or specifications of known attacks. However, existing approaches suffer from several limitations: 1) generating high volumes of false alarms, 2) relying on expert knowledge, or 3) producing coarse-grained detection signals. In this paper, we recognize the structural similarity between threat detection in cybersecurity and recommendation in information retrieval. By mapping security concepts of system entity interactions to recommendation concepts of user-item interactions, we identify cyber threats by predicting the preferences of a system entity on its interactive entities. Furthermore, inspired by the recent advances in modeling high-order connectivity via item side information in the recommendation, we transfer the insight to cyber threat analysis and customize an automated detection system, SHADEWATCHER. It fulfills the potential of high-order information in audit records via graph neural networks to improve detection effectiveness. Besides, we equip SHADEWATCHER with dynamic updates towards better generalization to false alarms. In our evaluation against both real-life and simulated cyber-attack scenarios, SHADEWATCHER shows its advantage in identifying threats with high precision and recall rates. Moreover, SHADEWATCHER is capable of pinpointing threats from nearly a million system entity interactions within seconds. Jun Zeng 0006, Xiang Wang 0010, Jiahao Liu 0005, Yinfang Chen, Zhenkai Liang, Tat-Seng Chua, Zheng Leong Chua |
SP | 5 |
| 2022 | FreeWill: Automatically Diagnosing Use-after-free Bugs via Reference Miscounting Detection on Binaries
Liang He 0011, Hong Hu 0004, Purui Su, Yan Cai 0001, Zhenkai Liang |
USENIX Security Symposium | 5 |
| 2022 | FlowMatrix: GPU-Assisted Information-Flow Analysis through Matrix-Based Representation
Kaihang Ji, Jun Zeng 0006, Yuancheng Jiang, Zhenkai Liang, Zheng Leong Chua, Prateek Saxena, Abhik Roychoudhury |
USENIX Security Symposium | 4 |
| 2022 | Semantic-Fuzzing-Based Empirical Analysis of Voice Assistant Systems of Asian Symbol LanguagesabstractRecently, smart voice assistants (VAs) are widely deployed to provide control services via voice commands in IoT systems, e.g., smart home, industrial IoT systems, etc. However, due to the complexity of the application environment and the diversity of voice commands, more and more attacks against VAs cause severe security problems. As voice development platforms allow third-party voice skills to be accessed, adversaries are able to obtain users’ private information by squatting attacks using confusing names. The existing work studied the exploitability of semantic misinterpretation in VA systems on phonetic languages such as English. However, due to the semantic structural difference between phonetic English and symbol-based Asian languages, such as Chinese, the linguistic-model-guided fuzzing tool proposed by the previous work is insufficient to conduct semantic analysis on the VAs of Asian Languages. In this article, we conduct a systematic analysis to evaluate the feasibility of voice misinterpretation attacks to typical Asian language VAs through semantic fuzzing. We develop Harmony-Fuzzer, the semantic fuzzing tool that the fuzzing process is under the guidance of fuzzing rules abstracted from phenomena of speech errors, disfluency, or semantically similar expressions in Chinese corpus. We use Bayesian networks to formulate fuzzing models statistically so that the fuzzing space can be controlled by the probability of fuzzing processing. We use our results to test VAs and design malicious skills to empirically verify the feasibility of squatting attacks. We found that squatting attacks on Chinese VAs are feasible when attackers leverage some linguistic phenomena delicately. Qixiao Lin, Zhenkai Liang |
IEEE Internet Things J. | 4 |
| 2021 | Identifying privacy weaknesses from multi-party trigger-action integration platformsabstractWith many trigger-action platforms that integrate Internet of Things (IoT) systems and online services, rich functionalities transparently connecting digital and physical worlds become easily accessible for the end users. On the other hand, such facilities incorporate multiple parties whose data control policies may radically differ and even contradict each other, and thus privacy violations may arise throughout the lifecycle (e.g., generation and transmission) of triggers and actions. In this work, we conduct an in-depth study on the privacy issues in multi-party trigger-action integration platforms (TAIPs). We first characterize privacy violations that may arise with the integration of heterogeneous systems and services. Based on this knowledge, we propose Taifu, a dynamic testing approach to identify privacy weaknesses from the TAIP. The key insight of Taifu is that the applets which actually program the trigger-action rules can be used as test cases to explore the behavior of the TAIP. We evaluate the effectiveness of our approach by applying it on the TAIPs that are built around the IFTTT platform. To our great surprise, we find that privacy violations are prevalent among them. Using the automatically generated 407 applets, each from a different TAIP, Taifu detects 194 cases with access policy breaches, 218 access control missing, 90 access revocation missing, 15 unintended flows, and 73 over-privilege access. Kulani Mahadewa, Yanjun Zhang 0002, Guangdong Bai, Lei Bu, Zhiqiang Zuo 0002, Dileepa Fernando, Zhenkai Liang, Jin Song Dong 0001 |
ISSTA | 7 |
| 2021 | WATSON: Abstracting Behaviors from Audit Logs via Aggregation of Contextual Semantics
Jun Zeng 0006, Zheng Leong Chua, Yinfang Chen, Kaihang Ji, Zhenkai Liang |
NDSS | 5 |
| 2021 | Scrutinizing Implementations of Smart Home IntegrationsabstractA key feature of the booming smart home is the integration of a wide assortment of technologies, including various standards, proprietary communication protocols and heterogeneous platforms. Due to customization, unsatisfied assumptions and incompatibility in the integration, critical security vulnerabilities are likely to be introduced by the integration. Hence, this work addresses the security problems in smart home systems from anintegrationperspective, as a complement to numerous studies that focus on the analysis of individual techniques. We propose HomeScan, an approach that examines the security of the implementations of smart home systems. It extracts the abstract specification of application-layer protocols and internal behaviors of entities, so that it is able to conduct an end-to-end security analysis against various attack models. Applying HomeScanon three extensively-used smart home systems, we have found twelve non-trivial security issues, which may lead to unauthorized remote control and credential leakage. Kulani Mahadewa, Kailong Wang 0001, Guangdong Bai, Ling Shi 0002, Yan Liu 0012, Jin Song Dong 0001, Zhenkai Liang |
IEEE Trans. Software Eng. | 7 |
| 2020 | Robust P2P Primitives Using SGX EnclavesabstractPeer-to-peer (P2P) systems such as BitTorrent and Bitcoin are susceptible to serious attacks from byzantine nodes that join as peers. Due to well-known impossibility results for designing P2P primitives in unrestricted byzantine settings, research has explored many adversarial models with additional assumptions, ranging from mild (such as pre-established PKI) to strong (such as the existence of common random coins). One such widely-studied model is the general-omission model, which yields simple protocols with good efficiency, but has been considered impractical or unrealizable since it artificially limits the adversary only to omitting messages.In this work, we study the setting of a synchronous network wherein peer nodes have CPUs equipped with a recent trusted computing mechanism called Intel SGX. In this model, we observe that the byzantine adversary reduces to the adversary in the general-omission model. As a first result, we show that by leveraging SGX features, we eliminate any source of advantage for a byzantine adversary beyond that gained by omitting messages, making the general-omission model realizable. Our evaluation of 1000 nodes running on 40 DeterLab machines confirms theoretical efficiency claim. Yaoqi Jia, Shruti Tople, Tarik Moataz, Deli Gong, Prateek Saxena, Zhenkai Liang |
ICDCS | 6 |
| 2020 | Robust P2P Primitives Using SGX Enclaves
Yaoqi Jia, Shruti Tople, Tarik Moataz, Deli Gong, Prateek Saxena, Zhenkai Liang |
RAID | 6 |
| 2019 | Neural Network Inversion in Adversarial Setting via Background Knowledge AlignmentabstractThe wide application of deep learning technique has raised new security concerns about the training data and test data. In this work, we investigate the model inversion problem under adversarial settings, where the adversary aims at inferring information about the target model's training data and test data from the model's prediction values. We develop a solution to train a second neural network that acts as the inverse of the target model to perform the inversion. The inversion model can be trained with black-box accesses to the target model. We propose two main techniques towards training the inversion model in the adversarial settings. First, we leverage the adversary's background knowledge to compose an auxiliary set to train the inversion model, which does not require access to the original training data. Second, we design a truncation-based technique to align the inversion model to enable effective inversion of the target model from partial predictions that the adversary obtains on victim user's data. We systematically evaluate our approach in various machine learning tasks and model architectures on multiple image datasets. We also confirm our results on Amazon Rekognition, a commercial prediction API that offers "machine learning as a service". We show that even with partial knowledge about the black-box model's training data, and with only partial prediction values, our inversion approach is still able to perform accurate inversion of the target model, and outperform previous approaches. Jiyi Zhang, Ee-Chien Chang, Zhenkai Liang |
CCS | 4 |
| 2019 | LightSense: A Novel Side Channel for Zero-permission Mobile User Tracking
Quanqi Ye, Guangdong Bai, Naipeng Dong, Zhenkai Liang, Jin Song Dong 0001, Haoyu Wang 0001 |
ISC | 5 |
| 2019 | One Engine To Serve 'em All: Inferring Taint Rules Without Architectural Semantics
Zheng Leong Chua, Teodora Baluta, Prateek Saxena, Zhenkai Liang, Purui Su |
NDSS | 5 |
| 2019 | Detecting Android Side Channel Probing Attacks Based on System States
Qixiao Lin, Futian Shi, Shishi Zhu, Zhenkai Liang |
WASA | 5 |
| 2019 | Fuzzing Program Logic Deeply Hidden in Binary Program StagesabstractFuzzing is an effective method to identify bugs and security vulnerabilities in software. One particular difficulty faced by fuzzing is how to effectively generate inputs to cover program paths, especially for programs with complex logic. We observe that complex programs are often composed of components, which is a natural result of software engineering principles. The components interface with each other using memory buffers, forming stages of processing in the program logic. Program logic in later stages is difficult to reach by fuzzers. In this paper, we develop a novel solution to fuzz such program logic, called STAGEFUZZER. It identifies the stages and memory interfaces from program binaries, and fuzzes later stages of the program effectively. In our evaluation with a suite of typical binaries, STAGEFUZZER correctly identifies the program structure and effectively increases the coverage of program logic compared to AFL fuzzer. Zheng Leong Chua, Yuwei Liu 0001, Purui Su, Zhenkai Liang |
SANER | 5 |
| 2019 | I Can See Your Brain: Investigating Home-Use Electroencephalography System SecurityabstractHealth-related Internet of Things (IoT) devices are becoming more popular in recent years. On the one hand, users can access information of their health conditions more conveniently; on the other hand, they are exposed to new security risks. In this paper, we presented, to the best of our knowledge, the first in-depth security analysis on home-use electroencephalography (EEG) IoT devices. Our key contributions are twofold. First, we reverse-engineered the home-use EEG system framework via which we identified the design and implementation flaws. By exploiting these flaws, we developed two sets of novel easy-to-exploit PoC attacks, which consist of four remote attacks and one proximate attack. In a remote attack, an attacker can steal a user's brain wave data through a carefully crafted program while in the proximate attack, the attacker can steal a victim's brain wave data over-the-air without accessing the victim's device on any sense when he is close to the victim. As a result, all the 156 brain-computer interface (BCI) apps in the NeuroSky App store are vulnerable to the proximate attack. We also discovered that all the 31 free apps in the NeuroSky App store are vulnerable to at least one remote attack. Second, we proposed a novel deep learning model of a joint recurrent convolutional neural network (RCNN) to infer a user's activities based on the reduced-featured EEG data stolen from the home-use EEG IoT devices, and our evaluation over the real-world EEG data indicates that the inference accuracy of the proposed RCNN is can reach 70.55%. Yinhao Xiao, Xiuzhen Cheng, Jiguo Yu, Zhenkai Liang, Zhi Tian |
IEEE Internet Things J. | 5 |
| 2018 | DTaint: Detecting the Taint-Style Vulnerability in Embedded Device FirmwareabstractA rising number of embedded devices are reachable in the cyberspace, such as routers, cameras, printers, etc. Those devices usually run firmware whose code is proprietary with few public documents. Furthermore, most of the firmware images cannot be analyzed in dynamic analysis due to various hardware-specific peripherals. As a result, it hinders traditional static analysis and dynamic analysis techniques. In this paper, we propose a static binary analysis approach, DTaint, to detect taint-style vulnerabilities in the firmware. The taint-style vulnerability is a typical class of weakness, where the input data reaches a sensitive sink through an unsafe path. Specifically, we generate data dependency in a bottom-up manner through traversing callees before callers. To reduce the influence of the binary firmware, DTaint identifies pointer aliasing, interprocedural data flow, and similarity of the data structure layout. We have implemented a prototype of DTaint and conducted experiments to evaluate its performance. Our results show that DTaint discovers more vulnerabilities in less time, compared with the existing techniques. Furthermore, we illustrate the effectiveness of DTaint through applying it over six firmware images from four manufacturers. We have found 21 vulnerabilities, where 13 of them are previously-unknown and zero-day vulnerabilities. Qiang Li 0007, Yaowen Zheng, Limin Sun 0001, Zhenkai Liang |
DSN | 7 |
| 2018 | Robust Detection of Android UI SimilarityabstractThe similarity of the user interfaces (UIs) among Android apps is an important factor to indicate abnormal behaviors in Android apps. For example, phishing apps use UIs similar to those of their target apps to lure users to input sensitive information. As another example, repackaged apps preserve the UIs of the original apps while adding malicious code. In this paper, we propose a novel solution, GeminiScope, to robustly detect similar UIs among Android apps. Our approach analyzes the UI layouts of Android apps, extracts the fundamental features of the UIs' visual appearance, and rates the similarity among the UIs of apps. We evaluated GeminiScope using a few sets of apps from various sources. It reliably detected apps with similar UIs, showing the potential of using visual similarity as a unique feature to detect malicious apps. Jingdong Bian, Hanjun Ma, Yaoqi Jia, Zhenkai Liang, Xuxian Jiang |
ICC | 5 |
| 2018 | HOMESCAN: Scrutinizing Implementations of Smart Home IntegrationsabstractA key feature of the booming smart home is the integration of a wide assortment of technologies, including various standards, proprietary communication protocols and heterogeneous platforms. Due to customization, unsatisfied assumptions and incompatibility in the integration, critical security vulnerabilities are likely to be introduced by the integration. Hence, this work addresses the security problems in smart home systems from an integration perspective, as a complement to numerous studies that focus on the analysis of individual techniques. We propose HOMESCAN, an approach that examines the security of the implementations of smart home systems. It extracts the abstract specification of application-layer protocols and internal behaviors of participants, so that it is able to conduct an end-to-end security analysis against various attack models. Applying HOMESCAN on three extensively-used smart home systems, we have found twelve non-trivial security vulnerabilities, which may lead to unauthorized remote control and credential leakage. Kulani Mahadewa, Kailong Wang 0001, Guangdong Bai, Ling Shi 0002, Jin Song Dong 0001, Zhenkai Liang |
ICECCS | 6 |
| 2018 | A Novel Graph-based Mechanism for Identifying Traffic Vulnerabilities in Smart Home IoTabstractSmart home IoT devices have been more prevalent than ever before but the relevant security considerations fail to keep up with due to device and technology heterogeneity and resource constraints, making IoT systems susceptible to various attacks. In this paper, we propose a novel graph-based mechanism to identify the vulnerabilities in communication of IoT devices for smart home systems. Our approach takes one or more packet capture files as inputs to construct a traffic graph by passing the captured messages, identify the correlated subgraphs by examining the attribute-value pairs associated with each message, and then quantify their vulnerabilities based on the sensitivity levels of different keywords. To test the effectiveness of our approach, we setup a smart home system that can control a smart bulb LB100 via either the smartphone APP for LB100 or the Google Home speaker. We collected and analyzed 58,714 messages and exploited 6 vulnerable correlated sub graphs, based on which we implemented 6 attack cases that can be easily reproduced by attackers with little knowledge of IoT. This study is novel as our approach takes only the collected traffic files as inputs without requiring the knowledge of the device firmware while being able to identify new vulnerabilities. With this approach, we won the third prize out of 20 teams in a hacking competition. Yinhao Xiao, Jiguo Yu, Xiuzhen Cheng, Zhenkai Liang, Zhiguo Wan |
INFOCOM | 5 |
| 2018 | Automated Identification of Sensitive Data via Flexible User Requirements
Zhenkai Liang |
SecureComm (1) | 2 |
| 2018 | Automated identification of sensitive data from implicit user specificationabstractThe sensitivity of information is dependent on the context of application and user preference. Protecting sensitive data in the cloud era requires identifying them in the first place. It typically needs intensive manual efforts. More importantly, users may specify sensitive information only through an implicit manner. Existing research efforts on identifying sensitive data from its descriptive texts focus on keyword/phrase searching. These approaches can have high false positives/negatives as they do not consider the semantics of the descriptions. In this paper, we propose S3 , an automated approach to identify sensitive data based on users’ implicit specifications. Our approach considers semantic, syntactic and lexical information comprehensively, aiming to identify sensitive data by the semantics of its descriptive texts. We introduce the notion concept space to represent the user’s notion of privacy, by which our approach can support flexible user requirements in defining sensitive data. Our approach is able to learn users’ preferences from readable concepts initially provided by users, and automatically identify related sensitive data. We evaluate our approach on over 18,000 top popular applications from Google Play Store. S3 achieves an average precision of 89.2%, and average recall 95.8% in identifying sensitive data. Zhenkai Liang |
Cybersecur. | 2 |
| 2018 | SplitPass: A Mutually Distrusting Two-Party Password Manager
Dong Du 0003, Yubin Xia, Haibo Chen 0001, Binyu Zang, Zhenkai Liang |
J. Comput. Sci. Technol. | 6 |
| 2017 | Automatically assessing crashes from heap overflowsabstractHeap overflow is one of the most widely exploited vulnerabilities, with a large number of heap overflow instances reported every year. It is important to decide whether a crash caused by heap overflow can be turned into an exploit. Efficient and effective assessment of exploitability of crashes facilitates to identify severe vulnerabilities and thus prioritize resources. In this paper, we propose the first metrics to assess heap overflow crashes based on both the attack aspect and the feasibility aspect. We further present HCSIFTER, a novel solution to automatically assess the exploitability of heap overflow instances under our metrics. Given a heap-based crash, HCSIFTER accurately detects heap overflows through dynamic execution without any source code or debugging information. Then it uses several novel methods to extract program execution information needed to quantify the severity of the heap overflow using our metrics. We have implemented a prototype HCSIFTER and applied it to assess nine programs with heap overflow vulnerabilities. HCSIFTER successfully reports that five heap overflow vulnerabilities are highly exploitable and two overflow vulnerabilities are unlikely exploitable. It also gave quantitatively assessments for other two programs. On average, it only takes about two minutes to assess one heap overflow crash. The evaluation result demonstrates both effectiveness and efficiency of HC Sifter. Liang He 0011, Yan Cai 0001, Hong Hu 0004, Purui Su, Zhenkai Liang, Yi Yang 0040, Huafeng Huang, Jia Yan 0004, Xiangkun Jia, Dengguo Feng |
ASE | 5 |
| 2017 | Neural Nets Can Learn Function Type Signatures From Binaries
Zheng Leong Chua, Shiqi Shen, Prateek Saxena, Zhenkai Liang |
USENIX Security Symposium | 4 |
| 2017 | Phishing Website Detection Based on Effective CSS Features of Web Pages
Wenqian Tian, Tao Wei 0002, Zhenkai Liang |
WASA | 5 |
| 2017 | RoppDroid: Robust permission re-delegation prevention in Android inter-component communicationabstractAndroid is designed such that Android applications (Apps) can provide functions to each other by providing a complex inter-component communication (ICC) model. While app interactions make it convenient and easy for one app to delegate functionality to another app, it also leads to permission re-delegation among Android apps which can cause privilege escalation . One approach taken by existing work tries to mitigate privilege escalation by enforcing tightened permissions. Unfortunately, preventing privilege escalation often renders the recipient apps unusable (for example, causing the app to crash). In this work, we propose another approach to address the privilege escalation problem from Android app ICC which intends to better preserve app functionality. We propose a context specific resource virtualization to eliminate privilege escalation by taking into account the interaction of ICCs among apps. We evaluated our prototype system, R opp D roid , on real-world Android apps and showed the effectiveness in providing robust protection for those apps. Our prototype also has low performance overheads. Behnaz Hassanshahi, Roland H. C. Yap, Zhenkai Liang |
Comput. Secur. | 5 |
| 2017 | Monet: A User-Oriented Behavior-Based Malware Variants Detection System for AndroidabstractAndroid, the most popular mobile OS, has around 78% of the mobile market share. Due to its popularity, it attracts many malware attacks. In fact, people have discovered around 1 million new malware samples per quarter, and it was reported that over 98% of these new malware samples are in fact “derivatives” (or variants) from existing malware families. In this paper, we first show that runtime behaviors of malware's core functionalities are in fact similar within a malware family. Hence, we propose a framework to combine “runtime behavior” with “static structures” to detect malware variants. We present the design and implementation of Monet, which has a client and a backend server module. The client module is a lightweight, in-device app for behavior monitoring and signature generation, and we realize this using two novel interception techniques. The backend server is responsible for large scale malware detection. We collect 3723 malware samples and top 500 benign apps to carry out extensive experiments of detecting malware variants and defending against malware transformation. Our experiments show that Monet can achieve around 99% accuracy in detecting malware variants. Furthermore, it can defend against ten different obfuscation and transformation techniques, while only incurs around 7% performance overhead and about 3% battery overhead. More importantly, Monet will automatically alert users with intrusion details so to prevent further malicious behaviors. Mingshen Sun, John C. S. Lui, Richard T. B. Ma, Zhenkai Liang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2016 | "The Web/Local" Boundary Is Fuzzy: A Security Study of Chrome's Process-based SandboxingabstractProcess-based isolation, suggested by several research prototypes, is a cornerstone of modern browser security architectures. Google Chrome is the first commercial browser that adopts this architecture. Unlike several research prototypes, Chrome's process-based design does not isolate different web origins, but primarily promises to protect "the local system" from "the web". However, as billions of users now use web-based cloud services (e.g., Dropbox and Google Drive), which are integrated into the local system, the premise that browsers can effectively isolate the web from the local system has become questionable. In this paper, we argue that, if the process-based isolation disregards the same-origin policy as one of its goals, then its promise of maintaining the "web/local system (local)" separation is doubtful. Specifically, we show that existing memory vulnerabilities in Chrome's renderer can be used as a stepping-stone to drop executables/scripts in the local file system, install unwanted applications and misuse system sensors. These attacks are purely data-oriented and do not alter any control flow or import foreign code. Thus, such attacks bypass binary-level protection mechanisms, including ASLR and in-memory partitioning. Finally, we discuss various full defenses and present a possible way to mitigate the attacks presented. Yaoqi Jia, Zheng Leong Chua, Hong Hu 0004, Shuo Chen 0001, Prateek Saxena, Zhenkai Liang |
CCS | 6 |
| 2016 | Data-Oriented Programming: On the Expressiveness of Non-control Data AttacksabstractAs control-flow hijacking defenses gain adoption, it is important to understand the remaining capabilities of adversaries via memory exploits. Non-control data exploits are used to mount information leakage attacks or privilege escalation attacks program memory. Compared to control-flow hijacking attacks, such non-control data exploits have limited expressiveness, however, the question is: what is the real expressive power of non-control data attacks? In this paper we show that such attacks are Turing-complete. We present a systematic technique called data-oriented programming (DOP) to construct expressive non-control data exploits for arbitrary x86 programs. In the experimental evaluation using 9 programs, we identified 7518 data-oriented x86 gadgets and 5052 gadget dispatchers, which are the building blocks for DOP. 8 out of 9 real-world programs have gadgets to simulate arbitrary computations and 2 of them are confirmed to be able to build Turing-complete attacks. We build 3 end-to-end attacks to bypass randomization defenses without leaking addresses, to run a network bot which takes commands from the attacker, and to alter the memory permissions. All the attacks work in the presence of ASLR and DEP, demonstrating how the expressiveness offered by DOP significantly empowers the attacker. Hong Hu 0004, Shweta Shinde, Sendroiu Adrian, Zheng Leong Chua, Prateek Saxena, Zhenkai Liang |
IEEE Symposium on Security and Privacy | 6 |
| 2016 | Toward Exposing Timing-Based Probing Attacks in Web Applications
Futian Shi, Yaoqi Jia, Zhenkai Liang |
WASA | 5 |
| 2016 | Anonymity in Peer-assisted CDNs: Inference Attacks and MitigationabstractAbstract The peer-assisted CDN is a new content distribution paradigm supported by CDNs (e.g., Akamai), which enables clients to cache and distribute web content on behalf of a website. Peer-assisted CDNs bring significant bandwidth savings to website operators and reduce network latency for users. In this work, we show that the current designs of peer-assisted CDNs expose clients to privacy-invasive attacks, enabling one client to infer the set of browsed resources of another client. To alleviate this, we propose an anonymous peer-assisted CDN (APAC), which employs content delivery while providing initiator anonymity (i.e., hiding who sends the resource request) and responder anonymity (i.e., hiding who responds to the request) for peers. APAC can be a web service, compatible with current browsers and requiring no client-side changes. Our anonymity analysis shows that our APAC design can preserve a higher level of anonymity than state-of-the-art peer-assisted CDNs. In addition, our evaluation demonstrates that APAC can achieve desired performance gains. Yaoqi Jia, Guangdong Bai, Prateek Saxena, Zhenkai Liang |
Proc. Priv. Enhancing Technol. | 4 |
| 2016 | A Framework for Practical Dynamic Software UpdatingabstractDynamic software updating (DSU) enables a program to be patched on the fly without being shutdown. This paper addresses the practicality problem of the recent research on DSU systems, and presents Replus, a new DSU system that balances practicality and functionality. Replus aims to retain backward binary compatibility and support multi-threaded programs. In addition, it does not require customers to have developer-level software knowledge. More importantly, without specific compiler support, Replus can patch programs that are difficult to be updated at runtime, as well as programs that may incur an indefinite delay in DSU. The key technique of our solution is to update the stack elements for the patched program using two new mechanisms:Immediate Stack Updating, which immediately updates the stack of a thread, andtimely stack updating, which only updates the stack frames of the necessary functions without affecting others. Replus also develops anInstruction Level Updatingmechanism, which is more efficient for certain security patches. We used popular server applications as test suites to evaluate the effectiveness of Replus. The experimental results demonstrated that Replus can successfully update all the test suites with negligible impact on application performance. Hai Jin 0001, Deqing Zou, Zhenkai Liang, Bing Bing Zhou |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2015 | Web-to-Application Injection Attacks on Android: Characterization and Detection
Behnaz Hassanshahi, Yaoqi Jia, Roland H. C. Yap, Prateek Saxena, Zhenkai Liang |
ESORICS (2) | 5 |
| 2015 | Identifying Arbitrary Memory Access Vulnerabilities in Privilege-Separated Software
Hong Hu 0004, Zheng Leong Chua, Zhenkai Liang, Prateek Saxena |
ESORICS (2) | 3 |
| 2015 | Automatic Generation of Data-Oriented Exploits
Hong Hu 0004, Zheng Leong Chua, Sendroiu Adrian, Prateek Saxena, Zhenkai Liang |
USENIX Security Symposium | 5 |
| 2015 | Man-in-the-browser-cache: Persisting HTTPS attacks via browser cache poisoning
Yaoqi Jia, Xinshu Dong, Prateek Saxena, Zhenkai Liang |
Comput. Secur. | 6 |
| 2014 | CCS'14 Co-Located Workshop Summary for SPSM 2014abstractSecurity and privacy in smartphones and mobile devices is an emerging area which has received significant attention from the research community during the past few years. The SPSM workshop was created to bring together these researchers and practitioners. Following the success of the three previous editions, we present this fourth edition of the workshop which has attracted a significant number of great submissions and benefited from the expertise of an international program committee comprising of mobile security experts across the academia and the industry. Kapil Singh, Zhenkai Liang |
CCS | 2 |
| 2014 | TrustFound: Towards a Formal Foundation for Model Checking Trusted Computing Platforms
Guangdong Bai, Jianan Hao, Jianliang Wu 0002, Yang Liu 0003, Zhenkai Liang, Andrew P. Martin |
FM | 5 |
| 2014 | Understanding Complex Binary Loading BehaviorsabstractBinary loading is used extensively in many operating systems, e.g. Program execution usually involves loading dynamically linked libraries (binaries in DLL form). In Windows, binary loading is used heavily, but the process is complex and is affected by many factors - this flexibility turns out to be a rich source of attacks. When a typical Windows executable runs, many binaries are loaded, possibly from third parties. It is not uncommon for Windows programs to have binary loading vulnerabilities. However, it is difficult for software developers to identify if their programs have such vulnerabilities, how they arise, and how to fix them. We propose LDRSCOPE, to explain why binaries are loaded and detect the factors that affect the loading. This allows developers to better identify the problems and secure their code. We also deal with vulnerabilities arising from software configuration such as configuration files. Some vulnerabilities can also be due to third party libraries, we clearly identify and explain their effects. Roland H. C. Yap, Zhenkai Liang |
ICECCS | 4 |
| 2014 | DroidVault: A Trusted Data Vault for Android DevicesabstractMobile OSes and applications form a large, complex and vulnerability-prone software stack. In such an environment, security techniques to strongly protect sensitive data in mobile devices are important and challenging. To address such challenges, we introduce the concept of the trusted data vault, a small trusted engine that securely manages the storage and usage of sensitive data in an untrusted mobile device. In this paper, we design and build Droid Vault - the first realization of a trusted data vault on the Android platform. Droid Vault establishes a secure channel between data owners and data users while allowing data owners to enforce strong control over the sensitive data with a minimal trusted computing base (TCB). We prototype Droid Vault via the novel use of hardware security features of ARM processors, i.e., Trust Zone. Our evaluation demonstrates its functionality for processing sensitive data and its practicality for adoption in the real world. Hong Hu 0004, Guangdong Bai, Yaoqi Jia, Zhenkai Liang, Prateek Saxena |
ICECCS | 5 |
| 2014 | SQLR: Grammar-Guided Validation of SQL Injection SanitizersabstractThe SQL injection attack is one of the major threats to web applications. Through malicious inputs, attackers can cause data leakage and damage, and even remote code execution on the victim servers. A common solution is to use input sanitizers to filter out inputs that can result in SQL injection attacks. In this paper, we propose a novel solution, SQLR, to validate SQL sanitizers by systematically generating SQL injection attack patterns. Our approach uses the SQL grammar to guide the enumeration of malicious SQL queries efficiently, and summarizes the queries into patterns that can be used by existing solutions. SQLR successfully identified new attack patterns and weaknesses in sanitizers used in several real-world web applications. Sai Sathyanarayan, Dawei Qi, Zhenkai Liang, Abhik Roychoudary |
ICECCS | 3 |
| 2014 | AirBag: Boosting Smartphone Resistance to Malware Infection
Chiachih Wu, Yajin Zhou, Kunal Patel, Zhenkai Liang, Xuxian Jiang |
NDSS | 4 |
| 2014 | You Can't Be Me: Enabling Trusted Paths and User Sub-origins in Web Browsers
Enrico Budianto, Yaoqi Jia, Xinshu Dong, Prateek Saxena, Zhenkai Liang |
RAID | 5 |
| 2013 | CRYPTSERVER: strong data protection in commodity LAMP serversabstractModern web applications store sensitive data on their servers. Such data is prone to theft resulting from exploits against vulnerabilities in the server software stacks. In this work, we propose a new architecture for web servers, called CryptServer, in which we pre-determine and fix a small amount of application code that can compute over sensitive data. By encrypting sensitive data before making it available to the rest of untrusted application code, CryptServer provides strong defense against all malicious code that an attacker may run in the server software stack. As a step towards making this approach practical, we develop an assistance tool to identify the portion of server-side logic that requires computation over sensitive data. Our preliminary results show that the size of such logic is small in six popular web applications we study. To the extent of our evaluation, converting these applications to a CryptServer architecture requires modest developer effort. Zhaofeng Chen, Xinshu Dong, Prateek Saxena, Zhenkai Liang |
CCS | 4 |
| 2013 | Protecting sensitive web content from client-side vulnerabilities with CRYPTONSabstractWeb browsers isolate web origins, but do not provide direct abstractions to isolate sensitive data and control computation over it within the same origin. As a result, guaranteeing security of sensitive web content requires trusting all code in the browser and client-side applications to be vulnerability-free. In this paper, we propose a new abstraction, called Crypton, which supports intra-origin control over sensitive data throughout its life cycle. To securely enforce the semantics of Cryptons, we develop a standalone component called Crypton-Kernel, which extensively leverages the functionality of existing web browsers without relying on their large TCB. Our evaluation demonstrates that the Crypton abstraction supported by the Crypton-Kernel is widely applicable to popular real-world applications with millions of users, including webmail, chat, blog applications, and Alexa Top 50 websites, with low performance overhead. Xinshu Dong, Zhaofeng Chen, Hossein Siadati, Shruti Tople, Prateek Saxena, Zhenkai Liang |
CCS | 6 |
| 2013 | Enforcing system-wide control flow integrity for exploit detection and diagnosisabstractModern malware like Stuxnet is complex and exploits multiple vulnerabilites in not only the user level processes but also the OS kernel to compromise a system. A main trait of such exploits is manipulation of control flow. There is a pressing need to diagnose such exploits. Existing solutions that monitor control flow either have large overhead or high false positives and false negatives, hence making their deployment impractical. In this paper, we present Total-CFI, an efficient and practical tool built on a software emulator, capable of exploit detection by enforcing system-wide Control Flow Integrity (CFI). Total-CFI performs punctual guest OS view reconstruction to identify key guest kernel semantics like processes, code modules and threads. It incorporates a novel thread stack identification algorithm that identifies the stack boundaries for different threads in the system. Furthermore, Total-CFI enforces a CFI policy - a combination of whitelist based and shadow call stack based approaches to monitor indirect control flows and detect exploits. We provide a proof-of-concept implementation of Total-CFI on DECAF, built on top of Qemu. We tested 25 commonly used programs and 7 recent real world exploits on Windows OS and found 0 false positives and 0 false negatives respectively. The boot time overhead was found to be no more than 64.1% and the average memory overhead was found to be 7.46KB per loaded module, making it feasible for hardware integration. Aravind Prakash, Heng Yin 0001, Zhenkai Liang |
AsiaCCS | 3 |
| 2013 | A Quantitative Evaluation of Privilege Separation in Web Browser Designs
Xinshu Dong, Hong Hu 0004, Prateek Saxena, Zhenkai Liang |
ESORICS | 4 |
| 2013 | A Comprehensive Client-Side Behavior Model for Diagnosing Attacks in Ajax ApplicationsabstractBehavior models of applications are widely used for diagnosing security incidents in complex web-based systems. However, Ajax techniques that enable better web experiences also make it fairly challenging to model Ajax application behaviors in the complex browser environment. In Ajax applications, server-side states are no longer synchronous with the views to end users at the client side. Therefore, to model the behaviors of Ajax applications, it is indispensable to incorporate client-side application states into the behavior models, as being explored by prior work. Unfortunately, how to leverage behavior models to perform security diagnosis in Ajax applications has yet been thoroughly examined. Existing models extracted from Ajax application behaviors are insufficient in a security context. In this paper, we propose a new behavior model for diagnosing attacks in Ajax applications, which abstracts both client-side state transitions as well as their communications to external servers. Our model articulates different states with the browser events or user actions that trigger state transitions. With a prototype implementation, we demonstrate that the proposed model is effective in attack diagnosis for real-world Ajax applications. Xinshu Dong, Kailas Patil, Zhenkai Liang |
ICECCS | 4 |
| 2013 | A Software Environment for Confining Malicious Android Applications via Resource VirtualizationabstractIn the Android system, applications (apps) execute on the same platform that manages all system resources, where resource accesses are regulated through a permission-based mechanism. As a result, malicious apps get chances to abuse resources that are available on the Android platform. In this paper, we propose resource virtualization as a security mechanism to confine resource-abusing Android apps. The physical resources on a mobile device are virtualized to a different virtual view for selected Android apps. Resource virtualization simulates a partial but consistent virtual view of the Android resources. Therefore, it can not only confine the resource-abusing apps effectively, but also ensure the usability of them. We implement a system prototype, RVDroid, and evaluate it with real-world apps of various types. Our results demonstrate its effectiveness on malicious Android apps and its compatibility and usability on benign ones. Guangdong Bai, Zhenkai Liang, Heng Yin 0001 |
ICECCS | 3 |
| 2013 | Rating Web Pages Using Page-Transition Evidence
Xinshu Dong, Tao Wei 0002, Zhenkai Liang |
ICICS | 5 |
| 2013 | SafeStack: Automatically Patching Stack-Based Buffer Overflow VulnerabilitiesabstractBuffer overflow attacks still pose a significant threat to the security and availability of today's computer systems. Although there are a number of solutions proposed to provide adequate protection against buffer overflow attacks, most of existing solutions terminate the vulnerable program when the buffer overflow occurs, effectively rendering the program unavailable. The impact on availability is a serious problem on service-oriented platforms. This paper presents SafeStack, a system that can automatically diagnose and patch stack-based buffer overflow vulnerabilities. The key technique of our solution is to virtualize memory accesses and move the vulnerable buffer into protected memory regions, which provides a fundamental and effective protection against recurrence of the same attack without stopping normal system execution. We developed a prototype on a Linux system, and conducted extensive experiments to evaluate the effectiveness and performance of the system using a range of applications. Our experimental results showed that SafeStack can quickly generate runtime patches to successfully handle the attack's recurrence. Furthermore, SafeStack only incurs acceptable overhead for the patched applications. Hai Jin 0001, Deqing Zou, Bing Bing Zhou, Zhenkai Liang, Weide Zheng, Xuanhua Shi |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2012 | Tracking the Trackers: Fast and Scalable Dynamic Analysis of Web Content for Privacy Violations
Xinshu Dong, Zhenkai Liang, Xuxian Jiang |
ACNS | 3 |
| 2012 | Codejail: Application-Transparent Isolation of Libraries with Tight Program Interactions
Yongzheng Wu, Sai Sathyanarayan, Roland H. C. Yap, Zhenkai Liang |
ESORICS | 4 |
| 2012 | Detecting and Preventing ActiveX API-Misuse Vulnerabilities in Internet Explorer
Sai Sathyanarayan, Roland H. C. Yap, Zhenkai Liang |
ICICS | 4 |
| 2012 | An Empirical Study of Dangerous Behaviors in Firefox Extensions
Xiaohong Li 0001, Xuhui Liu, Xinshu Dong, Junjie Wang 0001, Zhenkai Liang, Zhiyong Feng 0002 |
ISC | 6 |
| 2012 | Identifying and Analyzing Pointer Misuses for Sophisticated Memory-corruption Exploit Diagnosis
Aravind Prakash, Zhenkai Liang, Heng Yin 0001 |
NDSS | 4 |
| 2012 | A Framework to Eliminate Backdoors from Response-Computable AuthenticationabstractResponse-computable authentication (RCA) is a two-party authentication model widely adopted by authentication systems, where an authentication system independently computes the expected user response and authenticates a user if the actual user response matches the expected value. Such authentication systems have long been threatened by malicious developers who can plant backdoors to bypass normal authentication, which is often seen in insider-related incidents. A malicious developer can plant backdoors by hiding logic in source code, by planting delicate vulnerabilities, or even by using weak cryptographic algorithms. Because of the common usage of cryptographic techniques and code protection in authentication modules, it is very difficult to detect and eliminate backdoors from login systems. In this paper, we propose a framework for RCA systems to ensure that the authentication process is not affected by backdoors. Our approach decomposes the authentication module into components. Components with simple logic are verified by code analysis for correctness, components with cryptographic/ obfuscated logic are sand boxed and verified through testing. The key component of our approach is NaPu, a native sandbox to ensure pure functions, which protects the complex and backdoor-prone part of a login module. We also use a testing-based process to either detect backdoors in the sand boxed component or verify that the component has no backdoors that can be used practically. We demonstrated the effectiveness of our approach in real-world applications by porting and verifying several popular login modules into this framework. Shuaifu Dai, Tao Wei 0002, Chao Zhang 0008, Tielei Wang, Zhenkai Liang |
IEEE Symposium on Security and Privacy | 6 |
| 2012 | DARWIN: An approach to debugging evolving programsabstractBugs in programs are often introduced when programs evolve from a stable version to a new version. In this article, we propose a new approach called DARWIN for automatically finding potential root causes of such bugs. Given two programs—a reference program and a modified program—and an input that fails on the modified program, our approach uses symbolic execution to automatically synthesize a new input that (a) is very similar to the failing input and (b) does not fail. We find the potential cause(s) of failure by comparing control-flow behavior of the passing and failing inputs and identifying code fragments where the control flows diverge. A notable feature of our approach is that it handles hard-to-explain bugs, like code missing errors, by pointing to code in the reference program. We have implemented this approach and conducted experiments using several real-world applications, such as the Apache Web server, libPNG (a library for manipulating PNG images), and TCPflow (a program for displaying data sent through TCP connections). In each of these applications, DARWIN was able to localize bugs with high accuracy. Even though these applications contain several thousands of lines of code, DARWIN could usually narrow down the potential root cause(s) to less than ten lines. In addition, we find that the inputs synthesized by DARWIN provide additional value by revealing other undiscovered errors. Dawei Qi, Abhik Roychoudhury, Zhenkai Liang, Kapil Vaswani |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2011 | AdSentry: comprehensive and flexible confinement of JavaScript-based advertisementsabstractInternet advertising is one of the most popular online business models. JavaScript-based advertisements (ads) are often directly embedded in a web publisher's page to display ads relevant to users (e.g., by checking the user's browser environment and page content). However, as third-party code, the ads pose a significant threat to user privacy. Worse, malicious ads can exploit browser vulnerabilities to compromise users' machines and install malware. To protect users from these threats, we propose AdSentry, a comprehensive confinement solution for JavaScript-based advertisements. The crux of our approach is to use a shadow JavaScript engine to sandbox untrusted ads. In addition, AdSentry enables flexible regulation on ad script behaviors by completely mediating its access to the web page (including its DOM) without limiting the JavaScript functionality exposed to the ads. Our solution allows both web publishers and end users to specify access control policies to confine ads' behaviors. We have implemented a proof-of-concept prototype of AdSentry that transparently supports the Mozilla Firefox browser. Our experiments with a number of ads-related attacks successfully demonstrate its practicality and effectiveness. The performance measurement indicates that our system incurs a small performance overhead. Xinshu Dong, Zhenkai Liang, Xuxian Jiang |
ACSAC | 3 |
| 2011 | Jump-oriented programming: a new class of code-reuse attackabstractReturn-oriented programming is an effective code-reuse attack in which short code sequences ending in a ret instruction are found within existing binaries and executed in arbitrary order by taking control of the stack. This allows for Turing-complete behavior in the target program without the need for injecting attack code, thus significantly negating current code injection defense efforts (e.g., W⊕X). On the other hand, its inherent characteristics, such as the reliance on the stack and the consecutive execution of return-oriented gadgets, have prompted a variety of defenses to detect or prevent it from happening. Tyler K. Bletsch, Xuxian Jiang, Vincent W. Freeh, Zhenkai Liang |
AsiaCCS | 4 |
| 2011 | Towards Fine-Grained Access Control in JavaScript ContextsabstractA typical Web 2.0 application usually includes JavaScript from various sources with different trust. It is critical to properly regulate JavaScript's access to web application resources. Unfortunately, existing protection mechanisms in web browsers do not provide enough granularity in JavaScript access control. Specifically, existing solutions partially mitigate this sort of threat by only providing access control for certain types of JavaScript objects, or by unnecessarily restricting the functionality of untrusted JavaScript. In this paper, we systematically analyze the complete access control requirements in a web browser's JavaScript environment and identify the fundamental lack of fine-grained JavaScript access control mechanisms in modern web browsers. As our solution, we propose a reference monitor called JCShadow that enables fine-grained access control in JavaScript contexts without unnecessarily restricting the functionality of JavaScript. We have developed a proof-of-concept prototype in the Mozilla Firefox browser and the evaluation with real-world attacks indicates that JCShadow effectively prevents such attacks with low performance overhead. Kailas Patil, Xinshu Dong, Zhenkai Liang, Xuxian Jiang |
ICDCS | 4 |
| 2010 | Heap Taichi: exploiting memory allocation granularity in heap-spraying attacksabstractHeap spraying is an attack technique commonly used in hijacking browsers to download and execute malicious code. In this attack, attackers first fill a large portion of the victim process's heap with malicious code. Then they exploit a vulnerability to redirect the victim process's control to attackers' code on the heap. Because the location of the injected code is not exactly predictable, traditional heap-spraying attacks need to inject a huge amount of executable code to increase the chance of success. Injected executable code usually includes lots of NOP-like instructions leading to attackers' shellcode. Targeting this attack characteristic, previous solutions detect heap-spraying attacks by searching for the existence of such large amount of NOP sled and other shellcode. Tao Wei 0002, Tielei Wang, Zhenkai Liang |
ACSAC | 4 |
| 2010 | Test generation to expose changes in evolving programsabstractSoftware constantly undergoes changes throughout its life cycle, and thereby it evolves. As changes are introduced into a code base, we need to make sure that the effect of the changes is thoroughly tested. For this purpose, it is important to generate test cases that can stress the effect of a given change. In this paper, we propose an automatic test generation solution to this problem. Given a change c, we use dynamic symbolic execution to generate a test input t, which stresses the change. This is done by ensuring (i) the change c is executed by t, and (ii) the effect of c is observable in the output produced by the test t. To construct a change-reaching input, our technique uses distance in control-dependency graph to guide path exploration towards the change. Then, our technique identifies the common programming patterns that may prevent a given change from affecting the program's output. For each of these patterns we propose methods to tune the change-reaching input into an input that reaches the change and propagates the effect of the change to the output. Our experimental results show that our test generation technique is effective in generating change-exposing inputs for real-world programs. Dawei Qi, Abhik Roychoudhury, Zhenkai Liang |
ASE | 3 |
| 2010 | Transparent Protection of Commodity OS Kernels Using Hardware Virtualization
Michael C. Grace, Zhi Wang 0004, Deepa Srinivasan, Jinku Li, Xuxian Jiang, Zhenkai Liang, Siarhei Liakh |
SecureComm | 6 |
| 2010 | Golden implementation driven software debuggingabstractThe presence of a functionally correct golden implementation has a significant advantage in the software development life cycle. Such a golden implementation is exploited for software development in several domains, including embedded software --- a low resource-consuming version of the golden implementation. The golden implementation gives the functionality that the program is supposed to implement, and is used as a guide during the software development process. In this paper, we investigate the possibility of using the golden implementation as a reference model in software debugging. We perform a substantial case study involving the Busybox embedded Linux utilities while treating the GNU Core Utilities as the golden or reference implementation. Our debugging method consists of dynamic slicing with respect to the observable error in both the implementations (the golden implementation as well as the buggy software). During dynamic slicing we also perform a step-by-step weakest precondition computation of the observable error with respect to the statements in the dynamic slice. The formulae computed as weakest pre-condition in the two implementations are then compared to accurately locate the root cause of a given observable error. Experimental results obtained from Busybox suggest that our method performs well in practice and is able to pinpoint all the bugs recently published in [8] that could be reproduced on Busybox version 1.4.2. The bug report produced by our approach is concise and pinpoints the program locations inside the Busybox source that contribute to the difference in behavior. Ansuman Banerjee, Abhik Roychoudhury, Johannes A. Harlie, Zhenkai Liang |
SIGSOFT FSE | 4 |
| 2009 | Towards Generating High Coverage Vulnerability-Based Signatures with Protocol-Level Constraint-Guided Exploration
Juan Caballero, Zhenkai Liang, Pongsin Poosankam, Dawn Song |
RAID | 2 |
| 2009 | Darwin: an approach for debugging evolving programsabstractDebugging refers to the laborious process of finding causes of program failures. Often, such failures are introduced when a program undergoes changes and evolves from a stable version to a new, modified version. In this paper, we propose an automated approach for debugging evolving programs. Given two programs (a reference, stable program and a new, modified program) and an input that fails on the modified program, our approach uses concrete as well as symbolic execution to synthesize new inputs that differ marginally from the failing input in their control flow behavior. A comparison of the execution traces of the failing input and the new inputs provides critical clues to the root-cause of the failure. A notable feature of our approach is that it handles hard-to-explain bugs like code missing errors by pointing to the relevant code in the reference program. We have implemented our approach in a tool called DARWIN. We have conducted experiments with several real-life case studies, including real-world web servers and the libPNG library for manipulating PNG images. Our experience from these experiments points to the efficacy of DARWIN in pinpointing bugs. Moreover, while localizing a given observable error, the new inputs synthesized by DARWIN can reveal other undiscovered errors. Dawei Qi, Abhik Roychoudhury, Zhenkai Liang, Kapil Vaswani |
ESEC/SIGSOFT FSE | 3 |
| 2009 | Alcatraz: An Isolated Environment for Experimenting with Untrusted SoftwareabstractIn this article, we present an approach for realizing a safe execution environment (SEE) that enables users to “try out” new software (or configuration changes to existing software) without the fear of damaging the system in any manner. A key property of our SEE is that it faithfully reproduces the behavior of applications, as if they were running natively on the underlying (host) operating system. This is accomplished via one-way isolation: processes running within the SEE are given read-access to the environment provided by the host OS, but their write operations are prevented from escaping outside the SEE. As a result, SEE processes cannot impact the behavior of host OS processes, or the integrity of data on the host OS. SEEs support a wide range of tasks, including: study of malicious code, controlled execution of untrusted software, experimentation with software configuration changes, testing of software patches, and so on. It provides a convenient way for users to inspect system changes made within the SEE. If these changes are not accepted, they can be rolled back at the click of a button. Otherwise, the changes can be committed so as to become visible outside the SEE. We provide consistency criteria that ensure semantic consistency of the committed results. We develop two different implementation approaches, one in user-land and the other in the OS kernel , for realizing a safe-execution environment. Our implementation results show that most software, including fairly complex server and client applications, can run successfully within our SEEs. It introduces low performance overheads, typically below 10 percent. Zhenkai Liang, Weiqing Sun, V. N. Venkatakrishnan, R. Sekar 0001 |
ACM Trans. Inf. Syst. Secur. | 1 |
| 2008 | Expanding Malware Defense by Securing Software Installations
Weiqing Sun, R. Sekar 0001, Zhenkai Liang, V. N. Venkatakrishnan |
DIMVA | 3 |
| 2008 | AGIS: Towards automatic generation of infection signaturesabstractAn important yet largely uncharted problem in malware defense is how to automate generation of infection signatures for detecting compromised systems, i.e., signatures that characterize the behavior of malware residing on a system. To this end, we develop AGIS, a host-based technique that detects infections by malware and automatically generates an infection signature of the malware. AGIS monitors the runtime behavior of suspicious code according to a set of security policies to detect an infection, and then identifies its characteristic behavior in terms of system or API calls. AGIS then statically analyzes the corresponding executables to extract the instructions important to the infectionpsilas mission. These instructions can be used to build a template for a static-analysis-based scanner, or a regular-expression signature for legacy scanners. AGIS also detects encrypted malware and generates a signature from its plaintext decryption loop. We implemented AGIS on Windows XP and evaluated it against real-life malware, including keyloggers, mass-mailing worms, and a well-known mutation engine. The experimental results demonstrate the effectiveness of our technique in detecting new infections and generating high-quality signatures. Zhuowei Li 0001, XiaoFeng Wang 0001, Zhenkai Liang, Michael K. Reiter |
DSN | 3 |
| 2008 | HookFinder: Identifying and Understanding Malware Hooking Behaviors
Heng Yin 0001, Zhenkai Liang, Dawn Song |
NDSS | 2 |
| 2007 | Polyglot: automatic extraction of protocol message format using dynamic binary analysisabstractProtocol reverse engineering, the process of extracting the application-level protocol used by an implementation, without access to the protocol specification, is important for many network security applications. Recent work [17] has proposed protocol reverse engineering by using clustering on network traces. That kind of approach is limited by the lack of semantic information on network traces. In this paper we propose a new approach using program binaries. Our approach, shadowing, uses dynamic analysis and is based on a unique intuition - the way that an implementation of the protocol processes the received application data reveals a wealth of information about the protocol message format. We have implemented our approach in a system called Polyglot and evaluated it extensively using real-world implementations of five different protocols: DNS, HTTP, IRC, Samba and ICQ. We compare our results with the manually crafted message format, included in Wireshark, one of the state-of-the-art protocol analyzers. The differences we find are small and usually due to different implementations handling fields in different ways. Finding such differences between implementations is an added benefit, as they are important for problems such as fingerprint generation, fuzzing, and error detection. Juan Caballero, Heng Yin 0001, Zhenkai Liang, Dawn Song |
CCS | 3 |
| 2007 | Towards Automatic Discovery of Deviations in Binary Implementations with Applications to Error Detection and Fingerprint Generation
David Brumley, Juan Caballero, Zhenkai Liang, James Newsome |
USENIX Security Symposium | 3 |
| 2005 | Automatic Generation of Buffer Overflow Attack Signatures: An Approach Based on Program Behavior ModelsabstractBuffer overflows have become the most common target for network-based attacks. They are also the primary mechanism used by worms and other forms of automated attacks. Although many techniques have been developed to prevent server compromises due to buffer overflows, these defenses still lead to server crashes. When attacks occur repeatedly, as is common with automated attacks, these protection mechanisms lead to repeated restarts of the victim application, rendering its service unavailable. To overcome this problem, we develop a new approach that can learn the characteristics of a particular attack, and filter out future instances of the same attack or its variants. By doing so, our approach significantly increases the availability of servers subjected to repeated attacks. The approach is fully automatic, does not require source code, and has low runtime overheads. In our experiments, it was effective against most attacks, and did not produce any false positives. Zhenkai Liang, R. Sekar 0001 |
ACSAC | 1 |
| 2005 | Fast and automated generation of attack signatures: a basis for building self-protecting serversabstractLarge-scale attacks, such as those launched by worms and zombie farms, pose a serious threat to our network-centric society. Existing approaches such as software patches are simply unable to cope with the volume and speed with which new vulnerabilities are being discovered. In this paper, we develop a new approach that can provide effective protection against a vast majority of these attacks that exploit memory errors in C/C++ programs. Our approach, called COVERS, uses a forensic analysis of a victim server's memory to correlate attacks to inputs received over the network, and automatically develop a signature that characterizes inputs that carry attacks. The signatures tend to capture characteristics of the underlying vulnerability (e.g., a message field being too long) rather than the characteristics of an attack, which makes them effective against variants of attacks. Our approach introduces low overheads (under 10%), does not require access to source code of the protected server, and has successfully generated signatures for the attacks studied in our experiments, without producing false positives. Since the signatures are generated in tens of milliseconds, they can potentially be distributed quickly over the Internet to filter out (and thus stop) fast-spreading worms. Another interesting aspect of our approach is that it can defeat guessing attacks reported against address-space randomization and instruction set randomization techniques. Finally, it increases the capacity of servers to withstand repeated attacks by a factor of 10 or more. Zhenkai Liang, R. Sekar 0001 |
CCS | 1 |
| 2005 | One-Way Isolation: An Effective Approach for Realizing Safe Execution Environments
Weiqing Sun, Zhenkai Liang, V. N. Venkatakrishnan, R. Sekar 0001 |
NDSS | 2 |
| 2005 | Automatic Synthesis of Filters to Discard Buffer Overflow Attacks: A Step Towards Realizing Self-Healing Systems
Zhenkai Liang, R. Sekar 0001, Daniel C. DuVarney |
USENIX ATC, General Track | 1 |
| 2003 | Isolated Program Execution: An Application Transparent Approach for Executing Untrusted ProgramsabstractWe present a new approach for safe execution of untrusted programs by isolating their effects from the rest of the system. Isolation is achieved by intercepting file operations made by untrusted processes, and redirecting any change operations to a "modification cache" that is invisible to other processes in the system. File read operations performed by the untrusted process are also correspondingly modified, so that the process has a consistent view of system state that incorporates the contents of the file system as well as the modification cache. On termination of the untrusted process, its user is presented with a concise summary of the files modified by the process. Additionally, the user can inspect these files using various software utilities (e.g., helper applications to view multimedia files) to determine if the modifications are acceptable. The user then has the option to commit these modifications, or simply discard them. Essentially, our approach provides "play" and "rewind" buttons for running untrusted software. Key benefits of our approach are that it requires no changes to the untrusted programs (to be isolated) or the underlying operating system; it cannot be subverted by malicious programs; and it achieves these benefits with acceptable runtime overheads. We describe a prototype implementation of this system for Linux called Alcatraz and discuss its performance and effectiveness. Zhenkai Liang, V. N. Venkatakrishnan, R. Sekar 0001 |
ACSAC | 1 |
| 2002 | An Approach for Secure Software Installation
V. N. Venkatakrishnan, R. Sekar 0001, T. Kamat, S. Tsipa, Zhenkai Liang |
LISA | 5 |