EDBT 2026 Demo / reviewers in the wild / expert
Tao Wei 0002
dblp:64/5099-2 · also Tao (Lenx) Wei
· DBLP profile ↗
67ranked-venue papers
3as first author
31since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 39 · 1 first-author · 19 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Computer networks · 8 · 1 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revisiting the Reliability of Language Models in Instruction-FollowingabstractAdvanced LLMs have achieved near-ceiling instruction-following accuracy on benchmarks such as IFEVAL.However, these impressive scores do not necessarily translate to reliable services in real-world use, where users often vary their phrasing, contextual framing, and task formulations.In this paper, we study nuance-oriented reliability: whether models exhibit consistent competence across cousin prompts that convey analogous user intents but with subtle nuances.To quantify this, we introduce a new metric, reliable@k, and develop an automated pipeline that generates high-quality cousin prompts via data augmentation.Building upon this, we construct IFE-VAL++ for systematic evaluation.Across 20 proprietary and 26 open-source LLMs, we find that current models exhibit substantial insufficiency in nuance-oriented reliability-their performance can drop by up to 61.8% with nuanced prompt modifications.What's more, we characterize it and explore three potential improvement recipes.Our findings highlight nuance-oriented reliability as a crucial yet underexplored next step toward more dependable and trustworthy LLM behavior.Our code and benchmark are accessible: https: //github.com/jianshuod/IFEval-pp. Jianshuo Dong, Liu Yan, Zhenyu Zhong, Tao Wei 0002, Chao Zhang 0008, Han Qiu 0001 |
ACL (1) | 5 |
| 2026 | SoK: Analysis of Accelerator TEE Designs
Chenxu Wang 0005, Yujun Liang, Xuanyao Peng, Yuqun Zhang, Fengwei Zhang, Jiannong Cao 0001, Rui Hou 0001, Shoumeng Yan, Tao Wei 0002, Zhengyu He |
NDSS | 11 |
| 2026 | LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning
Wei Huang 0039, Anda Cheng, Yinggui Wang, Lei Wang 0251, Tao Wei 0002 |
Proc. VLDB Endow. | 5 |
| 2026 | Building Confidential Accelerator Computing Environment for Arm CCA
Chenxu Wang 0005, Fengwei Zhang, Yunjie Deng 0001, Kevin Leach, Jiannong Cao 0001, Zhenyu Ning, Shoumeng Yan, Tao Wei 0002, Zhengyu He |
IEEE Trans. Dependable Secur. Comput. | 9 |
| 2025 | Panther: Private Approximate Nearest Neighbor Search in the Single Server SettingabstractApproximate nearest neighbor search (ANNS), also known as vector search, is an important building block for various applications, such as recommendation systems, biometric authentication, and machine learning. In this work, we are interested in the private ANNS problem, where the client wants to learn (and can only learn) the ANNS results without revealing the query to the server. Previous private ANNS works either suffer from high communication cost (Chen et al., USENIX Security 2020) or work under a stronger security assumption of two non-colluding servers (Servan-Schreiber et al., SP 2022). We present Panther, an efficient private ANNS framework under the single server setting. Panther achieves its high performance via several novel co-designs of private information retrieval, secret-sharing, garbled circuits, and homomorphic encryption. We made extensive experiments using Panther on four public datasets, showing that Panther could answer an ANNS query on 10 million points in 18 seconds with 284 MB of communication. This is more than 7.8× faster and 20× more compact than Chen et al. Min Zhang 0043, Cheng Hong 0001, Jian Liu 0012, Tao Wei 0002 |
CCS | 6 |
| 2025 | PromeFuzz: A Knowledge-Driven Approach to Fuzzing Harness Generation with Large Language ModelsabstractAPI-level fuzzing has become increasingly important for discovering subtle bugs in modern software, yet generating effective fuzzing harnesses remains a complex and error-prone task. Existing approaches often rely on limited consumer code or shallow program analysis, which fail to capture deep API semantics and interdependencies, resulting in poor coverage and high false positive rates. Recent methods incorporating Large Language Models (LLMs) have improved harness generation by leveraging pretrained knowledge, but they still struggle with hallucinations and lack domain-specific understanding. Yuwei Liu 0001, Junquan Deng, Xiangkun Jia, Lin Huang 0005, Tao Wei 0002, Purui Su |
CCS | 7 |
| 2025 | "I've Decided to Leak": Probing Internals Behind Prompt Leakage IntentsabstractLarge language models (LLMs) exhibit prompt leakage vulnerabilities, where they may be coaxed into revealing system prompts embedded in LLM services, raising intellectual property and confidentiality concerns.An intriguing question arises: Do LLMs genuinely internalize prompt leakage intents in their hidden states before generating tokens?In this work, we use probing techniques to capture LLMs' intent-related internal representations and confirm that the answer is yes.We start by comprehensively inducing prompt leakage behaviors across diverse system prompts, attack queries, and decoding methods.We develop a hybrid labeling pipeline, enabling the identification of broader prompt leakage behaviors beyond mere verbatim leaks.Our results show that a simple linear probe can predict prompt leakage risks from pre-generation hidden states without generating any tokens.Across all tested models, linear probes consistently achieve 90%+ AUROC, even when applied to new system prompts and attacks.Understanding the model internals behind prompt leakage drives practical applications, including intention-based detection of prompt leakage risks. Jianshuo Dong, Liu Yan, Zhenyu Zhong, Tao Wei 0002, Ke Xu 0002, Minlie Huang, Chao Zhang 0008, Han Qiu 0001 |
EMNLP | 5 |
| 2025 | ccAI: A Compatible and Confidential System for AI ComputingabstractConfidential xPU computing has emerged as a prominent technique for effectively securing users' AI computing workloads on heterogeneous systems equipped with xPUs.Although the industry adopts this technology in cutting-edge hardware (e.g.NVIDIA H100 GPU) to safeguard high-performance AI computing, most clouds still rely on legacy xPUs and suffer from data leakage problems. Chenxu Wang 0005, Danqing Tang, Changxu Ci, Yankai Xu, Fengwei Zhang, Jiannong Cao 0001, Shoumeng Yan, Tao Wei 0002, Zhengyu He |
MICRO | 10 |
| 2025 | AnchorSync: Global Consistency Optimization for Long Video Editing
Zichi Liu, Yinggui Wang, Tao Wei 0002, Chao Ma 0004 |
ACM Multimedia | 3 |
| 2025 | SCRUTINIZER: Towards Secure Forensics on Compromised TrustZone
Yiming Zhang 0030, Fengwei Zhang, Xiapu Luo, Rui Hou 0001, Xuhua Ding, Zhenkai Liang, Shoumeng Yan, Tao Wei 0002, Zhengyu He |
NDSS | 8 |
| 2025 | RACONTEUR: A Knowledgeable, Insightful, and Portable LLM-Powered Shell Command Explainer
Jiangyi Deng, Xinfeng Li, Yanjiao Chen, Yijie Bai, Haiqin Weng, Yan Liu 0069, Tao Wei 0002, Wenyuan Xu 0001 |
NDSS | 7 |
| 2025 | BumbleBee: Secure Two-party Inference Framework for Large Transformers
Jian Liu 0012, Cheng Hong 0001, Kui Ren 0001, Tao Wei 0002 |
NDSS | 8 |
| 2025 | AegisGuard: RL-Guided Adapter Tuning for TEE-Based Efficient & Secure On-Device InferenceabstractOn-device large models (LMs) reduce cloud dependency but expose proprietary model weights to the end-user, making them vulnerable to white-box model stealing (MS) attacks. A common defense is TEE-Shielded DNN Partition (TSDP), which places all trainable LoRA adapters (fine tuned on private data) inside a trusted execution environment (TEE). However, this design suffers from excessive host-to-TEE communication latency. We propose AegisGuard, a fine tuning and deployment framework that selectively shields the MS sensitive adapters while offloading the rest to the GPU, balancing security and efficiency. AegisGuard integrates two key components: i) RL-based Sensitivity Measurement (RSM), which injects Gaussian noise during training and applies a lightweight reinforcement learning to rank adapters based on their impact on model stealing; and (ii) Shielded-Adapter Compression (SAC), which structurally prunes the selected adapters to reduce both parameter size and intermediate feature maps, further lowering TEE computation and data transfer costs. Extensive experiments demonstrate that AegisGuard achieves black-box level MS resilience (surrogate accuracy around 39%, matching fully shielded baselines), while reducing end-to-end inference latency by 2–3× and cutting TEE memory usage by 4× compared to state-of-the-art TSDP methods. Ziqi Zhang 0017, Yinggui Wang, Tiantong Wang, Yurong Hao, Tao Wei 0002, Yang Cao 0011, Wei Yang Bryan Lim |
NeurIPS | 7 |
| 2025 | MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM InferenceabstractPrivate large language model (LLM) inference based on secure multi-party computation (MPC) achieves formal data privacy protection but suffers from significant latency overhead, especially for long input sequences. While key-value (KV) cache eviction and sparse attention algorithms have been proposed for efficient LLM inference in plaintext, they are not designed for MPC and cannot benefit private LLM inference directly. In this paper, we propose an accurate and MPC-friendly KV cache eviction framework, dubbed MPCache, building on the observation that historical tokens in a long sequence may have different effects on the downstream decoding. Hence, MPCache combines a look-once static eviction algorithm to discard unimportant KV cache and a query-aware dynamic selection algorithm to activate only a small subset of KV cache for attention computation. MPCache further incorporates a series of optimizations for efficient dynamic KV cache selection, including MPC-friendly similarity approximation, hierarchical KV cache clustering, and cross-layer index-sharing strategy. Extensive experiments demonstrate that MPCache consistently outperforms prior-art KV cache eviction baselines across different generation tasks and achieves 1.8 ~ 2.01x and 3.39 ~ 8.37x decoding latency and communication reduction on different sequence lengths, respectively. Wenxuan Zeng, Ye Dong, Jinjin Zhou, Lei Wang 0251, Tao Wei 0002, Runsheng Wang, Meng Li 0004 |
NeurIPS | 6 |
| 2025 | ZHE: Efficient Zero-Knowledge Proofs for HE EvaluationsabstractHomomorphic Encryption (HE) allows computations on encrypted data without decryption. It can be used where the users' information are to be processed by an untrustful server, and has been a popular choice in privacy-preserving applications. However, in order to obtain meaningful results, we have to assume an honest-but-curious server, i.e., it will faithfully follow what was asked to do. If the server is malicious, there is no guarantee that the computed result is correct. The notion of verifiable HE (vHE) is introduced to detect malicious server's behaviors, but current vHE schemes are either more than four orders of magnitude slower than the underlying HE operations (Atapoor et. al, CIC 2024) or fast but incompatible with server-side private inputs (Chatel et. al, CCS 2024). In this work, we propose a vHE framework ZHE: efficient Zero-Knowledge Proofs (ZKPs) that prove the correct execution of HE evaluations while protecting the server's private inputs. More precisely, we first design two new highly-efficient ZKPs for modulo operations and (Inverse) Number Theoretic Transforms (NTTs), two of the basic operations of HE evaluations. Then we build a customized ZKP for HE evaluations, which is scalable, enjoys a fast prover time and has a non-interactive online phase. Our ZKP is applicable to all Ring-LWE based HE schemes, such as BGV and CKKS. Finally, we implement our protocols for both BGV and CKKS and conduct extensive experiments on various HE workloads. Compared to the state-of-the-art works, both of our prover time and verifier time are improved; especially, our prover cost is only roughly 27–36× more expensive than the underlying HE operations, this is two to three orders of magnitude cheaper than state-of-the-arts. Zhelei Zhou, Yun Li 0010, Zhaomin Yang, Bingsheng Zhang, Cheng Hong 0001, Tao Wei 0002 |
SP | 7 |
| 2025 | PoiSAFL: Scalable Poisoning Attack Framework to Byzantine-resilient Semi-asynchronous Federated Learning
Xiaoyi Pang, Zhibo Wang 0001, Jiahui Hu 0001, Yinggui Wang, Lei Wang 0251, Tao Wei 0002, Kui Ren 0001, Chun Chen 0001 |
USENIX Security Symposium | 7 |
| 2025 | On the Interoperability of Encrypted DatabasesabstractEncrypted database is an emerging and promising technology. It is able to run SQL operations on encrypted data. However, most existing encrypted databases haveno data interoperability, i.e., the output of an operator (e.g., addition) cannot be taken as input of another (e.g., comparison). As a result, these encrypted databases can only support simple queries like addition, multiplication and comparison, but unable to support a composition of these simple queries (e.g.,SELECT user_id FROM salary WHERE$V_{1} + V_{2} > 5000$V1+V2>5000). In SIGMOD ’14, Wong et al. propose SDB, which to the best of our knowledge is the only encrypted database that achieves data interoperability. Unfortunately, it has recently been broken (VLDB ’21). In this paper, we propose a novel encrypted database namedSDB+. It achieves data interoperability based on a suit of sophisticated designs. We formally prove thatSDB+achieves indistinguishability under chosen query attacks (IND-CQA). We provide a full-fledged implementation and run it on three benchmarks. Our experimental results show thatSDB+achieves comparable efficiency with SDB, even though the latter is insecure. Xinle Cao, Jian Liu 0012, Yan Liu 0069, Tao Wei 0002, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | Towards Sample-Specific Backdoor Attack With Clean Labels via Attribute TriggerabstractCurrently, sample-specific backdoor attacks (SSBAs) are the most advanced and malicious methods since they can easily circumvent most of the current backdoor defenses. In this paper, we reveal that SSBAs are not sufficiently stealthy due to their poisoned-label nature, where users can discover anomalies if they check the image-label relationship. In particular, we demonstrate that it is ineffective to directly generalize existing SSBAs to their clean-label variants by poisoning samples solely from the target class. We reveal that it is primarily due to two reasons, including(1)the ‘antagonistic effects’ of ground-truth features and(2)the learning difficulty of sample-specific features. Accordingly, trigger-related features of existing SSBAs cannot be effectively learned under the clean-label setting due to their mild trigger intensity required for ensuring stealthiness. We argue that the intensity constraint of existing SSBAs is mostly because their trigger patterns are ‘content-irrelevant’ and therefore act as ‘noises’ for both humans and DNNs. Motivated by this understanding, we propose to exploit content-relevant features,$a.k.a.$(human-relied) attributes, as the trigger patterns to design clean-label SSBAs. This new attack paradigm is dubbed backdoor attack with attribute trigger (BAAT). Extensive experiments are conducted on benchmark datasets, which verify the effectiveness of our BAAT and its resistance to existing defenses. Mingyan Zhu 0001, Yiming Li 0004, Tao Wei 0002, Shutao Xia, Zhan Qin |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | Constructing SDN Covert Timing Channels Between Hosts With Unprivileged AttackersabstractSoftware-defined networking (SDN) has been widely deployed due to its centralization and programmable features. However, these new features bring new threats at the same time. Previous studies have shown that SDN covert channels can be built with a privileged adversary that controls SDN key components, such as controller applications or SDN switches. In this paper, we propose new SDN covert timing channels between hosts without controlling applications, controllers, or having access to switches. Experiments in a real SDN testbed demonstrate the feasibility and effectiveness of our covert channels. To defend against the covert timing channels, we design a defense system named CovertGuard, which utilizes the timing characteristics of the covert channels’ delays to detect and eliminate covert channels effectively. Yixiong Ji, Jiahao Cao 0001, Qi Li 0002, Yan Liu 0069, Tao Wei 0002, Ke Xu 0002 |
IEEE Trans. Netw. | 5 |
| 2024 | Coral: Maliciously Secure Computation Framework for Packed and Mixed CircuitsabstractAchieving malicious security with high efficiency in dishonest-majority secure multiparty computation is a formidable challenge. The milestone works SPDZ and TinyOT have spawn a large family of protocols in this direction. For boolean circuits, state-of-the-art works (Cascudo et. al, TCC 2020 and Escudero et. al, CRYPTO 2022) have proposed schemes based on reverse multiplication-friendly embedding (RMFE) to reduce the amortized cost. However, these protocols are theoretically described and analyzed, resulting in a significant gap between theory and concrete efficiency. Cheng Hong 0001, Tao Wei 0002 |
CCS | 5 |
| 2024 | ProFake: Detecting Deepfakes in the Wild against Quality Degradation with Progressive Quality-adaptive LearningabstractDespite the promising advances in deepfake detection on current datasets, detecting visual deepfakes in real-world scenarios (e.g., deepfake videos and live streaming on YouTube) remains a challenge due to the inherent quality degradation such as unpredictable compression employed by social media platforms. Such degradation perturbs discernible forgery clues and diminishes the effectiveness of deepfake detection methods, raising a critical safety concern to the misuse of forgery faces in real-world scenarios. In this paper, we aim to understand the impacts of real-world degradation on the robustness of deepfake detection. Particularly, we investigate the risk of degraded deepfakes towards their detection on two real-world scenarios (i.e., deepfake videos and deepfake live streaming on social media platforms). By measuring the effects of real-world degradations on the performance and representation capabilities of detection models, we reveal that real-world deepfakes can be simulated via common degradation operations (e.g., JPEG compression) as they are perceptually similar to deepfake detectors. By analyzing the training dynamics under different sequences of training samples, we observe that the training order of deepfakes progressing from non-degraded (easy) to heavily degraded (hard) enhances the adaptability of detection models to various degradation in real-world scenarios. Drawing from these observations, we present a novel deepfake detection method ProFake to enhance the robustness of deepfake detection against real-world quality degradations. ProFake enables quality-adaptive learning via progressively degrade, detect and assign weights for the training samples driven by the feedback of model performance and image quality, which ensures that our model gradually focuses on more challenging samples to achieve quality-adaptive deepfake detection. Extensive experiments show that compared with existing methods, ProFake improves deepfake detection accuracy by an average of over 10 % in real-world scenarios and by an average of over 30 % in heavily degraded scenarios, while maintaining comparable performance in detecting high-quality deepfakes. Huiyu Xu, Yaopeng Wang, Zhibo Wang 0001, Zhongjie Ba, Haiqin Weng, Tao Wei 0002, Kui Ren 0001 |
CCS | 8 |
| 2024 | UnsafeCop: Towards Memory Safety for Real-World Unsafe Rust Code with Practical Bounded Model CheckingabstractAbstract Rust has gained popularity as a safer alternative to C/C++ for low-level programming due to its memory-safety features and minimal runtime overhead. However, the use of the “unsafe” keyword allows developers to bypass safety guarantees, posing memory-safety risks. Bounded Model Checking (BMC) is commonly used to detect memory-safety problems, but it has limitations for large-scale programs, as it can only detect bugs within a bounded number of executions. In this paper, we introduce UnsafeCop that utilizes and enhances BMC for analyzing memory safety in real-world unsafe Rust code. Our methodology incorporates harness design, loop bound inference, and both loop and function stubbing for comprehensive analysis. We optimize verification efficiency through a strategic function verification order, leveraging both types of stubbing. We conducted a case study on TECC (Trusted-Environment-based Cryptographic Computing), a proprietary framework consisting of 30,174 lines of Rust code, including 3,019 lines of unsafe Rust code, developed by Ant Group. Experimental results demonstrate that UnsafeCop effectively detects and verifies dozens of memory safety issues, reducing verification time by 73.71% compared to the traditional non-stubbing approach, highlighting its practical effectiveness. Jingling Xue, Lin Huang 0005, Yuan Zi, Tao Wei 0002 |
FM (2) | 5 |
| 2023 | Counterfactual-based Saliency Map: Towards Visual Contrastive Explanations for Neural NetworksabstractExplaining deep models in a human-understandable way has been explored by many works that mostly explain why an input causes a corresponding prediction (i.e., Why P?). However, seldom they could handle those more complex causal questions like "Why P rather than Q?" and "Why one is P, while another is Q?", which would better help humans understand the behavior of deep models. Considering the insufficient study on such complex causal questions, we make the first attempt to explain different causal questions by contrastive explanations in a unified framework, i.e., Counterfactual Contrastive Explanation (CCE), which visually and intuitively explains the aforementioned questions via a novel positive-negative saliency-based explanation scheme. More specifically, we propose a content-aware counterfactual perturbing algorithm to stimulate contrastive examples, from which a pair of positive and negative saliency maps could be derived to contrastively explain why P (positive class) rather than Q (negative class). Beyond existing works, our counterfactual perturbation meets the principles of validity, sparsity, and data distribution closeness at the same time. In addition, by slightly adjusting the objective of perturbation, our framework can adapt to different causal questions. Extensive experimental evaluation demonstrates the effectiveness and superior performance of the proposed CCE on different benchmark metrics for interpretability, including Sanity Check, Class Deviation Score and Insertion-Deletion tests. A user study is conducted and the results show that user confidence is increasing significantly when presented with CCE compared to standard saliency map baselines. Zhibo Wang 0001, Haiqin Weng, Hengchang Guo, Tao Wei 0002, Kui Ren 0001 |
ICCV | 7 |
| 2023 | ODDFuzz: Discovering Java Deserialization Vulnerabilities via Structure-Aware Directed Greybox FuzzingabstractJava deserialization vulnerability is a severe threat in practice. Researchers have proposed static analysis solutions to locate candidate vulnerabilities and fuzzing solutions to generate proof-of-concept (PoC) serialized objects to trigger them. However, existing solutions have limited effectiveness and efficiency.In this paper, we propose a novel hybrid solution ODDFuzz to efficiently discover Java deserialization vulnerabilities. First, ODDFuzz performs lightweight static taint analysis to identify candidate gadget chains that may cause deserialization vulnerabilities. In this step, ODDFuzz tries to locate all candidates and avoid false negatives. Then, ODDFuzz performs directed greybox fuzzing (DGF) to explore those candidates and generate PoC testcases to mitigate false positives. Specifically, ODDFuzz applies a structure-aware seed generation method to guarantee the validity of the testcases, and adopts a novel hybrid feedback and a step-forward strategy to guide the directed fuzzing.We implemented a prototype of ODDFuzz and evaluated it on the popular Java deserialization repository ysoserial. Results show that, ODDFuzz could discover 16 out of 34 known gadget chains, while two state-of-the-art baselines only identify three of them. In addition, we evaluated ODDFuzz on real-world applications including Oracle WebLogic Server, Apache Dubbo, Sonatype Nexus, and protostuff, and found six previously unreported exploitable gadget chains with five CVEs assigned. Sicong Cao, Biao He 0002, Xiaobing Sun 0001, Yu Ouyang, Chao Zhang 0008, Xiaoxue Wu 0001, Ting Su 0001, Lili Bo, Bin Li 0006, Chuanlei Ma, Tao Wei 0002 |
SP | 12 |
| 2023 | Aegis: Mitigating Targeted Bit-flip Attacks against Deep Neural Networks
Jialai Wang, Han Qiu 0001, Tianwei Zhang 0004, Qi Li 0002, Zongpeng Li, Tao Wei 0002, Chao Zhang 0008 |
USENIX Security Symposium | 8 |
| 2023 | CVTEE: A Compatible Verified TEE Architecture With Enhanced SecurityabstractSensitive resources in Trusted Execution Environment (TEE) have suffered serious security threats in recent years. Previous protection approaches either lack a strong assurance of TEE security properties or are limited to a single platform. We propose a compatible verified TEE architecture, calledCVTEE, which delegates a security monitor to manage TEE resources securely. This architecture has two key advantages: i) its functional correctness and security are guaranteed by a machine-checkable proof of security objectives of Trusted Application (TA) isolation, runtime confidentiality, and runtime integrity, and ii) it is applicable to different TEE platforms and implementation-independent due to its high level of abstraction and non-determinism of data types. Note that access control policy and information flow control policy are the core for security management of resources. After formally specifying the security attributes of TEE resources, we develop these policies based on Common Criteria (CC) in the security monitor and provide atomic interfaces.CVTEEis formally verified with 386 lemmas/theorems and$\sim$10,000 LOC of Isabelle/HOL. In addition, we implement a proof of concept for the access control module of Teaclave, and prove that the constructed access control model meets the security requirements through 5 theorems. Xinliang Miao, Jianhong Zhao, Yongwang Zhao, Shuang Cao, Tao Wei 0002, Liehui Jiang, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2023 | Black-Box Dataset Ownership Verification via Backdoor WatermarkingabstractDeep learning, especially deep neural networks (DNNs), has been widely and successfully adopted in many critical applications for its high effectiveness and efficiency. The rapid development of DNNs has benefited from the existence of some high-quality datasets (e.g., ImageNet), which allow researchers and developers to easily verify the performance of their methods. Currently, almost all existing released datasets require that they can only be adopted for academic or educational purposes rather than commercial purposes without permission. However, there is still no good way to ensure that. In this paper, we formulate the protection of released datasets as verifying whether they are adopted for training a (suspicious) third-party model, where defenders can only query the model while having no information about its parameters and training details. Based on this formulation, we propose to embed external patterns via backdoor watermarking for the ownership verification to protect them. Our method contains two main parts, including dataset watermarking and dataset verification. Specifically, we exploit poison-only backdoor attacks (e.g., BadNets) for dataset watermarking and design a hypothesis-test-guided method for dataset verification. We also provide some theoretical analyses of our methods. Experiments on multiple benchmark datasets of different tasks are conducted, which verify the effectiveness of our method. The code for reproducing main experiments is available at https://github.com/THUYimingLi/DVBW. Yiming Li 0004, Mingyan Zhu 0001, Xue Yang 0003, Yong Jiang 0001, Tao Wei 0002, Shutao Xia |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | Fairness-aware Adversarial Perturbation Towards Bias Mitigation for Deployed Deep ModelsabstractPrioritizing fairness is of central importance in artificial intelligence (AI) systems, especially for those societal applications, e.g., hiring systems should recommend applicants equally from different demographic groups, and risk assessment systems must eliminate racism in criminal justice. Existing efforts towards the ethical development of AI systems have leveraged data science to mitigate biases in the training set or introduced fairness principles into the training process. For a deployed AI system, however, it may not allow for retraining or tuning in practice. By contrast, we propose a more flexible approach, i.e., fairness-aware adversarial perturbation (FAAP), which learns to perturb input data to blind deployed models on fairness-related features, e.g., gender and ethnicity. The key advantage is that FAAP does not modify deployed models in terms of param-eters and structures. To achieve this, we design a discriminator to distinguish fairness-related attributes based on latent representations from deployed models. Meanwhile, a perturbation generator is trained against the discriminator, such that no fairness-related features could be extracted from perturbed inputs. Exhaustive experimental evaluation demonstrates the effectiveness and superior performance of the proposed FAAP. In addition, FAAP is validated on real-world commercial deployments (inaccessible to model pa-rameters), which shows the transferability of FAAP, foreseeing the potential of black-box adaptation. Zhibo Wang 0001, Xiaowei Dong, Henry Xue, Weifeng Chiu, Tao Wei 0002, Kui Ren 0001 |
CVPR | 6 |
| 2022 | DetectS ec: Evaluating the robustness of object detection models to adversarial attacksabstractDespite their tremendous success in various machine learning tasks, deep neural networks (DNNs) are inherently vulnerable to adversarial examples, which are maliciously crafted inputs to cause DNNs to misbehave. Intensive research has been conducted on this phenomenon in simple tasks (e.g., image classification). However, little is known about this adversarial vulnerability for object detection, a much more complicated task, which often requires specialized DNNs and multiple additional components. In this paper, we present DetectSec, a uniform platform for robustness analysis of object detection models. Currently, DetectSec implements 13 representative adversarial attacks with 7 utility metrics and 13 defenses on 18 standard object detection models. Leveraging DetectSec, we conduct the first rigorous evaluation of adversarial attacks on the state-of-the-art object detection models. We analyze the impact of the factors including DNN architecture and capacity on the model robustness. We show that many conclusions about adversarial attacks and defenses in image classification tasks do not transfer to object detection tasks, for example, the targeted attack is stronger than the untargeted attack for two-stage detectors. Our findings will aid future efforts in understanding and defending against adversarial attacks in complicated tasks. In addition, we compare the robustness of different detection models and discuss their relative strengths and weaknesses. The platform DetectSec will be open source as a unique facility for further research on adversarial attacks and defenses in object detection tasks. Tianyu Du, Shouling Ji, Bo Li 0026, Tao Wei 0002, Yunhan Jia, Raheem A. Beyah, Ting Wang 0006 |
Int. J. Intell. Syst. | 7 |
| 2021 | Privilege-Escalation Vulnerability Discovery for Large-scale RPC Services: Principle, Design, and DeploymentabstractRPCs are fundamental to our large-scale distributed system. From a security perspective, the blast radius of RPCs is worryingly big since each RPC often interacts with tens of internal system components. Thus, discovering RPC vulnerabilities is often a top priority in the software quality assurance process for production systems. In this paper, we present the design, implementation, and deployment experiences of PAIR, a fully automated system for privilege-escalation vulnerability discovery in Ant Group's large-scale RPC system. The design of PAIR centers around the live replay design principle where the vulnerability discovery is driven by the live RPC requests collected from production, rather than relying on any engineered testing requests. This ensures that PAIR is able to provide complete coverage to our production RPC requests in a privacy-preserving manner, despite the manifest of scale (billions of daily requests), complexity (hundreds of system-services involved) and heterogeneity (RPC protocols are highly customized). However, the live replay design principle is not a panacea. We made a couple of critical design decisions (and addressed their corresponding challenges) along the way to realize the principle in production. First, to avoid inspecting the responses of user-facing RPCs (due to privacy concerns), PAIR designs a universal and privacy-preserving mechanism, via profiling the end-to-end system invocation, to represent the RPC handling logic. Second, to ensure that PAIR provides proactive defense (rather than reactive defense that is often limited by known vulnerabilities), PAIR designs an empirical vulnerability labeling mechanism to effectively identify a group of potentially insecure RPCs while safely excluding other RPCs. During the course of three-year production development, PAIR in total helped locate 133 truly insecure RPCs, from billions of requests, while maintaining a zero false negative rate per our production observations. Zhuotao Liu, Sainan Li, Qi Li 0002, Tao Wei 0002, Yu Wang 0096 |
AsiaCCS | 5 |
| 2021 | SpecTaint: Speculative Taint Analysis for Discovering Spectre Gadgets
Zhenxiao Qi, Yueqiang Cheng, Mengjia Yan 0001, Heng Yin 0001, Tao Wei 0002 |
NDSS | 7 |
| 2020 | COIN Attacks: On Insecurity of Enclave Untrusted Interfaces in SGXabstractIntel SGX is a hardware-based trusted execution environment (TEE), which enables an application to compute on confidential data in a secure enclave. SGX assumes a powerful threat model, in which only the CPU itself is trusted; anything else is untrusted, including the memory, firmware, system software, etc. An enclave interacts with its host application through an exposed, enclave-specific, (usually) bi-directional interface. This interface is the main attack surface of the enclave. The attacker can invoke the interface in any order and inputs. It is thus imperative to secure it through careful design and defensive programming. Mustakimur Khandaker, Yueqiang Cheng, Zhi Wang 0004, Tao Wei 0002 |
ASPLOS | 4 |
| 2020 | Fooling Detection Alone is Not Enough: Adversarial Attack against Multiple Object Tracking
Yunhan Jia, Yantao Lu, Junjie Shen 0001, Qi Alfred Chen, Hao Chan, Zhenyu Zhong, Tao Wei 0002 |
ICLR | 7 |
| 2020 | ZeroWall: Detecting Zero-Day Web Attacks through Encoder-Decoder Recurrent Neural NetworksabstractThe following topics are dealt with: learning (artificial intelligence); optimisation; telecommunication traffic; Internet; cloud computing; computational complexity; mobile computing; resource allocation; security of data; and telecommunication network routing. Ruming Tang, Zeyan Li 0001, Weibin Meng, Haixin Wang 0003, Qi Li 0002, Yongqian Sun, Dan Pei, Tao Wei 0002, Yanfei Xu, Yan Liu 0069 |
INFOCOM | 9 |
| 2020 | SAVIOR: Towards Bug-Driven Hybrid TestingabstractHybrid testing combines fuzz testing and concolic execution. It leverages fuzz testing to test easy-to-reach code regions and uses concolic execution to explore code blocks guarded by complex branch conditions. As a result, hybrid testing is able to reach deeper into program state space than fuzz testing or concolic execution alone. Recently, hybrid testing has seen significant advancement. However, its code coverage-centric design is inefficient in vulnerability detection. First, it blindly selects seeds for concolic execution and aims to explore new code continuously. However, as statistics show, a large portion of the explored code is often bug-free. Therefore, giving equal attention to every part of the code during hybrid testing is a non-optimal strategy. It slows down the detection of real vulnerabilities by over 43%. Second, classic hybrid testing quickly moves on after reaching a chunk of code, rather than examining the hidden defects inside. It may frequently miss subtle vulnerabilities despite that it has already explored the vulnerable code paths.We propose SAVIOR, a new hybrid testing framework pioneering a bug-driven principle. Unlike the existing hybrid testing tools, SAVIOR prioritizes the concolic execution of the seeds that are likely to uncover more vulnerabilities. Moreover, SAVIOR verifies all vulnerable program locations along the executing program path. By modeling faulty situations using SMT constraints, SAVIOR reasons the feasibility of vulnerabilities and generates concrete test cases as proofs. Our evaluation shows that the bug-driven approach outperforms mainstream automated testing techniques, including state-of-the-art hybrid testing systems driven by code coverage. On average, SAVIOR detects vulnerabilities 43.4% faster than DRILLER and 44.3% faster than QSYM, leading to the discovery of 88 and 76 more unique bugs, respectively. According to the evaluation on 11 well fuzzed benchmark programs, within the first 24 hours, SAVIOR triggers 481 UBSAN violations, among which 243 are real bugs. Yaohui Chen 0001, Jun Xu 0024, Shengjian Guo, Rundong Zhou, Tao Wei 0002, Long Lu |
SP | 7 |
| 2019 | Towards Memory Safe Enclave Programming with Rust-SGXabstractIntel Software Guard eXtension (SGX), a hardware supported trusted execution environment (TEE), is designed to protect security critical applications. However, it does not terminate traditional memory corruption vulnerabilities for the software running inside enclave, since enclave software is still developed with type unsafe languages such as C/C++. This paper presents RUST-SGX, an efficient and layered approach to exterminating memory corruption for software running inside SGX enclaves. The key idea is to enable the development of enclave programs with an efficient memory safe system language Rust with a RUST-SGX SDK by solving the key challenges of how to (1) make the SGX software memory safe and (2) meanwhile run as efficiently as with the SDK provided by Intel. We therefore propose to build RUST-SGX atop Intel SGX SDK, and tame unsafe components with formally proven memory safety. We have implemented RUST-SGX and tested with a series of benchmark programs. Our evaluation results show that RUST-SGX imposes little extra overhead (less than 5% with respect to the SGX specific features and services compared to software developed by Intel SGX SDK), and meanwhile have stronger memory safety. Huibo Wang, Pei Wang 0007, Mingshen Sun, Yiming Jing, Tao Wei 0002, Zhiqiang Lin 0001 |
CCS | 9 |
| 2019 | Field experience with obfuscating million-user iOS apps in large enterprise mobile developmentabstractSummary In recent years, mobile apps have become the infrastructure of many popular Internet services. It is now common that a mobile app serves millions of users across the globe. By examining the code of these apps, reverse engineers can learn various knowledge about the design and implementation of the apps. Real‐world cases have shown that the disclosed critical information allows malicious parties to abuse or exploit the app‐provided services for unrightful profits, leading to significant financial losses. One of the most viable mitigations against malicious reverse engineering is to obfuscate the apps. Despite that security by obscurity is typically considered to be an unsound protection methodology, software obfuscation can indeed increase the cost of reverse engineering, thus delivering practical merits for protecting mobile apps. In this paper, we share our experience of applying obfuscation to multiple commercial iOS apps, each of which has millions of users. We discuss the necessity of adopting obfuscation for protecting modern mobile business, the challenges of software obfuscation on the iOS platform, and our efforts in overcoming these obstacles. We especially focus on factors that are unique to mobile software development that may affect the design and deployment of obfuscation techniques. We report the outcome of our obfuscation with empirical experiments. We additionally elaborate on the follow‐up case studies about how our obfuscation affected the app publication process and how we responded to the negative impacts. This experience report can benefit mobile developers, security service providers, and Apple as the administrator of the iOS ecosystem. Pei Wang 0007, Dinghao Wu, Zhaofeng Chen, Tao Wei 0002 |
Softw. Pract. Exp. | 4 |
| 2018 | Software protection on the go: a large-scale empirical study on mobile app obfuscationabstractThe prosperity of smartphone markets has raised new concerns about software security on mobile platforms, leading to a growing demand for effective software obfuscation techniques. Due to various differences between the mobile and desktop ecosystems, obfuscation faces both technical and non-technical challenges when applied to mobile software. Although there have been quite a few software security solution providers launching their mobile app obfuscation services, it is yet unclear how real-world mobile developers perform obfuscation as part of their software engineering practices. Pei Wang 0007, Qinkun Bao, Shuai Wang 0011, Zhaofeng Chen, Tao Wei 0002, Dinghao Wu |
ICSE | 6 |
| 2017 | POSTER: Rust SGX SDK: Towards Memory Safety in Intel SGX EnclaveabstractIntel SGX is the next-generation trusted computing infrastructure. It can e effctively protect data inside enclaves from being stolen. Similar to traditional programs, SGX enclaves are likely to have security vulnerabilities and can be exploited as well. This gives an adversary a great opportunity to steal secret data or perform other malicious operations. Rust is one of the system programming languages with promising security properties. It has powerful checkers and guarantees memory-safety and thread-safety. In this paper, we show Rust SGX SDK, which combines Intel SGX and Rust programming language together. By using Rust SGX SDK, developers could write memory-safe secure enclaves easily, eliminating the most possibility of being pwned through memory vulnerabilities. What's more, the Rust enclaves are able to run as fast as the ones written in C/C++. Yueqiang Cheng, Tanghui Chen, Tao Wei 0002, Huibo Wang |
CCS | 7 |
| 2017 | Adaptive Android Kernel Live Patching
Zhi Wang 0004, Liangzhao Xia, Chenfu Bao, Tao Wei 0002 |
USENIX Security Symposium | 6 |
| 2017 | Phishing Website Detection Based on Effective CSS Features of Web Pages
Wenqian Tian, Tao Wei 0002, Zhenkai Liang |
WASA | 4 |
| 2017 | Accurate and efficient exploit capture and classification
Tao Wei 0002, Hui Xue 0003, Chao Zhang 0008, Xinhui Han |
Sci. China Inf. Sci. | 2 |
| 2016 | TaintART: A Practical Multi-level Information-Flow Tracking System for Android RunTimeabstractMobile operating systems like Android failed to provide sufficient protection on personal data, and privacy leakage becomes a major concern. To understand the security risks and privacy leakage, analysts have to carry out data-flow analysis. In 2014, Android upgraded with a fundamentally new design known as Android RunTime (ART) environment in Android 5.0. ART adopts ahead-of-time compilation strategy and replaces previous virtual-machine-based Dalvik. Unfortunately, many data-flow analysis systems like TaintDroid were designed for the legacy Dalvik environment. This makes data-flow analysis of new apps and malware infeasible. We design a multi-level information-flow tracking system for the new Android system called TaintART. TaintART employs a multi-level taint analysis technique to minimize the taint tag storage. Therefore, taint tags can be stored in processor registers to provide efficient taint propagation operations. We also customize the ART compiler to maximize performance gains of the ahead-of-time compilation optimizations. Based on the general design of TaintART, we also implement a multi-level privacy enforcement to prevent sensitive data leakage. We demonstrate that TaintART only incurs less than 15% overheads on a CPU-bound microbenchmark and negligible overhead on built-in or third-party applications. Compared to legacy Dalvik environment in Android 4.4, TaintART achieves about 99.7% faster performance for Java runtime benchmark. Mingshen Sun, Tao Wei 0002, John C. S. Lui |
CCS | 2 |
| 2015 | Enpublic Apps: Security Threats Using iOS Enterprise and Developer CertificatesabstractCompared with Android, the conventional wisdom is that iOS is more secure. However, both jailbroken and non-jailbroken iOS devices have number of vulnerabilities. For iOS, apps need to interact with the underlying system using Application Programming Interfaces (APIs). Some of these APIs remain undocumented and Apple forbids apps in App Store from using them. These APIs, also known as "private APIs", provide powerful features to developers and yet they may have serious security consequences if misused. Furthermore, apps which use private APIs can bypass the App Store and use the "Apple's Enterprise/Developer Certificates" for distribution. This poses a significant threat to the iOS ecosystem. So far, there is no formal study to understand these apps and how private APIs are being encapsulated. We call these iOS apps which distribute to the public using enterprise certificates as "enpublic" apps. In this paper, we present the design and implementation of iAnalytics, which can automatically analyze "enpublic" apps' private API usages and vulnerabilities. Using iAnalytics, we crawled and analyzed 1,408 enpublic iOS apps. We discovered that: 844 (60%) out of the 1408 apps do use private APIs, 14 (1%) apps contain URL scheme vulnerabilities, 901 (64%) enpublic apps transport sensitive information through unencrypted channel or store the information in plaintext on the phone. In addition, we summarized 25 private APIs which are crucial and security sensitive on iOS 6/7/8, and we have filed one CVE (Common Vulnerabilities and Exposures) for iOS devices. Hui Xue 0003, Tao Wei 0002, John C. S. Lui |
AsiaCCS | 4 |
| 2015 | Towards Discovering and Understanding Task Hijacking in Android
Chuangang Ren, Hui Xue 0003, Tao Wei 0002, Peng Liu 0005 |
USENIX Security Symposium | 4 |
| 2015 | SF-DRDoS: The store-and-flood distributed reflective denial of service attack
Bingshuang Liu, Jun Li 0001, Tao Wei 0002, Skyler Berg, Jiayi Ye, Chao Zhang 0008, Xinhui Han |
Comput. Commun. | 3 |
| 2015 | Improving lookup reliability in Kad
Bingshuang Liu, Tao Wei 0002, Chao Zhang 0008, Jun Li 0001 |
Peer-to-Peer Netw. Appl. | 2 |
| 2014 | Splider: A split-based crawler of the BT-DHT network and its applicationsabstractCapturing accurate snapshots of peer-to-peer (P2P) networks, especially those with millions of users, is essential to many P2P-based applications, including those monitoring and analyzing P2P networks. The large scale and dynamic nature of P2P networks, however, make this task very challenging. Existent crawlers of P2P networks, for example, often miss a substantial portion of the ID space while unnecessarily crawling numerous nodes repeatedly. In this paper, we design and evaluate a new crawler called Splider. Unlike traditional crawling algorithms that adopt an iterative approach, Splider recursively splits the ID space of P2P nodes to crawl even tiny corners of the ID space, while avoiding crawling repeated nodes. We further implement a Splider prototype for BT-DHT, a Kademlia-based distributed hash table (DHT) P2P network, that exploits the structure of routing tables at BT-DHT nodes. Experiments show that Splider is able to gather more than 16 million nodes with a 100% recall ratio, whereas a traditional iterative crawler can at best capture only about 8 million nodes with a 66% recall ratio while its traffic-cost effectiveness is 50% less than Splider. Splider can further support distributed deployment; without any synchronization overhead, it reduces the time of capturing a full snapshot to be only about 3 minutes. We finally report and analyze the captured BT-DHT snapshots, including the spatial and temporal distribution of BT-DHT nodes and the existence of sybil and eclipse attacks in BT-DHT. Bingshuang Liu, Shidong Wu, Tao Wei 0002, Chao Zhang 0008, Jun Li 0001 |
CCNC | 3 |
| 2014 | The store-and-flood distributed reflective denial of service attackabstractDistributed reflective denial of service (DRDoS) attacks, especially those based on UDP reflection and amplification, can generate hundreds of gigabits per second of attack traffic, and have become a significant threat to Internet security. In this paper we show that an attacker can further make the DRDoS attack more dangerous. In particular, we describe a new DRDoS attack called store-and-flood DRDoS, or SF-DRDoS. By leveraging peer-to-peer (P2P) file-sharing networks, SF-DRDoS becomes more surreptitious and powerful than traditional DRDoS. An attacker can store carefully prepared data on reflector nodes before the flooding phase to greatly increase the amplification factor of an attack. We implemented a prototype of SF-DRDoS on Kad, a popular Kademlia-based P2P file-sharing network. With real-world experiments, this attack achieved an amplification factor of 2400 on average, with the upper bound of attack bandwidth at 670 Gbps in Kad. Finally, we discuss possible defenses to mitigate the threat of SF-DRDoS. Bingshuang Liu, Skyler Berg, Jun Li 0001, Tao Wei 0002, Chao Zhang 0008, Xinhui Han |
ICCCN | 4 |
| 2014 | A Website Credibility Assessment Scheme Based on Page Association
Ruilong Wang, Tao Wei 0002 |
ISPEC | 5 |
| 2014 | Drawbridge: software-defined DDoS-resistant traffic engineeringabstractEnd hosts in today's Internet have the best knowledge of the type of traffic they should receive, but they play no active role in traffic engineering. Traffic engineering is conducted by ISPs, which unfortunately are blind to specific user needs. End hosts are therefore subject to unwanted traffic, particularly from Distributed Denial of Service (DDoS) attacks. This research proposes a new system called DrawBridge to address this traffic engineering dilemma. By realizing the potential of software-defined networking (SDN), in this research we investigate a solution that enables end hosts to use their knowledge of desired traffic to improve traffic engineering during DDoS attacks. Jun Li 0001, Skyler Berg, Mingwei Zhang 0004, Peter L. Reiher, Tao Wei 0002 |
SIGCOMM | 5 |
| 2013 | Protecting function pointers in binaryabstractFunction pointers have recently become an important attack vector for control-flow hijacking attacks. However, no protection mechanisms for function pointers have yet seen wide adoption. Methods proposed in the literature have high overheads, are not compatible with existing development process, or both. In this paper, we investigate several protection methods and propose a new method called FPGate (i.e., Function Pointer Gate). FPGate rewrites x86 binary executables and implements a novel method to overcome compatibility issues. All these protection methods are then evaluated and compared from the perspectives of performance and ease of deployment. Experiments show that FPGate achieves a good balance between performance, robustness and compatibility. Chao Zhang 0008, Tao Wei 0002, Zhaofeng Chen, Lei Duan, Stephen McCamant, Laszlo Szekeres |
AsiaCCS | 2 |
| 2013 | Rating Web Pages Using Page-Transition Evidence
Xinshu Dong, Tao Wei 0002, Zhenkai Liang |
ICICS | 4 |
| 2013 | SoK: Eternal War in MemoryabstractMemory corruption bugs in software written in low-level languages like C or C++ are one of the oldest problems in computer security. The lack of safety in these languages allows attackers to alter the program's behavior or take full control over it by hijacking its control flow. This problem has existed for more than 30 years and a vast number of potential solutions have been proposed, yet memory corruption attacks continue to pose a serious threat. Real world exploits show that all currently deployed protections can be defeated. This paper sheds light on the primary reasons for this by describing attacks that succeed on today's systems. We systematize the current knowledge about various protection techniques by setting up a general model for memory corruption attacks. Using this model we show what policies can stop which attacks. The model identifies weaknesses of currently deployed techniques, as well as other proposed protections enforcing stricter policies. We analyze the reasons why protection mechanisms implementing stricter polices are not deployed. To achieve wide adoption, protection mechanisms must support a multitude of features and must satisfy a host of requirements. Especially important is performance, as experience shows that only solutions whose overhead is in reasonable bounds get deployed. A comparison of different enforceable policies helps designers of new protection mechanisms in finding the balance between effectiveness (security) and efficiency. We identify some open research problems, and provide suggestions on improving the adoption of newer techniques. Laszlo Szekeres, Mathias Payer, Tao Wei 0002, Dawn Song |
IEEE Symposium on Security and Privacy | 3 |
| 2013 | Practical Control Flow Integrity and Randomization for Binary ExecutablesabstractControl Flow Integrity (CFI) provides a strong protection against modern control-flow hijacking attacks. However, performance and compatibility issues limit its adoption. We propose a new practical and realistic protection method called CCFIR (Compact Control Flow Integrity and Randomization), which addresses the main barriers to CFI adoption. CCFIR collects all legal targets of indirect control-transfer instructions, puts them into a dedicated "Springboard section" in a random order, and then limits indirect transfers to flow only to them. Using the Springboard section for targets, CCFIR can validate a target more simply and faster than traditional CFI, and provide support for on-site target-randomization as well as better compatibility. Based on these approaches, CCFIR can stop control-flow hijacking attacks including ROP and return-into-libc. Results show that ROP gadgets are all eliminated. We observe that with the wide deployment of ASLR, Windows/x86 PE executables contain enough information in relocation tables which CCFIR can use to find all legal instructions and jump targets reliably, without source code or symbol information. We evaluate our prototype implementation on common web browsers and the SPEC CPU2000 suite: CCFIR protects large applications such as GCC and Firefox completely automatically, and has low performance overhead of about 3.6%/8.6% (average/max) using SPECint2000. Experiments on real-world exploits also show that CCFIR-hardened versions of IE6, Firefox 3.6 and other applications are protected effectively. Chao Zhang 0008, Tao Wei 0002, Zhaofeng Chen, Lei Duan, Laszlo Szekeres, Stephen McCamant, Dawn Song |
IEEE Symposium on Security and Privacy | 2 |
| 2012 | Revisiting why Kad lookup failsabstractKad is one of the most popular peer-to-peer (P2P) networks deployed on today's Internet. Its reliability is dependent on not only to the usability of the file-sharing service, but also to the capability to support other Internet services. However, Kad can only attain around a 91% lookup success ratio today. We build a measurement system called Anthill to analyze Kad's performance quantitatively, and find that Kad's failures can be classified into four types: packet loss, selective Denial of Service (sDoS) nodes, search sequence miss, and publish/search space miss. The first two are due to environment changes, the third is caused by the detachment of routing and content operations in Kad, and the last one shows the limitations of the Kademlia DHT algorithm under Kad's current configuration. Based on the analysis, we propose corresponding approaches for Kad, which achieve a success ratio of 99.8%, with only moderate communication overhead. Bingshuang Liu, Tao Wei 0002, Jun Li 0001 |
P2P | 2 |
| 2012 | A Framework to Eliminate Backdoors from Response-Computable AuthenticationabstractResponse-computable authentication (RCA) is a two-party authentication model widely adopted by authentication systems, where an authentication system independently computes the expected user response and authenticates a user if the actual user response matches the expected value. Such authentication systems have long been threatened by malicious developers who can plant backdoors to bypass normal authentication, which is often seen in insider-related incidents. A malicious developer can plant backdoors by hiding logic in source code, by planting delicate vulnerabilities, or even by using weak cryptographic algorithms. Because of the common usage of cryptographic techniques and code protection in authentication modules, it is very difficult to detect and eliminate backdoors from login systems. In this paper, we propose a framework for RCA systems to ensure that the authentication process is not affected by backdoors. Our approach decomposes the authentication module into components. Components with simple logic are verified by code analysis for correctness, components with cryptographic/ obfuscated logic are sand boxed and verified through testing. The key component of our approach is NaPu, a native sandbox to ensure pure functions, which protects the complex and backdoor-prone part of a login module. We also use a testing-based process to either detect backdoors in the sand boxed component or verify that the component has no backdoors that can be used practically. We demonstrated the effectiveness of our approach in real-world applications by porting and verifying several popular login modules into this framework. Shuaifu Dai, Tao Wei 0002, Chao Zhang 0008, Tielei Wang, Zhenkai Liang |
IEEE Symposium on Security and Privacy | 2 |
| 2011 | Using type analysis in compiler to mitigate integer-overflow-to-buffer-overflow threatabstractOne of the top two causes of software vulnerabilities in operating systems is the integer overflow. A typical integer overflow vulnerability is the Integer Overflow to Buffer Overflow (IO2BO for short) vulnerability. IO2BO is an underestimated threat. Many programmers have not realized the existenc e of IO2BO and its harm. Even for those who are aware of IO2BO, locating and fixing IO2BO vulnerabilities are still tedious and error-prone. Automatically identifying and fixing this kind of vulnerability are critical for software security. In this article, we present the design and implementation of IntPatch, a compiler extension for automatically fixing IO2BO vulnerabilities in C/C++ programs at compile time. IntPatch utilizes classic type theory and a dataflow analysis framework to identify potential IO2BO vulnerabilities, and then uses backward slicing to find out related vulnerable arithmetic operations, and finally instruments programs with runtime checks. Moreover, IntPatch provides an interface for programmers who want to check integer overflows manually. We evaluated IntPatch on a few real-world applications. It caught all 46 previously known IO2BO vulnerabilities in our test suite and found 21 new bugs. Applications patched by IntPatch have negligible runtime performance losses which are on average 1%. Chao Zhang 0008, Tielei Wang, Tao Wei 0002 |
J. Comput. Secur. | 5 |
| 2011 | Checksum-Aware Fuzzing Combined with Dynamic Taint Analysis and Symbolic ExecutionabstractFuzz testing has proven successful in finding security vulnerabilities in large programs. However, traditional fuzz testing tools have a well-known common drawback: they are ineffective if most generated inputs are rejected at the early stage of program running, especially when target programs employ checksum mechanisms to verify the integrity of inputs. This article presents TaintScope, an automatic fuzzing system using dynamic taint analysis and symbolic execution techniques, to tackle the above problem. TaintScope has several novel features: (1) TaintScope is a checksum-aware fuzzing tool. It can identify checksum fields in inputs, accurately locate checksum-based integrity checks by using branch profiling techniques, and bypass such checks via control flow alteration. Furthermore, it can fix checksum values in generated inputs using combined concrete and symbolic execution techniques. (2) TaintScope is a taint-based fuzzing tool working at the x86 binary level. Based on fine-grained dynamic taint tracing, TaintScope identifies the “hot bytes” in a well-formed input that are used in security-sensitive operations (e.g., invoking system/library calls), and then focuses on modifying such bytes with random or boundary values. (3) TaintScope is also a symbolic-execution-based fuzzing tool. It can symbolically evaluate a trace, reason about all possible values that can execute the trace, and then detect potential vulnerabilities on the trace. We evaluate TaintScope on a number of large real-world applications. Experimental results show that TaintScope can accurately locate the checksum checks in programs and dramatically improve the effectiveness of fuzz testing. TaintScope has already found 30 previously unknown vulnerabilities in several widely used applications, including Adobe Acrobat, Flash Player, Google Picasa, and Microsoft Paint. Most of these severe vulnerabilities have been confirmed by Secunia and oCERT, and assigned CVE identifiers (such as CVE-2009-1882, CVE-2009-2688). Vendor patches have been released or are in preparation based on our reports. Tielei Wang, Tao Wei 0002, Guofei Gu |
ACM Trans. Inf. Syst. Secur. | 2 |
| 2010 | Heap Taichi: exploiting memory allocation granularity in heap-spraying attacksabstractHeap spraying is an attack technique commonly used in hijacking browsers to download and execute malicious code. In this attack, attackers first fill a large portion of the victim process's heap with malicious code. Then they exploit a vulnerability to redirect the victim process's control to attackers' code on the heap. Because the location of the injected code is not exactly predictable, traditional heap-spraying attacks need to inject a huge amount of executable code to increase the chance of success. Injected executable code usually includes lots of NOP-like instructions leading to attackers' shellcode. Targeting this attack characteristic, previous solutions detect heap-spraying attacks by searching for the existence of such large amount of NOP sled and other shellcode. Tao Wei 0002, Tielei Wang, Zhenkai Liang |
ACSAC | 2 |
| 2010 | Secure dynamic code generation against sprayingabstractDCG (Dynamic Code Generation) technologies have found widely applications in the Web 2.0 era, Dion Blazakis recently presented a Flash JIT-Spraying attack against Adobe Flash Player that easily circumvented DEP and ASLR protection mechanisms built in modern operating systems. We have generalized and extended JIT Spraying into DCG Spraying. Based our analyses on this abstract model of DCG Spraying, we have found that all mainstream DCG implementations (Java/ JavaScript/ Flash/ .Net/ SilverLight) are vulnerable against DCG Spraying attack, and none of the existing ad hoc defenses such as compilation optimization, random NOP padding and constant splitting provides effective protection. Furthermore, we propose a new protection method, INSeRT, which combines randomization of intrinsic elements of machine instructions and randomly planted special trapping snippets. INSeRT practically renders the "sprayed code" ineffective, while alerts the host program of ongoing attacking attempts. We implemented a prototype of INSeRT on the V8 JavaScript engine, and the performance overhead is less than 5%, which should be acceptable in practical application. Tao Wei 0002, Tielei Wang, Lei Duan |
CCS | 1 |
| 2010 | IntPatch: Automatically Fix Integer-Overflow-to-Buffer-Overflow Vulnerability at Compile-Time
Chao Zhang 0008, Tielei Wang, Tao Wei 0002 |
ESORICS | 3 |
| 2010 | TaintScope: A Checksum-Aware Directed Fuzzing Tool for Automatic Software Vulnerability DetectionabstractFuzz testing has proven successful in finding security vulnerabilities in large programs. However, traditional fuzz testing tools have a well-known common drawback: they are ineffective if most generated malformed inputs are rejected in the early stage of program running, especially when target programs employ checksum mechanisms to verify the integrity of inputs. In this paper, we present TaintScope, an automatic fuzzing system using dynamic taint analysis and symbolic execution techniques, to tackle the above problem. TaintScope has several novel contributions: 1) TaintScope is the first checksum-aware fuzzing tool to the best of our knowledge. It can identify checksum fields in input instances, accurately locate checksum-based integrity checks by using branch profiling techniques, and bypass such checks via control flow alteration. 2) TaintScope is a directed fuzzing tool working at X86 binary level (on both Linux and Window). Based on fine-grained dynamic taint tracing, TaintScope identifies which bytes in a well-formed input are used in security-sensitive operations (e.g., invoking system/library calls) and then focuses on modifying such bytes. Thus, generated inputs are more likely to trigger potential vulnerabilities. 3) TaintScope is fully automatic, from detecting checksum, directed fuzzing, to repairing crashed samples. It can fix checksum values in generated inputs using combined concrete and symbolic execution techniques. We evaluate TaintScope on a number of large real-world applications. Experimental results show that TaintScope can accurately locate the checksum checks in programs and dramatically improve the effectiveness of fuzz testing. TaintScope has already found 27 previously unknown vulnerabilities in several widely used applications, including Adobe Acrobat, Google Picasa, Microsoft Paint, and ImageMagick. Most of these severe vulnerabilities have been confirmed by Secunia and oCERT, and assigned CVE identifiers (such as CVE-2009-1882, CVE-2009-2688). Corresponding patches from vendors are released or in progress based on our reports. Tielei Wang, Tao Wei 0002, Guofei Gu |
IEEE Symposium on Security and Privacy | 2 |
| 2009 | IntScope: Automatically Detecting Integer Overflow Vulnerability in X86 Binary Using Symbolic Execution
Tielei Wang, Tao Wei 0002, Zhiqiang Lin 0001 |
NDSS | 2 |
| 2007 | Structuring 2-way Branches in Binary ExecutablesabstractOne of the major challenges of control flow analysis in decompilation is to structure 2-way branches into conditionals, loop conditionals and switches. In this paper, we propose a graph-based method to formally describe structures of 2-way branches via the introduction of concepts called "compound branch subgraph" and "cascade branch subgraph". We then present novel structuring algorithms based on such concepts. Compared with previous works, our algorithms are deterministic rather than heuristic, and they do not use complicated data structures such as Interval/DSG. We show that in theory our algorithm is more accurate and efficient than typical current approaches; furthermore, we have applied the algorithm to several real-world binary executables, and experimental results validate such theoretical analysis. Tao Wei 0002 |
COMPSAC (1) | 1 |
| 2007 | A New Algorithm for Identifying Loops in Decompilation
Tao Wei 0002 |
SAS | 1 |
| 2006 | Linkability of a Blind Signature Scheme and Its Improved Scheme
Tao Wei 0002 |
ICCSA (4) | 2 |