EDBT 2026 Demo / reviewers in the wild / expert
Wenhao Wang 0001
dblp:57/9813-1
· DBLP profile ↗
36ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0001-7294-2724ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 26 · 5 first-author · 16 since 2021Systems, architecture and hardware · 9 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CryptPEFT: Efficient and Private Neural Network Inference via Parameter-Efficient Fine-Tuning
Saisai Xia, Wenhao Wang 0001, Yuhui Zhang 0011, Yier Jin, Dan Meng 0002, Rui Hou 0001 |
NDSS | 2 |
| 2025 | The Road to Trust: Building Enclaves within Confidential VMs
Wenhao Wang 0001, Linke Song, Benshan Mei, Shijun Zhao, Shoumeng Yan, XiaoFeng Wang 0001, Dan Meng 0002, Rui Hou 0001 |
NDSS | 1 |
| 2025 | Comet: Accelerating Private Inference for Large Language Model by Predicting Activation SparsityabstractWith the growing use of large language models (LLMs) hosted on cloud platforms to offer inference services, privacy concerns about the potential leakage of sensitive information are escalating. Secure Multi-Party Computation (MPC) is a promising solution to protect the privacy in LLM inference. However, MPC requires frequent inter-server communication, causing high performance overhead. Inspired by the prevalent activation sparsity of LLMs, where most neuron are not activated after non-linear activation functions, we propose an efficient private inference system, Comet. This system employs an accurate and fast predictor to predict the sparsity distribution of activation function output. Additionally, we introduce a new private inference protocol. It efficiently and securely avoids computations involving zero values by exploiting the spatial locality of the predicted sparsity distribution. While this computation-avoidance approach impacts the spatiotemporal continuity of KV cache entries, we address this challenge with a low-communication overhead cache refilling strategy that merges miss requests and incorporates a prefetching mechanism. Finally, we evaluate Comet on four common LLMs and compare it with six state-of-the-art private inference systems. Comet achieves a$1.87\times-2.63\times$speedup and a$1.94\times-2.64\times$communication reduction. Guang Yan, Yuhui Zhang 0011, Zimu Guo, Lutan Zhao, Xiaojun Chen 0004, Wenhao Wang 0001, Dan Meng 0002, Rui Hou 0001 |
SP | 7 |
| 2025 | VMPL-KMI: Protecting Kernel Module Integrity within Confidential VMsabstractConfidential Virtual Machines (CVMs), such as AMD SEV, offer external protection but lack a privilege hierarchy, making them vulnerable to susceptible loadable kernel modules (LKMs). Although the Integrity Measurement Architecture (IMA) checks load-time integrity, it is prone to kernel exploits. While AMD SEV-SNP’s VM Privilege Levels (VMPLs) offer hardware-enforced intra-CVM isolation, they remain underexploited for kernel protection. This paper presents VMPL-KMI, utilizing VMPLs to enforce both load-time measurement and runtime protection of LKMs. By maintaining MList in the most privileged VMPL0 and synchronizing integrity checks with memory protection, VMPL-KMI prevents tampering and eliminates Time-Of-Check-to-Time-Of-Use (TOCTTOU) vulnerabilities. To minimize the overhead, we introduce a service protocol based on the Secure VM Service Module (SVSM) standard, reducing verification to a single interaction. Experimental results demonstrate practical efficiency: 5–10% overhead for small modules (over 20% for larger ones) during loading and less than 5% during unloading. Benshan Mei, Wenhao Wang 0001, Dongdai Lin |
TrustCom | 2 |
| 2025 | WhistleBlower: A System-Level Empirical Study on RowHammerabstractWith frequent software-induced activations on DRAM rows, bit flips can occur on their physically adjacent rows (i.e., RowHammer). Existing studies leverage FPGA platforms to characterize RowHammer, which have identified key factors that contribute to RowHammer bit flips, e.g., data pattern. As the FPGA-based studies have removed the interference of the OS and the memory controller, their findings on the identified contributing factors do not always work as reported in a real-world computing system, resulting in negative effects on system-level RowHammer attacks and defenses. In this paper, we carry out a system-level empirical study on factors from both the software side and the DRAM side that contribute to RowHammer. We conduct the study on 33 DRAM modules including both DDR4 and DDR3, with 292 DRAM chips from various vendors. Our experimental results from the software side show that some prior findings about existing factors are inconsistent with our observations, thus not applicable to a real-world system. Also, we contribute to identifying one new factor that effectively affects RowHammer bit flips. Our DRAM-side results identify three types of new contributing factors and indicate that DRAM modules are more vulnerable if they achieve better performance and lower power consumption. Particularly, Intel XMP, intended for improving DRAM performance, might be abused for RowHammer attacks. Zhi Zhang 0001, Yueqiang Cheng, Wenhao Wang 0001, Wei Song 0002, Yansong Gao 0001, Qifei Zhang 0001, Dongxi Liu, Surya Nepal |
IEEE Trans. Computers | 4 |
| 2025 | A Bidirectional Differential Evolution-Based Unknown Cyberattack Detection SystemabstractThe evolving unknown cyberattacks, compounded by the widespread emerging technologies (say 5G, Internet of Things, etc.), have rapidly expanded the cyber threat landscape. However, most existing intrusion detection systems (IDSs) are effective in detecting only known cyberattacks, because only known cyberattack samples are usually available for IDS training. Identifying unknown cyberattacks, therefore, remains a big challenging issue. To meet this gap, in this paper, motivated by artificial immunity (AIm) and differential evolution (DE), we propose a bidirectional differential evolution based unknown cyberattack detection system, coined BDE-IDS. Specifically, we first design a bidirectional differential evolution algorithm for known nonself antigens (abnormal data), where bidirectional evolutionary directions are considered for increasing or decreasing the differences between known nonself antigens and self antigens (normal data), to create new antigens possibly used for generating cyberattack detectors. Second, a novel tolerance training mechanism is developed to eliminate invalid newly-evolved antigens falling into the coverage of either known self or nonself antigens. Third, the remaining antigens are employed to generate detectors for unknown cyberattacks. Extensive experiments demonstrate that the proposed BDE-IDS achieves outperformance in detecting unknown cyberattacks (as well as known cyberattacks) compared to state-of-the-art studies, including those AIm-based, signature-based, and anomaly-based IDSs. Hanyuan Huang, Tao Li 0016, Beibei Li 0002, Wenhao Wang 0001, Yanan Sun 0001 |
IEEE Trans. Evol. Comput. | 4 |
| 2025 | Multivariate Template Attack Against NTT-Based Polynomial Multiplication of Dilithium
Haopeng Fan, Hailong Zhang 0001, Yongjuan Wang, Wenhao Wang 0001, Haojin Zhang, Qingjun Yuan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | The Early Bird Catches the Leak: Unveiling Timing Side Channels in LLM Serving SystemsabstractThe wide deployment of Large Language Models (LLMs) has given rise to strong demands for optimizing their inference performance. Today’s techniques serving this purpose primarily focus on reducing latency and improving throughput through algorithmic and hardware enhancements, while largely overlooking their privacy side effects, particularly in a multi-user environment. In our research, for the first time, we discovered a set of new timing side channels in LLM systems, arising from shared caches and GPU memory allocations, which can be exploited to infer both confidential system prompts and those issued by other users. These vulnerabilities echo security challenges observed in traditional computing systems, highlighting an urgent need to address potential information leakage in LLM serving infrastructures. In this paper, we report novel attack strategies designed to exploit such timing side channels inherent in LLM deployments, specifically targeting the Key-Value (KV) cache and semantic cache widely used to enhance LLM inference performance. Our approach leverages timing measurements and classification models to detect cache hits, allowing an adversary to infer private prompts with high accuracy. We also propose a token-by-token search algorithm to efficiently recover shared prompt prefixes in the caches, showing the feasibility of stealing system prompts and those produced by peer users. Our experimental studies on black-box testing of popular online LLM services demonstrate that such privacy risks are completely realistic, with significant consequences. Our findings underscore the need for robust mitigation to protect LLM systems against such emerging threats. Linke Song, Zixuan Pang, Wenhao Wang 0001, XiaoFeng Wang 0001, Wei Song 0002, Yier Jin, Dan Meng 0002, Rui Hou 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Verifying Rust Implementation of Page Tables in a Software Enclave HypervisorabstractAs trusted execution environments (TEE) have become the corner stone for secure cloud computing, it is critical that they are reliable and enforce proper isolation, of which a key ingredient is spatial isolation. Many TEEs are implemented in software such as hypervisors for flexibility, and in a memory-safe language, namely Rust to alleviate potential memory bugs. Still, even if memory bugs are absent from the TEE, it may contain semantic errors such as mis-configurations in its memory subsystem which breaks spatial isolation. Zhenyang Dai, Vilhelm Sjöberg, Xupeng Li, Yu Chen 0004, Wenhao Wang 0001, Yuekai Jia, Sean Noble Anderson, Laila Elbeheiry, Shubham Sondhi, Yu Zhang 0313, Zhaozhong Ni, Shoumeng Yan, Ronghui Gu, Zhengyu He |
ASPLOS (2) | 6 |
| 2024 | ThermalScope: A Practical Interrupt Side Channel Attack Based on Thermal Event InterruptsabstractWhile interrupts play a critical role in modern OSes, they have been exploited as a wide range of side channel attacks to break system confidentiality, such as keystroke interrupts, graphic interrupts and network interrupts. In this paper, we propose ThermalScope, a new side channel that exploits thermal event interrupts, which is adaptable for both native and browser scenarios and incorporates two heat amplifying techniques. The thermal event interrupts are activated only when the CPU package temperature reaches a fixed threshold that is determined by manufacturers. Our key observation is that workloads running on CPUs inevitably generates their distinct heat, which can be correlated with the thermal event interrupts. To demonstrate the viability of ThermalScope, we conduct a comprehensive evaluation on multiple Ubuntu OSes with different Intel-based CPUs. First, we show that the activation of thermal event interrupts correlates with the level of CPU temperature. We then apply ThermalScope to mount different side channel attacks, i.e., building covert channels with a transmission rate of 0.1 b/s, fingerprinting DNN model architectures with an accuracy of over 90% and breaking KASLR within 8.2 hours. Xin Zhang 0110, Zhi Zhang 0001, Qingni Shen, Wenhao Wang 0001, Yansong Gao 0001, Zhuoxi Yang, Zhonghai Wu |
DAC | 4 |
| 2024 | SegScope: Probing Fine-grained Interrupts via Architectural FootprintsabstractInterrupts are critical hardware resources for OS kernels to schedule processes. As they are related to system activities, interrupts can be used to mount various side-channel attacks (i.e., monitoring keystrokes, inferring website visits, detecting GPU activities, and fingerprinting processes). Given that all these attacks rely on system file interfaces or architectural timers to probe interrupts, various countermeasures have been proposed to either remove the unprivileged access to the file interfaces or detect/cripple architectural timers. In this work, we propose SegScope, a new technique that abuses segment protection to provision fine-grained interrupt observations without any timer. As segment protection is widely used on x86, SegScope works across a wide range of Intel-and AMD-based CPUs. Particularly, we observe that while segment protection preserves the confidentiality of high privileged domain, it leaves a footprint via the data segment registers values when an interrupt occurs. With this key observation, SegScope is crafted by capturing the footprints. To show its security implications, we evaluate it in four case studies. First, SegScope has inferred website visits with a respective success rate of 92.4% on Chrome and 87.4% on Tor Browser in default system settings. Second, SegScope successfully extracts the keys from Cloudflare's Interoperable Reusable Cryptographic Library (CIRCL) vl.l. Third, SegScope steals DNN model architectures with an accuracy of over 80%. Last, SegScope effectively reduces the noise of interrupts to improve the performance of other side channels. As an example, SegScope reduces the error rate of Spectral side channel by 56×. Compared with existing timer-based interrupt-probing techniques, SegScope is fine-grained without introducing false-positives. Further, we leverage SegScope to craft a fine-grained timer, as regular timer interrupts as clock edges contain timestamps. Our evaluation shows that it achieves the same level of timing granularity as the high-resolution timer, i.e., rdtsc and rdpru. We then leverage the timer to break KASLR in about 10 seconds and mount a Flush+Reload based Spectre attack. Xin Zhang 0110, Zhi Zhang 0001, Qingni Shen, Wenhao Wang 0001, Yansong Gao 0001, Zhuoxi Yang, Jiliang Zhang 0002 |
HPCA | 4 |
| 2024 | Cabin: Confining Untrusted Programs Within Confidential VMs
Benshan Mei, Saisai Xia, Wenhao Wang 0001, Dongdai Lin |
ICICS (1) | 3 |
| 2024 | Tossing in the Dark: Practical Bit-Flipping on Gray-box Deep Neural Networks for Runtime Trojan Injection
Di Tang 0001, XiaoFeng Wang 0001, Zhaoyang Geng, Wenhao Wang 0001 |
USENIX Security Symposium | 6 |
| 2024 | Screening Least Square Technique Assisted Multivariate Template Attack Against the Random Polynomial Generation of DilithiumabstractIn recent years, the security of Dilithium against side-channel attacks (SCA) has attracted great attentions from the cryptographic engineering community. However, existing power analysis attacks cannot fully utilize the side-channel leakages of the Dilithium reference implementation to efficiently recover the private key. In light of this, a screening least square technique assisted multivariate template attack (SLST assisted MTA) is proposed in this paper. In SLST assisted MTA, side-channel leakages of coefficient$y_{i}$of random polynomial y, unsigned number$x_{i}$and random byte string$a_{i^{\prime }}$can be utilized simultaneously to recover coefficient$y_{i}$of random polynomial y with MTA. Then, one can build error-tolerant equations, and the private key$\mathbf {s_{1}}$can be solved with SLST efficiently. We evaluate the private key recovery efficiency of SLST assisted MTA with real traces measured from the Cortex-M4 processor based Dilithium reference implementation, and the evaluation results show that with MTA, 19.41%, 15.70% and 16.88% of the coefficients of y can be accurately recovered in cases of Dilithium 2, 3 and 5. Besides, using SLST, after five times screening, only 38, 40 and 39 power traces are enough to recover private key$\mathbf {s_{1}}$of Dilithium 2, 3 and 5 with 100% of success rate. Haopeng Fan, Hailong Zhang 0001, Yongjuan Wang, Wenhao Wang 0001, Yanbei Zhu, Haojin Zhang, Qingjun Yuan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | The Danger of Minimum Exposures: Understanding Cross-App Information Leaks on iOS through Multi-Side-Channel LearningabstractResearch on side-channel leaks has long been focusing on the information exposure from a single channel (memory, network traffic, power, etc.). Less studied is the risk of learning from multiple side channels related to a target activity (e.g., website visits) even when individual channels are not informative enough for an effective attack. Although the prior research made the first step on this direction, inferring the operations of foreground apps on iOS from a set of global statistics, still less clear are how to determine the maximum information leaks from all target-related side channels on a system, what can be learnt about the target from such leaks and most importantly, how to control information leaks from the whole system, not just from an individual channel. To answer these fundamental questions, we performed the first systematic study on multi-channel inference, focusing on iOS as the first step. Our research is based upon a novel attack technique, called Mischief, which given a set of potential side channels related to a target activity (e.g., foreground apps), utilizes probabilistic search to approximate an optimal subset of the channels exposing most information, as measured by Merit Score, a metric for correlation-based feature selection. On such an optimal subset, an inference attack is modeled as a multivariate time series classification problem, so the state-of-the-art deep-learning based solution, InceptionTime in particular, can be applied to achieve the best possible outcome. Mischief is found to work effectively on today's iOS (16.2), identifying foreground apps, website visits, sensitive IoT operations (e.g., opening the door) with a high confidence, even in an open-world scenario, which demonstrates that the protection Apple puts in place against the known attack is inadequate. Also importantly, this new understanding enables us to develop more comprehensive protection, which could elevate today's side-channel research from suppressing leaks from individual channels to controlling information exposure across the whole system. Jiale Guan, XiaoFeng Wang 0001, Wenhao Wang 0001, Luyi Xing, Fares Fahad S. Alharbi |
CCS | 4 |
| 2023 | The Dynamic Paradox: How Layer-skipping DNNs Amplify Cache Side-Channel LeakagesabstractTraditional neural networks are static: their network structure and parameters remain fixed regardless of the input samples after the networks are trained. In contrast, dynamic neural networks are capable of dynamically adapting their structures and parameters when processing different input samples, e.g., with early exiting and layer skipping, to achieve more efficient use of resources and more effective feature representation. These capabilities make dynamic neural networks a promising area of development for deep learning, with the potential to improve both the efficiency and effectiveness of neural network models.Previous research has demonstrated that the search space of static neural network structures can be significantly reduced by analyzing cache access traces. To date, the side channel risks associated with dynamic neural networks remain unexplored. In this paper, we present the first study on the side-channel leakages of dynamic neural networks which are used to infer the network structure. We conduct both theoretical and experimental validation on a typical dynamic Layer-skipping DNNs structure, SkipNet, and show that the adaptive nature of dynamic DNNs leads to greater side-channel information leakage compared to traditional neural network structures. Specifically, our findings reveal that the network structure space of dynamic DNNs can be significantly reduced, from 36608057× 212× (264− 1)2to 11286. Jinze She, Wenhao Wang 0001 |
TrustCom | 2 |
| 2023 | Trust Beyond Border: Lightweight, Verifiable User Isolation for Protecting In-Enclave ServicesabstractDue to the absence of in-enclave isolation, today's trusted execution environment (TEE), specifically Intel's Software Guard Extensions (SGX), does not have the capability to securely run different users’ tasks within a single enclave, which is required for supporting real-world services, such as an in-enclave machine learning model that classifies the data from various sources, or a microservice (e.g., data search) that performs a very small task (within sub-seconds) for a user and therefore cannot afford the resources and the delay for creating a separate enclave for each user. To address this challenge, we developedLiveries, a technique that enables lightweight, verifiable in-enclave user isolation for protecting time-sharing services. Our approach restricts an in-enclave thread's privilege when configuring an enclave, and further performs integrity check and sanitization on critical enclave data upon user switches. For this purpose, we developed a novel technique that ensures the protection of sensitive user data (e.g., session keys) even in the presence of the adversary who may have compromised the enclave. Our study shows that the new technique is lightweight (1% overhead) and verifiable (about 3200 lines of code), making a step towards assured protection of real-world in-enclave services. Wenhao Wang 0001, Weijie Liu 0004, XiaoFeng Wang 0001, Hongliang Tian, Dongdai Lin |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2023 | Implicit Hammer: Cross-Privilege-Boundary Rowhammer Through Implicit AccessesabstractRowhammer is a hardware vulnerability in DRAM memory, where repeated access to hammer rows can induce bit flips in neighboringvictim rows. Rowhammer attacks have enabled privilege escalation, sandbox escape, cryptographic key disclosures, etc. A key requirement ofallexisting rowhammer attacks is that an attacker must have access to at least part of an exploitable hammer row. We term such rowhammer attacks as Explicit Hammer. Recently, several proposals leverage the spatial proximity between the accessed hammer rows and the location of the victim rows for a defense against rowhammer. These all aim to deny the attacker's permission to access hammer rows near sensitive data, thus defeating explicit hammer-based attacks. In this paper, we question the core assumption underlying these defenses. We present Implicit Hammer, a confused-deputy attack that causes accesses to hammer rows that the attacker is not allowed to access. It is a paradigm shift in rowhammer attacks since it crosses privilege boundary to stealthily rowhammer an inaccessible row by implicit DRAM accesses. Such accesses are achieved by abusing inherent features of modern hardware and/or software. We propose a generic model to rigorously formalize the necessary conditions to initiate implicit hammer and explicit hammer, respectively. Compared to explicit hammer, implicit hammer can defeat the advanced software-only defenses, stealthy in hiding itself and hard to be mitigated. To demonstrate the practicality of implicit hammer, we have created two implicit hammer's instances, called PThammer and SyscallHammer. Zhi Zhang 0001, Yueqiang Cheng, Wenhao Wang 0001, Yansong Gao 0001, Dongxi Liu, Surya Nepal, Anmin Fu, Yi Zou 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2022 | SoftTRR: Protect Page Tables against Rowhammer Attacks using Software-only Target Row Refresh
Zhi Zhang 0001, Yueqiang Cheng, Wenhao Wang 0001, Surya Nepal, Yansong Gao 0001, Zhe Wang 0017, Chenggang Wu 0002 |
USENIX ATC | 5 |
| 2022 | HyperEnclave: An Open and Cross-platform Trusted Execution Environment
Yuekai Jia, Wenhao Wang 0001, Yu Chen 0004, Zhengde Zhai, Shoumeng Yan, Zhengyu He |
USENIX ATC | 3 |
| 2021 | Constructive Use of Process Variations: Reconfigurable and High-Resolution Delay-LineabstractDelay-line is a critical circuit component for highspeed electronic design and testing, such as high-performance FPGA and ASICs, to provide timing signals of specific duration or duty cycle. However, the performance of existing CMOS-based delay-lines is limited by various practical issues. For example, the minimum propagation delay (resolution) of CMOS gates is limited by the process variations from circuit fabrication. This paper presents a novel delay-line scheme, which instead of mitigating the process variations from circuit fabrication, constructively leverages them to generate time signals of specific duration. Moreover, the resolution of the proposed delay-line method is reconfigurable, for which we propose a Machine Learning modeling method to assist such reconfiguration, i.e., to generate time duration of different scales. The performance of the proposed delay-line is validated with HSpice simulation and prototype on a Xilinx Virtex-6 FPGA evaluation kit. The experimental results demonstrate that the proposed delay-line method achieves an ultra-high resolution of sub-picosecond. Wenhao Wang 0001, Yukui Luo, Xiaolin Xu 0001 |
DATE | 1 |
| 2021 | Practical and Efficient in-Enclave Verification of Privacy ComplianceabstractA trusted execution environment (TEE) such as Intel Software Guard Extension (SGX) runs attestation to prove to a data owner the integrity of the initial state of an enclave, including the program to operate on her data. For this purpose, the data-processing program is supposed to be open to the owner or a trusted third party, so its functionality can be evaluated before trust being established. In the real world, however, increasingly there are application scenarios in which the program itself needs to be protected (e.g., proprietary algorithm). So its compliance with privacy policies as expected by the data owner should be verified without exposing its code. To this end, this paper presents Deflection, a new model for TEE-based delegated and flexible in-enclave code verification. Given that the conventional solutions do not work well under the resource-limited and TCB-frugal TEE, we come up with a new design inspired by Proof-Carrying Code. Our design strategically moves most of the workload to the code generator, which is responsible for producing easy-to-check code, while keeping the consumer simple. Also, the whole consumer can be made public and verified through a conventional attestation. We implemented this model on Intel SGX and demonstrate that it introduces a very small part of TCB. We also thoroughly evaluated its performance on micro- and macro- benchmarks and real-world applications, showing that the design only incurs a small overhead when enforcing several categories of security policies. Weijie Liu 0004, Wenhao Wang 0001, XiaoFeng Wang 0001, Yaosong Lu, Kai Chen 0012, Qintao Shen, Yi Chen 0024, Haixu Tang |
DSN | 2 |
| 2021 | Randomized Last-Level Caches Are Still Vulnerable to Cache Side-Channel Attacks! But We Can Fix ItabstractCache randomization has recently been revived as a promising defense against conflict-based cache side-channel attacks. As two of the latest implementations, CEASER-S and ScatterCache both claim to thwart conflict-based cache side-channel attacks using randomized skewed caches. Unfortunately, our experiments show that an attacker can easily find a usable eviction set within the chosen remap period of CEASER-S and increasing the number of partitions without dynamic remapping, such as ScatterCache, cannot eliminate the threat. By quantitatively analyzing the access patterns left by various attacks in the LLC, we have newly discovered several problems with the hypotheses and implementations of randomized caches, which are also overlooked by the research on conflict-based cache side-channel attacks.However, cache randomization is not a false hope and it is an effective defense that should be widely adopted in future processors. The newly discovered problems are corresponding to flaws associated with the existing implementation of cache randomization and are fixable. Several new defense ideas are proposed in this paper. Our experiments show that all the newly discovered problems are fixed within the current performance budget. We also argue that randomized set-associative caches can be sufficiently strengthened and possess a better chance to be actually adopted in commercial processors than their skewed counterparts because they introduce less overhaul to the existing cache structure. Wei Song 0002, Boya Li, Zihan Xue, Wenhao Wang 0001, Peng Liu 0005 |
SP | 5 |
| 2021 | BitMine: An End-to-End Tool for Detecting Rowhammer VulnerabilityabstractRowhammer is a destructive software-induced DRAM fault, which an attacker can leverage to break system security. Both individual customers and enterprise users (e.g., cloud providers) might refrain from using a computing system if it is vulnerable to rowhammer vulnerability. In this paper, we provide the first end-to-end tool, coined BitMine, that systematically assesses a DRAM chip’s vulnerability to rowhammer bit flips. BitMine is an extension of DRAMDig. As DRAM address mappings are proprietary techniques and critical in inducing rowhammer bit flips, DRAMDig, our prior work, leverages domain knowledge to efficiently and deterministically reverse-engineer DRAM address mappings on Intel machines. By incorporating DRAMDig, BitMine configures three key parameters, i.e., hammer methods, hammer patterns, data patterns, on the effectiveness of finding rowhammer bit flips. BitMine by default implements 13 hammer methods, 4 hammer patterns and 16 data patterns and is extensible to support more. We evaluate DRAMDig and BitMine against multiple machine models that combine different DRAM chips and Intel microarchitectures. Our experiment results show that DRAMDig efficiently uncovers a deterministic DRAM address mapping for each machine model, and every implemented parameter in BitMine has its distinct effectiveness in triggering bit flips for different machine models. Zhi Zhang 0001, Yueqiang Cheng, Wenhao Wang 0001, Yansong Gao 0001, Surya Nepal, Yang Xiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | A Privacy-Preserving-Oriented DNN Pruning and Mobile Acceleration FrameworkabstractWeight pruning of deep neural networks (DNNs) has been proposed to satisfy the limited storage and computing capability of mobile edge devices. However, previous pruning methods mainly focus on reducing the model size and/or improving performance without considering the privacy of user data. To mitigate this concern, we propose a privacy-preserving-oriented pruning and mobile acceleration framework that does not require the private training dataset. At the algorithm level of the proposed framework, a systematic weight pruning technique based on the alternating direction method of multipliers (ADMM) is designed to iteratively solve the pattern-based pruning problem for each layer with randomly generated synthetic data. In addition, corresponding optimizations at the compiler level are leveraged for inference accelerations on devices. With the proposed framework, users could avoid the time-consuming pruning process for non-experts and directly benefit from compressed models. Experimental results show that the proposed framework outperforms three state-of-art end-to-end DNN frameworks, i.e., TensorFlow-Lite, TVM, and MNN, with speedup up to 4.2×, 2.5×, and 2.0×, respectively, with almost no accuracy loss, while preserving data privacy. Yifan Gong 0004, Zheng Zhan 0001, Zhengang Li 0001, Wei Niu 0002, Wenhao Wang 0001, Bin Ren 0002, Caiwen Ding, Xue Lin 0001, Xiaolin Xu 0001, Yanzhi Wang 0001 |
ACM Great Lakes Symposium on VLSI | 6 |
| 2020 | Enabling Rack-scale Confidential Computing using Heterogeneous Trusted Execution EnvironmentabstractWith its huge real-world demands, large-scale confidential computing still cannot be supported by today's Trusted Execution Environment (TEE), due to the lack of scalable and effective protection of high-throughput accelerators like GPUs, FPGAs, and TPUs etc. Although attempts have been made recently to extend the CPU-like enclave to GPUs, these solutions require change to the CPU or GPU chips, may introduce new security risks due to the side-channel leaks in CPU-GPU communication and are still under the resource constraint of today's CPU TEE.To address these problems, we present the first Heterogeneous TEE design that can truly support large-scale compute or data intensive (CDI) computing, without any chip-level change. Our approach, called HETEE, is a device for centralized management of all computing units (e.g., GPUs and other accelerators) of a server rack. It is uniquely designed to work with today's data centres and clouds, leveraging modern resource pooling technologies to dynamically compartmentalize computing tasks, and enforce strong isolation and reduce TCB through hardware support. More specifically, HETEE utilizes the PCIe ExpressFabric to allocate its accelerators to the server node on the same rack for a non-sensitive CDI task, and move them back into a secure enclave in response to the demand for confidential computing. Our design runs a thin TCB stack for security management on a security controller (SC), while leaving a large set of software (e.g., AI runtime, GPU driver, etc.) to the integrated microservers that operate enclaves. An enclaves is physically isolated from others through hardware and verified by the SC at its inception. Its microserver and computing units are restored to a secure state upon termination.We implemented HETEE on a real hardware system, and evaluated it with popular neural network inference and training tasks. Our evaluations show that HETEE can easily support the CDI tasks on the real-world scale and incurred a maximal throughput overhead of 2.17% for inference and 0.95% for training on ResNet152. Rui Hou 0001, XiaoFeng Wang 0001, Wenhao Wang 0001, Jiangfeng Cao, Boyan Zhao, Zhongpu Wang, Yuhui Zhang 0011, Jiameng Ying, Lixin Zhang 0002, Dan Meng 0002 |
SP | 4 |
| 2020 | TEADS: A Defense-aware Framework for Synthesizing Transient Execution Attacks
Tianlin Huo, Wenhao Wang 0001, Tingting Wang 0010, Mingshu Li 0001 |
TrustCom | 2 |
| 2020 | Partial-SMT: Core-scheduling Protection Against SMT Contention-based AttacksabstractNumerous recent works in side-channel attacks have experimentally shown that Simultaneous Multi-Threading (SMT) inherently has a broader attack surface as it exposes more microarchitecture components per-core than cross-core. Existing mechanisms that protect against these attacks either incur high execution costs or are ineffective against certain attack variants. In this paper, we propose Partial-SMT, a system based on core-scheduling that protects security-critical programs from all contention-based attacks due to SMT. Partial-SMT allocates some complete physical cores for the exclusive use of the individual applications and provides a user-level threading library linked into each application to control the placement of their threads on dedicated cores, thereby preventing the attacker from accessing shared CPU resources simultaneously on the victim's core. The key insight is that by limiting ourselves to SMT contention-based side channels, we can translate the protection into an allocation policy that allocates or frees computing resources with a granularity of one physical core. Security-critical applications can be implemented on-demand and coexist with existing applications. We demonstrate that Partial-SMT effectively defeats typical SMT contention-based attacks. We modify AES and SPEC 2006 to use Partial-SMT, and they all incur the slight negligible performance overhead. Yeping He, Qiming Zhou, Hengtai Ma, Liang He 0011, Wenhao Wang 0001 |
TrustCom | 6 |
| 2018 | Beware of Your Screen: Anonymous Fingerprinting of Device Screens for Off-line Payment ProtectionabstractQR-code mobile payment becomes increasingly popular, being offered by major banks (e.g., ICBC) and payment service providers (e.g., PayPal). Unlike mobile payment solutions provided by hardware vendors (e.g., Apple Pay and Samsung Pay), QR code payment schemes do not rely on any hardware support and can therefore be easily deployed. However, the security guarantee of the new scheme is less clear: in the absence of hardware protection, users' digital wallet can be vulnerable to an OS-level adversary, who could steal her secret for generating payment tokens. Zhe Zhou 0001, Di Tang 0001, Wenhao Wang 0001, XiaoFeng Wang 0001, Zhou Li 0001, Kehuan Zhang |
ACSAC | 3 |
| 2018 | Correlation Cube Attacks: From Weak-Key Distinguisher to Key Recovery
Meicheng Liu, Jingchun Yang, Wenhao Wang 0001, Dongdai Lin |
EUROCRYPT (2) | 3 |
| 2018 | Racing in Hyperspace: Closing Hyper-Threading Side Channels on SGX with Contrived Data RacesabstractIn this paper, we present HYPERRACE, an LLVM-based tool for instrumenting SGX enclave programs to eradicate all side-channel threats due to Hyper-Threading. HYPERRACE creates a shadow thread for each enclave thread and asks the underlying untrusted operating system to schedule both threads on the same physical core whenever enclave code is invoked, so that Hyper-Threading side channels are closed completely. Without placing additional trust in the operating system's CPU scheduler, HYPERRACE conducts a physical-core co-location test: it first constructs a communication channel between the threads using a shared variable inside the enclave and then measures the communication speed to verify that the communication indeed takes place in the shared L1 data cache-a strong indicator of physical-core co-location. The key novelty of the work is the measurement of communication speed without a trustworthy clock; instead, relative time measurements are taken via contrived data races on the shared variable. It is worth noting that the emphasis of HYPERRACE's defense against Hyper-Threading side channels is because they are open research problems. In fact, HYPERRACE also detects the occurrence of exception-or interrupt-based side channels, the solution.s of which have been studied by several prior works. Guoxing Chen, Wenhao Wang 0001, Tianyu Chen 0018, Sanchuan Chen, Yinqian Zhang, XiaoFeng Wang 0001, Ten-Hwang Lai, Dongdai Lin |
IEEE Symposium on Security and Privacy | 2 |
| 2017 | Leaky Cauldron on the Dark Land: Understanding Memory Side-Channel Hazards in SGXabstractSide-channel risks of Intel's SGX have recently attracted great attention. Under the spotlight is the newly discovered page-fault attack, in which an OS-level adversary induces page faults to observe the page-level access patterns of a protected process running in an SGX enclave. With almost all proposed defense focusing on this attack, little is known about whether such efforts indeed raises the bar for the adversary, whether a simple variation of the attack renders all protection ineffective, not to mention an in-depth understanding of other attack surfaces in the SGX system. In the paper, we report the first step toward systematic analyses of side-channel threats that SGX faces, focusing on the risks associated with its memory management. Our research identifies 8 potential attack vectors, ranging from TLB to DRAM modules. More importantly, we highlight the common misunderstandings about SGX memory side channels, demonstrating that high frequent AEXs can be avoided when recovering EdDSA secret key through a new page channel and fine-grained monitoring of enclave programs (at the level of 64B) can be done through combining both cache and cross-enclave DRAM channels. Our findings reveal the gap between the ongoing security research on SGX and its side-channel weaknesses, redefine the side-channel threat model for secure enclaves, and can provoke a discussion on when to use such a system and how to use it securely. Wenhao Wang 0001, Guoxing Chen, Xiaorui Pan, Yinqian Zhang, XiaoFeng Wang 0001, Vincent Bindschaedler, Haixu Tang, Carl A. Gunter |
CCS | 1 |
| 2015 | Searching cubes for testing Boolean functions and its application to TriviumabstractIn this paper, we describe a sub-maximal degree monomial test and propose a heuristic algorithm for searching favourable cubes, for testing Boolean functions formed by stream ciphers. We apply them to Trivium, and mount a distinguisher on Trivium reduced to 839 rounds with 237complexity, which is so far the best distinguisher on reduced Trivium. Meicheng Liu, Dongdai Lin, Wenhao Wang 0001 |
ISIT | 3 |
| 2013 | Analysis of Multiple Checkpoints in Non-perfect and Perfect Rainbow Tradeoff Revisited
Wenhao Wang 0001, Dongdai Lin |
ICICS | 1 |
| 2012 | A New Variant of Time Memory Trade-Off on the Improvement of Thing and Ying's Attack
Zhenqi Li, Yao Lu 0002, Wenhao Wang 0001, Bin Zhang 0003, Dongdai Lin |
ICICS | 3 |
| 2011 | Improvement and Analysis of VDP Method in Time/Memory Tradeoff Applications
Wenhao Wang 0001, Dongdai Lin, Zhenqi Li, Tianze Wang |
ICICS | 1 |