Peinan Li

dblp:119/1110 · DBLP profile ↗
← Back
25ranked-venue papers
2as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 2 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RIFT: A Burst-Aware and Log-Horizon Transformer for Ransomware Detection
Zhilu Wang, Peinan Li, Lingbo Zhao, Fengkai Yuan, Dan Meng 0002, Rui Hou 0001
ICIC (11)2
2026 NXT: Sharable Trusted Execution Environment for Multi-Tenant NPU Cluster
abstract
Cloud AI services have experienced rapid development, raising concerns to privacy protection of cloud tenants. Many proposals have been made to use the NPU Trusted Execution Environment (TEE) to protect AI workloads on the cloud. However, existing designs typically bind NPUs exclusively to a single tenant, preventing multiple tenants from sharing the computing power of the TEE-NPUs.As such, we have designed a novel TEE for discrete NPUs, named NXT (NPU eXtension for Trust), which breaks the exclusive binding architecture and allows multiple tenants to securely share the TEE-NPU cluster. Firstly, we introduced an NPU Trusted Agent (NTA) to the most privileged level of the TEE system to assist in the global scheduling of the TEE-NPU cluster. Secondly, we implemented a flexible isolation mechanism to provide security for multi-tenant fine-grained sharing of NPU resources. Thirdly, we support efficient communication between workloads within TEE-NPUs and with legacy NPUs, accelerating multi-workload collaborative computing. We evaluated NXT by extending gem5 and a cycle-accurate NPU simulator to build a prototype. Results show that NXT improves overall utilization by up to 3.49× and ANTT by up to 7.24×, with only 6.38% average overhead for scheduling and isolation. It also boosts parallel inference performance by 58.3% for GPT-2(XL) and 63.6% for LLaMA-13B compared to static protection schemes.
Shiwen Wang 0002, Peinan Li, Yunkai Bai, Wu Luo, Guang Yan, Dan Meng 0002, Rui Hou 0001
IEEE Trans. Computers2
2025 LegoZK: A Dynamically Reconfigurable Accelerator for Zero-Knowledge Proof
abstract
Zero-knowledge proof (ZKP) allows a prover to convince a verifier of the truth of a statement without revealing any secret information. This property is utilized in numerous privacy-preserving applications. However, the huge overhead of proof generation impedes the widespread adoption of ZKP. As a result, many ZKP accelerators have been developed to speed up proof generation. However, existing accelerators are designed at the granularity of core operators and exhibit low hardware resource utilization and limited adaptability. In this paper, we identify the commonality of all computation stages in proof generation at the level of basic finite field arithmetic operations. Based on this insight, we propose LegoZK, a dynamically reconfigurable hardware accelerator for ZKP. LegoZK employs finite field arithmetic units (FAUs) as its fundamental components and integrates these FAUs with a hierarchical on-chip network (NoC). By dynamically configuring the FAUs and the NoC, LegoZK can effectively accelerate the entire proof generation process, achieving higher overall performance. Additionally, for the most time-consuming MSM, this paper proposes a fast, fully pipelined bucket reduction algorithm based on lookup tables, which significantly reduces the latency of MSM. Experimental results demonstrate that LegoZK achieves on average speedup of $31.96 \times$ and $11.30 \times$ in proof generation compared to the state-of-the-art ZKP ASIC accelerator PipeZK and the GPU accelerator GZKP, respectively. And compared to PipeZK, LegoZK achieves $\mathbf{5 0. 1 \%}$ area reduction and $\mathbf{3 7. 7 \%}$ power consumption reduction.
Zhengbang Yang, Lutan Zhao, Peinan Li, Boyan Zhao, Dan Meng 0002, Rui Hou 0001
HPCA3
2025 RanDoctor: System-Level Ransomware Detection with ProbSparse Self-Attention
abstract
Ransomware attacks pose significant threats and have caused substantial economic losses across various industries worldwide. Existing defense mechanisms typically focus on detecting ransomware in environments free from interference by other legitimate programs. However, in real-world applications, ransomware often coexists with normal programs, resulting in fragmented behavioral patterns that reduce detection accuracy. To address this issue, we propose a system-level ransomware detection approach, named RanDoctor. This method leverages long-time series analysis to capture the behavioral characteristics of ransomware, thereby improving the comprehensiveness and accuracy of detection. To further enhance system performance, we design the Ranformer model, incorporating the ProbSparse self-attention mechanism and a distillation process. Experimental results demonstrate that the RanDoctor system achieves a detection accuracy of 99.5%, representing a 8.0% improvement over state-of-the-art detection models.
Zhilu Wang, Peinan Li, Lingbo Zhao, Fengkai Yuan, Rui Hou 0001, Dan Meng 0002
ICASSP2
2025 Tips: Augment Memory Tagging to Defend Against Prefetcher Side Channels
abstract
Hardware prefetchers are essential for hiding memory latency and improving performance in commercial processors. However, recent studies have revealed that they can be exploited to launch side-channel attacks that leak sensitive data, recover cryptographic keys, and break the isolation of trusted execution environments. We observe that such attacks closely resemble classic memory safety violations, including buffer overflows, type confusion, use-after-free, and data race. This paper presents TIPS (Tag AugmentatIon for Prefetcher Security), a lightweight extension to the memory safety mechanisms already deployed in commercial processors. TIPS enhances memory tagging to protect prefetchers by enforcing tag-based array bounds, validating pointer types, associating prefetch patterns with their source threads or cores, and suppressing contentionbased interference. Experiments demonstrate that TIPS incurs less than 1.50 % performance overhead and 0.81 % area cost, while providing strong defense against a broad class of prefetcher side-channel attacks.
Yubiao Huang, Peinan Li, Huan Qiao, Yunkai Bai, Shiwen Wang 0002, Dan Meng 0002, Rui Hou 0001
ICCD2
2025 RanHunter: Advancing Ransomware Detection with Channel Attention and Multi-head Attention
Zhilu Wang, Peinan Li, Lingbo Zhao, Fengkai Yuan, Rui Hou 0001, Dan Meng 0002
ICIC (4)2
2025 Sonar: A Hardware Fuzzing Framework to Uncover Contention Side Channels in Processors
Kanqi Zhang, Peinan Li, Zelong Du, Quanchen Liu, Yongqiang Lyu 0001, Yu Jiang 0001, Dan Meng 0002, Rui Hou 0001
MICRO2
2025 An Improved Method for Monitoring Subglacial Lake Activity in Antarctica From ICESat-2
abstract
Subglacial lakes are an important part of the Antarctic basal hydrological system, with many active subglacial lakes distributed in the steep topography of eastern Antarctica. When the ice surface has a slope, the horizontal geolocation error and elevation error of ICESat-2 altimetry data introduce additional uncertainty in estimating the ice surface elevation change, which will affect our ability to precisely monitor the activities of subglacial lakes. Therefore, based on the classical repeat-track analysis, we use weighted total least-squares adjustment to quantify the impact of horizontal geolocation and elevation errors on the fitted ice surface elevation. Through an iterative process, we derive precise time series of elevation changes to meet the need for monitoring subglacial lake hydrological activities. Using this improved method, we obtained the volume changes of the CookE2 subglacial lake in East Antarctica from March 2019 to March 2023. The results showed that Lake CookE2 was continuously recharged by subglacial water during this period, with an average equivalent recharge rate of 0.054 km³/yr. The water supply in the main lake accounted for about 68.5% of the total. The SW lobe, located away from the main basin, exhibited hydrological activity again after the drainage event and was reclassified into the active lake domain. The new lake area is ~ 118.73 km2. Compared to the traditional method, the improved approach reduces the variance of unit weight by ~0.044 m in the lake basin and by ~0.062 m on the steep southern basin margin. The precision of elevation change in rugged topography or steep slope areas has been significantly improved, aiding in precisely monitoring of subglacial lakes water volume changes and in determining their outlines.
Jun Liu 0077, Denghui Tang, Xiangbin Cui, Huan Xie 0001, Peinan Li
IEEE Trans. Geosci. Remote. Sens.6
2025 Chameleon: An Efficient FHE Scheme Switching Acceleration on GPUs
abstract
Fully homomorphic encryption (FHE) enables direct computation on encrypted data, making it a crucial technology for privacy protection. However, FHE suffers from significant performance bottlenecks. In this context, GPU acceleration offers a promising solution to bridge the performance gap. Existing efforts primarily focus on single-class FHE schemes, which fail to meet the diverse requirements of data types and functions, prompting the development of hybrid multi-class FHE schemes. However, studies have yet to thoroughly investigate specific GPU optimizations for hybrid FHE schemes. In this paper, we present an efficient GPU-based FHE scheme switching acceleration named Chameleon. First, we propose a scalable NTT acceleration design that adapts to larger CKKS polynomials and smaller TFHE polynomials. Specifically, Chameleon tackles synchronization issues by fusing stages to reduce synchronization, employing polynomial coefficient shuffling to minimize synchronization scale, and utilizing an SM-aware combination strategy to identify the optimal switching point. Second, Chameleon is the first to comprehensively analyze and optimize critical switching operations. It introduces CMux-level parallelization to accelerate LUT evaluation and a homomorphic rotation-free matrixvector multiplication to improve repacking efficiency. Finally, Chameleon outperforms the state-of-the-art GPU implementations by 1.23× in CKKS HMUL and 1.15× in bootstrapping. It also achieves up to 4.87× and 1.51× speedups for TFHE bootstrapping compared to CPU and GPU versions, respectively, and delivers a 67.3× average speedup for scheme switching over CPU-based implementation.
Haoqi He, Lutan Zhao, Peinan Li, Zhihao Li 0001, Dan Meng 0002, Rui Hou 0001
IEEE Trans. Parallel Distributed Syst.4
2024 SecPaging: Secure Enclave Paging with Hardware-Enforced Protection against Controlled-Channel Attacks
abstract
As a prevalent privacy-preserving technology, Trusted Execution Environment has become widely adopted in numerous commercial processors. Nonetheless, they remain susceptible to various controlled-channel attacks. Untrusted operating systems can deduce enclave secrets by manipulating page tables or observing allocation- or swap-based page faults. In this paper, we propose SecPaging, a novel secure enclave paging mechanism based on hardware-enforced and microcode-supported protection to prevent these attacks. First, enclave PTEs are protected through hardware isolation, preventing privileged attackers from malicious tampering or observations. Second, an Eager-Allocation mechanism is employed to prevent allocation-based controlled-channel attacks. Besides, a Record-Reload mechanism is proposed to prevent swap-based controlled-channel attacks. We simulate SecPaging on real SGX. Experiments demonstrate that controlled channel attacks can be defended with minimal performance overhead.
Yunkai Bai, Peinan Li, Yubiao Huang, Shiwen Wang 0002, Xingbin Wang, Dan Meng 0002, Rui Hou 0001
DAC2
2024 EnTurbo: Accelerate Confidential Serverless Computing via Parallelizing Enclave Startup Procedure
abstract
Serverless computing has gained widespread attention, and Trusted Execution Environments (TEEs) are well-suited for safeguarding user privacy. However, the additional startup procedure introduced by TEEs imposes considerable performance overhead on confidential serverless workloads. This paper introduces a novel parallelized enclave startup design, EnTurbo, which eliminates the integrity dependence of the enclave startup procedure, accelerating it while ensuring its security. Additionally, EnTurbo parallelizes the measurement procedure, enabling multi-thread measurement for acceleration with provable security. We evaluate EnTurbo by running confidential serverless workloads on SGX simulation mode. Results show that EnTurbo effectively speeds up enclave serverless by 1.42x-6.48x (SGXv1) and 1.33x-3.76x (SGXv2).
Yifan Zhu 0008, Peinan Li, Yunkai Bai, Yubiao Huang, Shiwen Wang 0002, Xingbin Wang, Dan Meng 0002, Rui Hou 0001
DAC2
2024 HyperTEE: A Decoupled TEE Architecture with Secure Enclave Management
abstract
Trusted Execution Environment (TEE) architectures have been deployed in various commercial processors to provide secure environments for confidential programs and data. However, as a relatively new feature against security threats, existing designs still face a number of problems. Exploiting the management vulnerabilities, attackers can disclose secrets via controlled-channel or micro-architecture side-channel attacks. To address these problems, this paper proposes a novel TEE architecture, named HyperTEE. In our architecture, enclave management tasks are decoupled from the original computing subsystem to a dedicated, physically isolated Enclave Manage-ment Subsystem (EMS). A properly architected EMS prevents current management vulnerabilities and offers more secure enclave communication. We implemented the HyperTEE prototype on the FPGA platform. Experiments show that HyperTEE only introduces less than 1% area overhead, and 2.0 % and 1.9 % performance overhead on average for enclaves and non-enclave workloads, respectively.
Yunkai Bai, Peinan Li, Yubiao Huang, Michael C. Huang 0001, Shijun Zhao, Lutan Zhao, Fengwei Zhang, Dan Meng 0002, Rui Hou 0001
MICRO2
2024 Data-driven prediction for curved pipe jacking performance during underwater excavation of ancient shipwreck using an attention-based graph convolutional network approach
Zeyu Dai 0003, Peinan Li, Yi Rui, Yixin Zhai
Expert Syst. Appl.2
2023 ChaosINTC: A Secure Interrupt Management Mechanism against Interrupt-based Attacks on TEE
abstract
For Trusted Execution Environment (TEE), interrupt-based side-channel attacks are becoming significant threats. Malicious supervisors use interrupts to perform single-step side-channel attacks or to improve the accuracy of existing side-channel attacks. This paper proposes a secure interrupt handle mechanism dedicated to TEE, named ChaosINTC. (1) To prevent frequent interrupts, a dynamic interrupt response delay mechanism delays the interrupt delivery with a variable time. (2) To prevent maliciously modifying ISRs, an interrupt handler protecting mechanism performs isolation and integrity checking. We deployed ChaosINTC on an open-source RISC-V core and evaluated its performance via FPGA. Our design provides strong security with marginal hardware and performance costs.
Yifan Zhu 0008, Peinan Li, Lutan Zhao, Dan Meng 0002, Rui Hou 0001
DAC2
2023 NTTFusion: Efficient Number Theoretic Transform Acceleration on GPUs
abstract
Fully homomorphic encryption (FHE) holds great promise as an encryption technology for safeguarding privacy by enabling computations directly on encrypted data. However, FHE encounters significant performance bottlenecks due to the extensive utilization of number theoretic transform (NTT) and its inverse (INTT). Therefore, it is crucial to accelerate NTT to enhance the efficiency of FHE. Conventional NTT implementations rely on mandatory synchronization to maintain data consistency, which leads to two critical problems: excessive synchronizations and unexplored synchronization switching points. This paper presents NTTFusion, an efficient GPU-based NTT acceleration design that focuses on boosting the performance of NTT. (i) To reduce the number of synchronizations, we propose two types of stage fusion methods specifically designed for different polynomial lengths. For small polynomial lengths, we employ the butterfly decomposition approach, while for large polynomial lengths, we leverage the thread aggregation method. (ii) To explore the optimal synchronization switching point, we propose an SM-aware synchronization combination strategy to balance synchronization overhead and hardware utilization. Finally, we conduct experiments on a realistic NVIDIA GPU server and demonstrate that the butterfly decomposition method achieves up to 1.37× speedup compared to the state-of-the-art implementation. Furthermore, the thread aggregation method can yield up to 1.32× speedup for larger polynomial lengths. The optimal synchronization switching point-based NTT, which incorporates thread aggregation, can produce a maximum 1.6× performance boost under a typical large polynomial length.
Peinan Li, Rui Hou 0001, Dan Meng 0002
ICCD2
2023 Dynamic prediction for attitude and position of shield machine in tunneling: A hybrid deep learning method considering dual attention
Zeyu Dai 0003, Peinan Li, Mengqi Zhu, Hehua Zhu, Yixin Zhai
Adv. Eng. Informatics2
2023 Architecting the Autocuckoo Filter to Defend Against Cross-Core Cache Attacks
abstract
Cross-core cache timing side-channel attacks, which observe cache access behavior of victims running on different physical cores to infer sensitive information, have become a significant threat. Although the attacks are covert, they cause the attacked cachelines to frequently migrate among cache hierarchies, rendering abnormal traffic. Based on this observation, the proposed scheme PiPoMonitor records cache-memory access traffic and prefetch suspicious lines under attack to interfere with adversaries’ probes. In pursuit of security and performance, PiPoMonitor exploits a Cuckoo filter as the recording structure and introduces two features to it: 1) autonomic deletion and 2) relocation accelerating. The former exponentially increases the uncertainty of record eviction against reverse engineering attacks, while the latter leverages a pipelined architecture to alleviate the impact of intensive filter queries on the memory critical path. PiPoMonitor is not only able to effectively mitigate cross-core cache attacks and defeat sophisticated defense-aware attackers but also induces a negligible performance penalty and acceptable hardware overhead.
Fengkai Yuan, Kai Wang 0061, Jiameng Ying, Rui Hou 0001, Lutan Zhao, Peinan Li, Yifan Zhu 0008, Zhenzhou Ji, Dan Meng 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2023 HE-Booster: An Efficient Polynomial Arithmetic Acceleration on GPUs for Fully Homomorphic Encryption
abstract
Fully Homomorphic Encryption (FHE) enables secure offloading of computations to untrusted cloud servers as it allows computing on encrypted data. However, existing well-known FHE schemes suffer from heavy performance overheads. Thus numerous accelerations based on FPGAs, ASICs, and GPUs have been proposed. Compared to FPGAs and ASICs, GPUs have obvious advantages in productivity and development costs. And also, GPUs have already been widely deployed in commercial cloud or supercomputing centers. Therefore, we present HE-Booster, an efficient GPU-based FHE acceleration design. For single-GPU acceleration, a thorough systematic design is exploited to map five common phases in typical FHE schemes to the GPU parallel architecture. In particular, inspired by the regular architecture of NTT/INTT, a novel inter-thread local synchronization is proposed to exploit thread-level parallelism. For multi-GPU acceleration, we propose a scalable parallelization design that exploitsdata-level parallelismthrough fine-grained data partition under different representations. Finally, experiments on 1 NVIDIA GPU demonstrate that our work outperforms 251.7×, 78.5× and 164.9× than three mainstream CPU-based libraries HElib, SEAL, and PALISADE, and up to 170.5× speedup is obtained compared to the GPU-accelerated library cuHE. What's more, performing 8 homomorphic multiplications on 8 GPUs can deliver up to a 7.66× performance boost compared to a single-GPU implementation.
Peinan Li, Rui Hou 0001, Zhihao Li 0001, Jiangfeng Cao, XiaoFeng Wang 0001, Dan Meng 0002
IEEE Trans. Parallel Distributed Syst.2
2022 Conditional address propagation: an efficient defense mechanism against transient execution attacks
abstract
Speculative execution is a critical technique in modern high performance processors. However, continuously exposed transient execution attacks, including Spectre and Meltdown, disclosed a large attack surface in mispredicted execution. Current state-of-the-art defense strategy blocks all memory accesses that use addresses loaded speculatively. However, propagation of base addresses is common in general applications and we find that more than 60% blocked memory accesses use propagated base rather than offset addresses. Therefore, we propose a novel hardware defense mechanism, named Conditional Address Propagation, to identify safe base addresses through taint tracking and address checking by a History Table. Then, the safe base addresses are allowed to be propagated to retrieve performance. For remaining unsafe addresses, they cannot be propagated for security. We constructed experiments on cycle-accurate Gem5 simulator. Compared to the representative study, STT, our mechanism effectively decreases the performance overhead from 13.27% to 1.92% targeting Spectre-type and 19.66% to 5.23% targeting all-type cache-based transient execution attacks.
Peinan Li, Rui Hou 0001, Lutan Zhao, Yifan Zhu 0008, Dan Meng 0002
DAC1
2022 HyBP: Hybrid Isolation-Randomization Secure Branch Predictor
abstract
Recently exposed vulnerabilities reveal the necessity to improve the security of branch predictors. Branch predictors record history about the execution of different processes, and such information from different processes are stored in the same structure and thus accessible to each other. This leaves the attackers with the opportunities for malicious training and malicious perception. Physical or logical isolation mechanisms such as using dedicated tables and flushing during context-switch can provide security but incur non-trivial costs in space and/or execution time. Randomization mechanisms incurs the performance cost in a different way: those with higher securities add latency to the critical path of the pipeline, while the simpler alternatives leave vulnerabilities to more sophisticated attacks.This paper proposes HyBP, a practical hybrid protection and effective mechanism for building secure branch predictors. The design applies the physical isolation and randomization in the right component to achieve the best of both worlds. We propose to protect the smaller tables with physically isolation based on (thread, privilege) combination; and protect the large tables with randomization. Surprisingly, the physical isolation also significantly enhances the security of the last-level tables by naturally filtering out accesses, reducing the information flow to these bigger tables. As a result, key changes can happen less frequently and be performed conveniently at context switches. Moreover, we propose a latency hiding design for a strong cipher by precomputing the "code book" with a validated, cryptographically strong cipher. Overall, our design incurs a performance penalty of 0.5% compared to 5.1% of physical isolation under the default context switching interval in Linux.
Lutan Zhao, Peinan Li, Rui Hou 0001, Michael C. Huang 0001, Xuehai Qian, Lixin Zhang 0002, Dan Meng 0002
HPCA2
2021 A Lightweight Isolation Mechanism for Secure Branch Predictors
abstract
Recently exposed vulnerabilities reveal that branch predictors shared by different processes leave the attackers with the opportunities for malicious training and perception. Instead of flush-based or physical isolation of hardware resources, we want to achieve isolation of the content in these hardware tables with some lightweight processing using randomization as follows. (1) Content encoding. We propose to use hardware-based thread-private random numbers to encode the contents of the branch predictor tables. It achieves a similar effect of logical isolation but adds little in terms of space or time overheads. (2) Index encoding. We propose a randomized index mechanism of the branch predictor. This disrupts the correspondence between the branch instruction address and the branch predictor entry, thus increases the noise for malicious perception attacks. Our analyses using an FPGA-based RISC-V processor prototype and additional auxiliary simulations suggest that the proposed mechanisms incur a very small performance cost while providing strong protection.
Lutan Zhao, Peinan Li, Rui Hou 0001, Michael C. Huang 0001, Jiazhen Li, Lixin Zhang 0002, Xuehai Qian, Dan Meng 0002
DAC2
2021 PiPoMonitor: Mitigating Cross-core Cache Attacks Using the Auto-Cuckoo Filter
abstract
Cache side channel attacks obtain victim cache line access footprint to infer security-critical information. Among them, cross-core attacks exploiting the shared last level cache are more threatening as their simplicity to set up and high capacity. Stateful approaches of detection-based mitigation observe precise cache behaviors and protect specific cache lines that are suspected of being attacked. However, their recording structures incur large storage overhead and are vulnerable to reverse engineering attacks. Exploring the intrinsic non-determinate layout of a traditional Cuckoo filter, this paper proposes a space efficient Auto-Cuckoo filter to record access footprints, which succeed to decrease storage overhead and resist reverse engineering attacks at the same time. With Auto-Cuckoo filter, we propose PiPoMonitor to detect Ping-Pong patterns and prefetch specific cache line to interfere with adversaries' cache probes. Security analysis shows the PiPoMonitor can effectively mitigate cross-core attacks and the Auto-Cuckoo filter is immune to reverse engineering attacks. Evaluation results indicate PiPoMonitor has negligible impact on performance and the storage overhead is only 0.37%, an order of magnitude lower than previous stateful approaches.
Fengkai Yuan, Kai Wang 0061, Rui Hou 0001, Peinan Li, Lutan Zhao, Jiameng Ying, Amro Awad, Dan Meng 0002
DATE5
2021 A Novel Probabilistic Saturating Counter Design for Secure Branch Predictor
Lutan Zhao, Rui Hou 0001, Kai Wang 0061, Yu-Lan Su, Peinan Li, Dan Meng 0002
J. Comput. Sci. Technol.5
2021 Exploiting Security Dependence for Conditional Speculation Against Spectre Attacks
abstract
Speculative execution side-channel vulnerabilities such as Spectre reveal that conventional architecture designs lack security consideration. This article proposes a software transparent defense framework, named as Conditional Speculation, against Spectre vulnerabilities found on traditional out-of-order microprocessors. It introduces the concept of security dependence to mark speculative memory instructions which could leak information with potential security risks. More specifically, security-dependent instructions are detected and marked with suspect speculation flags in the Issue Queue. All the instructions can be speculatively issued for execution in accordance with the classic out-of-order pipeline. For those instructions with suspect speculation flags, they are considered as safe instructions if their speculative execution dose not refill new cache lines with unauthorized privilege data. Otherwise, they are considered as unsafe instructions and thus not allowed to execute speculatively. To pursue a balance of performance and security, we investigate two filtering mechanisms, Cache-hit-based Hazard Filter and Trusted Page Buffer-based Hazard Filter to filter out false security hazards. As for true security hazards, we have two approaches to prevent them from changing cache states. One is to block all unsafe access, the other is to fetch them from lower-level caches or memory to a speculative buffer temporarily, and refill them after confirming that they are on the correct execution path. Our design philosophy is to speculatively execute safe instructions to maintain the performance benefits of out-of-order execution while delaying the cache updates for speculative execution of unsafe instructions for security consideration. We evaluate Conditional Speculation in terms of performance, security, and area. The experimental results show that the hardware overhead is marginal and the performance overhead is minimal.
Lutan Zhao, Peinan Li, Rui Hou 0001, Michael C. Huang 0001, Peng Liu 0005, Lixin Zhang 0002, Dan Meng 0002
IEEE Trans. Computers2
2019 Conditional Speculation: An Effective Approach to Safeguard Out-of-Order Execution Against Spectre Attacks
abstract
Speculative execution side-channel vulnerabilities such as Spectre reveal that conventional architecture designs lack security consideration. This paper proposes a software transparent defense mechanism, named as Conditional Speculation, against Spectre vulnerabilities found on traditional out-of-order microprocessors. It introduces the concept of security dependence to mark speculative memory instructions which could leak information with potential security risk. More specifically, security-dependent instructions are detected and marked with suspect speculation flags in the Issue Queue. All the instructions can be speculatively issued for execution in accordance with the classic out-of-order pipeline. For those instructions with suspect speculation flags, they are considered as safe instructions if their speculative execution will not refill new cache lines with unauthorized privilege data. Otherwise, they are considered as unsafe instructions and thus not allowed to execute speculatively. To reduce the performance impact from not executing unsafe instructions speculatively, we investigate two filtering mechanisms, Cachehit based Hazard Filter and Trusted Page Buffer based Hazard Filter to filter out false security hazards. Our design philosophy is to speculatively execute safe instructions to maintain the performance benefits of out-of-order execution while blocking the speculative execution of unsafe instructions for security consideration. We evaluate Conditional Speculation in terms of performance, security and area. The experimental results show that the hardware overhead is marginal and the performance overhead is minimal.
Peinan Li, Lutan Zhao, Rui Hou 0001, Lixin Zhang 0002, Dan Meng 0002
HPCA1