VLDB 2026 Research / reviewers in the wild / expert
Lutan Zhao
dblp:238/2108
· DBLP profile ↗
28ranked-venue papers
4as first author
27since 2021 · last 2026
0000-0003-3672-6463ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 3 first-author · 20 since 2021Security and privacy · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sub-Millisecond Gate BootstrappingabstractGate bootstrapping is a core primitive that enables arbitrary circuit evaluation in fully homomorphic encryption (FHE), where blind rotation remains the dominant performance bottleneck. In this work, we present a sub-millisecond NTRU-based gate bootstrapping scheme that achieves state-of-the-art performance through coordinated algorithmic, software, and hardware-level optimizations. Chunling Chen, Zhihao Li 0001, Qingyun Niu, Xianhui Lu, Ruida Wang, Lutan Zhao, Rui Hou 0001 |
AsiaCCS | 6 |
| 2026 | Maverick: Rethinking TFHE Bootstrapping on GPUs via Algorithm-Hardware Co-DesignabstractFully homomorphic encryption (FHE) enables arbitrary computation over encrypted data (ciphertext) without compromising confidentiality. Within this family, TFHE features versatile bootstrapping mechanisms that is attractive for security-critical applications. However, its prohibitive computational cost severely limits practical deployment. While hardware acceleration is promising, mere compute scaling fails to overcome the inherent barriers. In particular, the combination of limited algorithmic parallelism and inadequate understanding of hardware behaviors prevents full exploitation of the available performance headroom. Haoqi He, Lutan Zhao, Qingyun Niu, Dan Meng 0002, Rui Hou 0001 |
ASPLOS (2) | 3 |
| 2026 | Thunder: Efficient Multi-node FHE Acceleration Framework via In-Transit Computation
Lutan Zhao, Qingyun Niu, Yinhang Zheng, Zhengbang Yang, Boyan Zhao, Rui Hou 0001 |
Euro-Par (1) | 2 |
| 2026 | OmniZK: A Versatile Accelerator Architecture for Zero-Knowledge Proofs
Zhengbang Yang, Lutan Zhao, Rui Hou 0001 |
Euro-Par (1) | 2 |
| 2026 | SwiftFL: Enabling Speculative Training for On-Device Federated Deep LearningabstractFederated deep learning (FDL) is a promising privacy-preserving approach for training deep neural networks on distributed datasets without raw data sharing. But the classical synchronous FDL faces straggler problem: slow trainers severely impede overall efficiency. Inspired by speculative execution techniques in modern processors, this paper proposes SwiftFL, a novel and efficient speculative training system for FDL. Instead of simply waiting for slower trainer, SwiftFL proactively updates the global model with predicted gradients, enabling faster trainers to speculatively initiate the next training round. Furthermore, a gradient compensation technique is proposed to correct mispredicted training without re-training. Finally, to overcome the model-drift problem caused by fast trainers perform more local training rounds, we propose a client selection strategy. This strategy determines whether trainers should perform speculative training by striking a balance between two metrics: model drift degree and local training efficiency. In the evaluation, we compare SwiftFL with four state-of-the-art FDL systems and demonstrate that SwiftFL achieves an average speedup of 6.08× while maintaining consistent final model accuracy. Yuhui Zhang 0011, Guang Yan, Xin Zhang 0110, Zimu Guo, Lutan Zhao, Jiangfeng Cao, Dan Meng 0002, Rui Hou 0001 |
EuroSys | 5 |
| 2026 | Peregrine: Accelerating TFHE Bootstrapping on GPUs via Multi-Level External Product Co-DesignabstractFully Homomorphic Encryption (FHE) is a ground-breaking cryptographic technology that enables computation directly on encrypted data, but its practical adoption continues to be hindered by high computational costs. GPUs have emerged as an increasingly attractive acceleration platform, offering massive parallelism and architectural flexibility to accommodate rapidly evolving FHE algorithms. Despite notable advances, most efforts remain confined to isolated, single-level optimizations, limiting their ability to push performance boundaries. In this paper, we present Peregrine, an efficient GPU-based TFHE acceleration built upon multi-level co-design of external product (EP) operations across parallelism, implementation, and scheduling. First, we propose a synchronization-free key unrolling technique that restructures the execution pipeline via operator decoupling, thereby unlocking greater EP-level parallelism. Second, we consolidate fragmented operators into a matrix-centric execution pattern, yielding a high-arithmetic-intensity kernel that substantially enhances external product efficiency. Third, we propose a hierarchical tiling strategy that reformats ring ciphertexts into the Module structure and schedules polynomiallevel tiles for fine-grained GPU mapping of EP operations. Experimental results show that Peregrine outperforms up to$176.9 \times$and$2.2 \times$over state-of-the-art CPU and GPU baselines, respectively, demonstrating its strong applicability to security-critical workloads. Haoqi He, Lutan Zhao, Dan Meng 0002, Rui Hou 0001 |
HPCA | 3 |
| 2026 | UniFHE: Faster Accelerator for FHE with Diverse Algebraic Structure and Balanced Memory SystemabstractFully homomorphic encryption (FHE) enables computations on encrypted data. Existing FHE schemes are primarily categorized into RLWE-based word-wise schemes and LWEbased bit-wise schemes. Efficient combination of different FHE schemes adapted to real-world applications has emerged as a research focus. This paper proposes UniFHE, the first FHE accelerator that supports diverse algebraic structures using general arithmetic units to achieve higher performance. UniFHE is compatible with both RLWE-based and LWE-based FHE schemes without modifications to their original algorithmic designs. To support both finite ring and complex field operations, UniFHE introduces a general arithmetic unit and further constructs core computation structures. To balance on-chip memory demands across different schemes, UniFHE adopts a multi-pipeline architecture for LWE-based schemes. The core functional units for RLWE-based schemes are spliced based on the LWE-based pipelines. Furthermore, an on-chip plaintext encoding mechanism significantly reduces off-chip memory bandwidth demands. Experimental results show that, beyond superior area and energy efficiency, UniFHE delivers up to$13.6 \times$higher performance compared to scheme-specific accelerator combinations. Moreover, in hybrid schemes, UniFHE achieves a$3.2 \times$speedup over the state-of-the-art unified FHE accelerator Trinity. Qingyun Niu, Lutan Zhao, Dan Meng 0002, Rui Hou 0001 |
HPCA | 2 |
| 2026 | A survey of optimization techniques for bootstrapping algorithms in FHEabstractAbstract Fully Homomorphic Encryption (FHE) enables arbitrary computation on encrypted data without decryption, making it a cornerstone of privacy-preserving outsourcing, such as cloud computing. However, homomorphic operations cause ciphertext noise to grow until decryption fails. The efficient solution is bootstrapping, which refreshes the noise in FHE ciphertexts to sustain arbitrary deep homomorphic evaluation. But in practice, bootstrapping consumes over 50% of total execution time, posing a serious obstacle to FHE adoption. This paper presents a systematic survey of FHE bootstrapping algorithms and their optimizations. We organize existing works into three main paradigms: word-wise bootstrapping for BGV, BFV, and CKKS schemes; bit-wise bootstrapping for FHEW and TFHE schemes; and hybrid bootstrapping, which leverages both word-wise schemes and bit-wise schemes. We analyze the evolution of crucial techniques, highlight latest advances in reducing latency, enhancing parallelism, and controlling noise growth, and compare the advantages and limitations of different schemes. Finally, we discuss emerging research trends. Lutan Zhao, Ruida Wang, Qingyun Niu, Xianhui Lu, Dan Meng 0002, Rui Hou 0001 |
Cybersecur. | 3 |
| 2025 | ShiftPIR: An Efficient PIR System with Gravity Shifting from Client to ServerabstractWe present ShiftPIR, a single-server Private Information Retrieval (PIR) protocol that gravity shifts both computation and communication overhead from the client to the server, thereby significantly improving overall efficiency. This shift is driven by the growing asymmetry between resource-constrained clients and compute-intensive servers, where server-side tasks can be effectively parallelized and scaled. To achieve this, ShiftPIR introduces a novel request generation method in which the client transmits only a compact plaintext offset derived from pre-uploaded seed ciphertexts. The server then reconstructs the full query ciphertexts using homomorphic rotations, eliminating the need for costly ciphertext generation and transmission on the client side. We further design a highly parallelizable query expansion mechanism that removes data dependencies between ciphertext rotations, enabling efficient GPU-based execution. Our experiments demonstrate that ShiftPIR reduces client-side latency to microseconds while maintaining a communication cost within 4X of the non-private baseline—far outperforming prior protocols with 104 -105 × overhead. Compared to the state-of-the-art protocol YPIR, ShiftPIR achieves up to 26X lower end-to-end latency. Lutan Zhao, Haoqi He, Wenzhe Lv, Dan Meng 0002, Rui Hou 0001 |
CCS | 2 |
| 2025 | LegoZK: A Dynamically Reconfigurable Accelerator for Zero-Knowledge ProofabstractZero-knowledge proof (ZKP) allows a prover to convince a verifier of the truth of a statement without revealing any secret information. This property is utilized in numerous privacy-preserving applications. However, the huge overhead of proof generation impedes the widespread adoption of ZKP. As a result, many ZKP accelerators have been developed to speed up proof generation. However, existing accelerators are designed at the granularity of core operators and exhibit low hardware resource utilization and limited adaptability. In this paper, we identify the commonality of all computation stages in proof generation at the level of basic finite field arithmetic operations. Based on this insight, we propose LegoZK, a dynamically reconfigurable hardware accelerator for ZKP. LegoZK employs finite field arithmetic units (FAUs) as its fundamental components and integrates these FAUs with a hierarchical on-chip network (NoC). By dynamically configuring the FAUs and the NoC, LegoZK can effectively accelerate the entire proof generation process, achieving higher overall performance. Additionally, for the most time-consuming MSM, this paper proposes a fast, fully pipelined bucket reduction algorithm based on lookup tables, which significantly reduces the latency of MSM. Experimental results demonstrate that LegoZK achieves on average speedup of $31.96 \times$ and $11.30 \times$ in proof generation compared to the state-of-the-art ZKP ASIC accelerator PipeZK and the GPU accelerator GZKP, respectively. And compared to PipeZK, LegoZK achieves $\mathbf{5 0. 1 \%}$ area reduction and $\mathbf{3 7. 7 \%}$ power consumption reduction. Zhengbang Yang, Lutan Zhao, Peinan Li, Boyan Zhao, Dan Meng 0002, Rui Hou 0001 |
HPCA | 2 |
| 2025 | Poseidon: A NAS-Based Ensemble Defense Method Against Multiple Perturbations
Yulan Su, Sisi Zhang, Zechao Lin, Xingbin Wang, Lutan Zhao, Dan Meng 0002, Rui Hou 0001 |
MMM (3) | 5 |
| 2025 | RobSparse: Automatic Search for GPU-Friendly Robust and Sparse Vision Transformers
Yulan Su, Sisi Zhang, Yan Wang 0122, Xingbin Wang, Lutan Zhao, Dan Meng 0002, Rui Hou 0001 |
MMM (3) | 5 |
| 2025 | Comet: Accelerating Private Inference for Large Language Model by Predicting Activation SparsityabstractWith the growing use of large language models (LLMs) hosted on cloud platforms to offer inference services, privacy concerns about the potential leakage of sensitive information are escalating. Secure Multi-Party Computation (MPC) is a promising solution to protect the privacy in LLM inference. However, MPC requires frequent inter-server communication, causing high performance overhead. Inspired by the prevalent activation sparsity of LLMs, where most neuron are not activated after non-linear activation functions, we propose an efficient private inference system, Comet. This system employs an accurate and fast predictor to predict the sparsity distribution of activation function output. Additionally, we introduce a new private inference protocol. It efficiently and securely avoids computations involving zero values by exploiting the spatial locality of the predicted sparsity distribution. While this computation-avoidance approach impacts the spatiotemporal continuity of KV cache entries, we address this challenge with a low-communication overhead cache refilling strategy that merges miss requests and incorporates a prefetching mechanism. Finally, we evaluate Comet on four common LLMs and compare it with six state-of-the-art private inference systems. Comet achieves a$1.87\times-2.63\times$speedup and a$1.94\times-2.64\times$communication reduction. Guang Yan, Yuhui Zhang 0011, Zimu Guo, Lutan Zhao, Xiaojun Chen 0004, Wenhao Wang 0001, Dan Meng 0002, Rui Hou 0001 |
SP | 4 |
| 2025 | Chameleon: An Efficient FHE Scheme Switching Acceleration on GPUsabstractFully homomorphic encryption (FHE) enables direct computation on encrypted data, making it a crucial technology for privacy protection. However, FHE suffers from significant performance bottlenecks. In this context, GPU acceleration offers a promising solution to bridge the performance gap. Existing efforts primarily focus on single-class FHE schemes, which fail to meet the diverse requirements of data types and functions, prompting the development of hybrid multi-class FHE schemes. However, studies have yet to thoroughly investigate specific GPU optimizations for hybrid FHE schemes. In this paper, we present an efficient GPU-based FHE scheme switching acceleration named Chameleon. First, we propose a scalable NTT acceleration design that adapts to larger CKKS polynomials and smaller TFHE polynomials. Specifically, Chameleon tackles synchronization issues by fusing stages to reduce synchronization, employing polynomial coefficient shuffling to minimize synchronization scale, and utilizing an SM-aware combination strategy to identify the optimal switching point. Second, Chameleon is the first to comprehensively analyze and optimize critical switching operations. It introduces CMux-level parallelization to accelerate LUT evaluation and a homomorphic rotation-free matrixvector multiplication to improve repacking efficiency. Finally, Chameleon outperforms the state-of-the-art GPU implementations by 1.23× in CKKS HMUL and 1.15× in bootstrapping. It also achieves up to 4.87× and 1.51× speedups for TFHE bootstrapping compared to CPU and GPU versions, respectively, and delivers a 67.3× average speedup for scheme switching over CPU-based implementation. Haoqi He, Lutan Zhao, Peinan Li, Zhihao Li 0001, Dan Meng 0002, Rui Hou 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | An Efficient Speculative Federated Tree Learning System With a Lightweight NN-Based PredictorabstractFederated tree-based models are popular in many real-world applications owing to their high accuracy and good interpretability. However, the classical synchronous method causes inefficient federated tree-based model training due to tree node dependencies. Inspired by speculative execution techniques in modern high-performance processors, this paper proposes FTSeir, a novel and efficient speculative federated learning system. Instead of simply waiting, FTSeir optimistically predicts the outcome of the prior tree node. By resolving tree node dependencies with a neural network-based split point predictor, the training tasks of child tree nodes can be executed speculatively in advance via separate threads. This speculation enables cross-layer concurrent training, thus significantly reducing the waiting time. Furthermore, we propose an eager verification mechanism to promptly identify mispredictions, thereby reducing wasted computing resources. On a misprediction, an incomplete rollback is triggered for quick recovery by reusing the output of the mis-speculative training, which reduces computational requirements. We implement FTSeir and evaluate its efficiency in a real-world federated learning setting with six public datasets. Evaluation results demonstrate that FTSeir achieves up to 3.45× and 3.60× speedup over the state-of-the-art gradient boosted decision trees and random forests implementations, respectively. Yuhui Zhang 0011, Hong Liao, Lutan Zhao, Yuncong Shao, Zhihong Tian 0001, Dan Meng 0002, Rui Hou 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2024 | SpecFL: An Efficient Speculative Federated Learning System for Tree-based Model TrainingabstractFederated tree-based models are popular in many real-world applications owing to their high accuracy and good interpretability. However, the classical synchronous method causes inefficient federated tree model training due to tree node dependencies. Inspired by speculative execution techniques in modern high-performance processors, this paper proposes SpecFL, a novel and efficient speculative federated learning system. Instead of simply waiting, SpecFL optimistically predicts the outcome of the prior tree node. By resolving tree node dependencies with a split point predictor, the training tasks of child tree nodes can be executed speculatively in advance via separate threads. This speculation enables cross-layer concurrent training, thus significantly reducing the waiting time. Furthermore, we propose a greedy speculation policy to exploit speculative training for deeper inter-layer concurrent training and an eager rollback mechanism for lossless model quality. We implement SpecFL and evaluate its efficiency in a real-world federated learning setting with six public datasets. The evaluation results demonstrate that SpecFL can be 2.08-3.33x and 2.14-3.44x faster than the state-of-the-art GBDT and RF implementations, respectively. Yuhui Zhang 0011, Lutan Zhao, Cheng Che, XiaoFeng Wang 0001, Dan Meng 0002, Rui Hou 0001 |
HPCA | 2 |
| 2024 | HyperTEE: A Decoupled TEE Architecture with Secure Enclave ManagementabstractTrusted Execution Environment (TEE) architectures have been deployed in various commercial processors to provide secure environments for confidential programs and data. However, as a relatively new feature against security threats, existing designs still face a number of problems. Exploiting the management vulnerabilities, attackers can disclose secrets via controlled-channel or micro-architecture side-channel attacks. To address these problems, this paper proposes a novel TEE architecture, named HyperTEE. In our architecture, enclave management tasks are decoupled from the original computing subsystem to a dedicated, physically isolated Enclave Manage-ment Subsystem (EMS). A properly architected EMS prevents current management vulnerabilities and offers more secure enclave communication. We implemented the HyperTEE prototype on the FPGA platform. Experiments show that HyperTEE only introduces less than 1% area overhead, and 2.0 % and 1.9 % performance overhead on average for enclaves and non-enclave workloads, respectively. Yunkai Bai, Peinan Li, Yubiao Huang, Michael C. Huang 0001, Shijun Zhao, Lutan Zhao, Fengwei Zhang, Dan Meng 0002, Rui Hou 0001 |
MICRO | 6 |
| 2023 | ChaosINTC: A Secure Interrupt Management Mechanism against Interrupt-based Attacks on TEEabstractFor Trusted Execution Environment (TEE), interrupt-based side-channel attacks are becoming significant threats. Malicious supervisors use interrupts to perform single-step side-channel attacks or to improve the accuracy of existing side-channel attacks. This paper proposes a secure interrupt handle mechanism dedicated to TEE, named ChaosINTC. (1) To prevent frequent interrupts, a dynamic interrupt response delay mechanism delays the interrupt delivery with a variable time. (2) To prevent maliciously modifying ISRs, an interrupt handler protecting mechanism performs isolation and integrity checking. We deployed ChaosINTC on an open-source RISC-V core and evaluated its performance via FPGA. Our design provides strong security with marginal hardware and performance costs. Yifan Zhu 0008, Peinan Li, Lutan Zhao, Dan Meng 0002, Rui Hou 0001 |
DAC | 3 |
| 2023 | Architecting the Autocuckoo Filter to Defend Against Cross-Core Cache AttacksabstractCross-core cache timing side-channel attacks, which observe cache access behavior of victims running on different physical cores to infer sensitive information, have become a significant threat. Although the attacks are covert, they cause the attacked cachelines to frequently migrate among cache hierarchies, rendering abnormal traffic. Based on this observation, the proposed scheme PiPoMonitor records cache-memory access traffic and prefetch suspicious lines under attack to interfere with adversaries’ probes. In pursuit of security and performance, PiPoMonitor exploits a Cuckoo filter as the recording structure and introduces two features to it: 1) autonomic deletion and 2) relocation accelerating. The former exponentially increases the uncertainty of record eviction against reverse engineering attacks, while the latter leverages a pipelined architecture to alleviate the impact of intensive filter queries on the memory critical path. PiPoMonitor is not only able to effectively mitigate cross-core cache attacks and defeat sophisticated defense-aware attackers but also induces a negligible performance penalty and acceptable hardware overhead. Fengkai Yuan, Kai Wang 0061, Jiameng Ying, Rui Hou 0001, Lutan Zhao, Peinan Li, Yifan Zhu 0008, Zhenzhou Ji, Dan Meng 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Conditional address propagation: an efficient defense mechanism against transient execution attacksabstractSpeculative execution is a critical technique in modern high performance processors. However, continuously exposed transient execution attacks, including Spectre and Meltdown, disclosed a large attack surface in mispredicted execution. Current state-of-the-art defense strategy blocks all memory accesses that use addresses loaded speculatively. However, propagation of base addresses is common in general applications and we find that more than 60% blocked memory accesses use propagated base rather than offset addresses. Therefore, we propose a novel hardware defense mechanism, named Conditional Address Propagation, to identify safe base addresses through taint tracking and address checking by a History Table. Then, the safe base addresses are allowed to be propagated to retrieve performance. For remaining unsafe addresses, they cannot be propagated for security. We constructed experiments on cycle-accurate Gem5 simulator. Compared to the representative study, STT, our mechanism effectively decreases the performance overhead from 13.27% to 1.92% targeting Spectre-type and 19.66% to 5.23% targeting all-type cache-based transient execution attacks. Peinan Li, Rui Hou 0001, Lutan Zhao, Yifan Zhu 0008, Dan Meng 0002 |
DAC | 3 |
| 2022 | HyBP: Hybrid Isolation-Randomization Secure Branch PredictorabstractRecently exposed vulnerabilities reveal the necessity to improve the security of branch predictors. Branch predictors record history about the execution of different processes, and such information from different processes are stored in the same structure and thus accessible to each other. This leaves the attackers with the opportunities for malicious training and malicious perception. Physical or logical isolation mechanisms such as using dedicated tables and flushing during context-switch can provide security but incur non-trivial costs in space and/or execution time. Randomization mechanisms incurs the performance cost in a different way: those with higher securities add latency to the critical path of the pipeline, while the simpler alternatives leave vulnerabilities to more sophisticated attacks.This paper proposes HyBP, a practical hybrid protection and effective mechanism for building secure branch predictors. The design applies the physical isolation and randomization in the right component to achieve the best of both worlds. We propose to protect the smaller tables with physically isolation based on (thread, privilege) combination; and protect the large tables with randomization. Surprisingly, the physical isolation also significantly enhances the security of the last-level tables by naturally filtering out accesses, reducing the information flow to these bigger tables. As a result, key changes can happen less frequently and be performed conveniently at context switches. Moreover, we propose a latency hiding design for a strong cipher by precomputing the "code book" with a validated, cryptographically strong cipher. Overall, our design incurs a performance penalty of 0.5% compared to 5.1% of physical isolation under the default context switching interval in Linux. Lutan Zhao, Peinan Li, Rui Hou 0001, Michael C. Huang 0001, Xuehai Qian, Lixin Zhang 0002, Dan Meng 0002 |
HPCA | 1 |
| 2022 | CPP: A lightweight memory page management extension to prevent code pointer leakageabstractProtecting code pointers (e.g., return address, function pointer) from leakage is desirable from a security perspective. Isolation mechanisms have been the favored candidate to protect code pointers. However, these mechanisms result in significant performance overhead as they need to instrument extra instructions for frequent permission switching or bound checking. In this paper, we propose CPP, a novel Code Pointer-only Memory Page Management to restrict attack-critical operations for code pointers by hardware. Our hardware–software co-design allows CPP mark code pointers at page granularity that requires minor hardware modification. CPP checks the legality of their operations in parallel with instruction execution. We implement a prototype system and our evaluation shows CPP can effectively mitigate the code pointer leakage attacks with less than 2.1% performance overhead. Jiameng Ying, Rui Hou 0001, Lutan Zhao, Fengkai Yuan, Penghui Zhao, Dan Meng 0002 |
J. Syst. Archit. | 3 |
| 2021 | A Lightweight Isolation Mechanism for Secure Branch PredictorsabstractRecently exposed vulnerabilities reveal that branch predictors shared by different processes leave the attackers with the opportunities for malicious training and perception. Instead of flush-based or physical isolation of hardware resources, we want to achieve isolation of the content in these hardware tables with some lightweight processing using randomization as follows. (1) Content encoding. We propose to use hardware-based thread-private random numbers to encode the contents of the branch predictor tables. It achieves a similar effect of logical isolation but adds little in terms of space or time overheads. (2) Index encoding. We propose a randomized index mechanism of the branch predictor. This disrupts the correspondence between the branch instruction address and the branch predictor entry, thus increases the noise for malicious perception attacks. Our analyses using an FPGA-based RISC-V processor prototype and additional auxiliary simulations suggest that the proposed mechanisms incur a very small performance cost while providing strong protection. Lutan Zhao, Peinan Li, Rui Hou 0001, Michael C. Huang 0001, Jiazhen Li, Lixin Zhang 0002, Xuehai Qian, Dan Meng 0002 |
DAC | 1 |
| 2021 | PiPoMonitor: Mitigating Cross-core Cache Attacks Using the Auto-Cuckoo FilterabstractCache side channel attacks obtain victim cache line access footprint to infer security-critical information. Among them, cross-core attacks exploiting the shared last level cache are more threatening as their simplicity to set up and high capacity. Stateful approaches of detection-based mitigation observe precise cache behaviors and protect specific cache lines that are suspected of being attacked. However, their recording structures incur large storage overhead and are vulnerable to reverse engineering attacks. Exploring the intrinsic non-determinate layout of a traditional Cuckoo filter, this paper proposes a space efficient Auto-Cuckoo filter to record access footprints, which succeed to decrease storage overhead and resist reverse engineering attacks at the same time. With Auto-Cuckoo filter, we propose PiPoMonitor to detect Ping-Pong patterns and prefetch specific cache line to interfere with adversaries' cache probes. Security analysis shows the PiPoMonitor can effectively mitigate cross-core attacks and the Auto-Cuckoo filter is immune to reverse engineering attacks. Evaluation results indicate PiPoMonitor has negligible impact on performance and the storage overhead is only 0.37%, an order of magnitude lower than previous stateful approaches. Fengkai Yuan, Kai Wang 0061, Rui Hou 0001, Peinan Li, Lutan Zhao, Jiameng Ying, Amro Awad, Dan Meng 0002 |
DATE | 6 |
| 2021 | A Novel Probabilistic Saturating Counter Design for Secure Branch Predictor
Lutan Zhao, Rui Hou 0001, Kai Wang 0061, Yu-Lan Su, Peinan Li, Dan Meng 0002 |
J. Comput. Sci. Technol. | 1 |
| 2021 | Exploiting Security Dependence for Conditional Speculation Against Spectre AttacksabstractSpeculative execution side-channel vulnerabilities such as Spectre reveal that conventional architecture designs lack security consideration. This article proposes a software transparent defense framework, named as Conditional Speculation, against Spectre vulnerabilities found on traditional out-of-order microprocessors. It introduces the concept of security dependence to mark speculative memory instructions which could leak information with potential security risks. More specifically, security-dependent instructions are detected and marked with suspect speculation flags in the Issue Queue. All the instructions can be speculatively issued for execution in accordance with the classic out-of-order pipeline. For those instructions with suspect speculation flags, they are considered as safe instructions if their speculative execution dose not refill new cache lines with unauthorized privilege data. Otherwise, they are considered as unsafe instructions and thus not allowed to execute speculatively. To pursue a balance of performance and security, we investigate two filtering mechanisms, Cache-hit-based Hazard Filter and Trusted Page Buffer-based Hazard Filter to filter out false security hazards. As for true security hazards, we have two approaches to prevent them from changing cache states. One is to block all unsafe access, the other is to fetch them from lower-level caches or memory to a speculative buffer temporarily, and refill them after confirming that they are on the correct execution path. Our design philosophy is to speculatively execute safe instructions to maintain the performance benefits of out-of-order execution while delaying the cache updates for speculative execution of unsafe instructions for security consideration. We evaluate Conditional Speculation in terms of performance, security, and area. The experimental results show that the hardware overhead is marginal and the performance overhead is minimal. Lutan Zhao, Peinan Li, Rui Hou 0001, Michael C. Huang 0001, Peng Liu 0005, Lixin Zhang 0002, Dan Meng 0002 |
IEEE Trans. Computers | 1 |
| 2021 | Mitigating Cross-Core Cache Attacks via Suspicious Traffic DetectionabstractContinuous Attacks are common cross-core cache side-channel attack scenarios that we observed, where adversaries frequently probe-target cache lines in a short time. Under Continuous Attacks, the attacked lines go through multiple load-evict processes between different cache (or memory) hierarchies, exhibiting Ping-Pong patterns. Identifying and obscuring these abnormal patterns effectively interfere with the attacker's probe and mitigate such attacks. Our recent proposal, Ping-Pong regulator (PPR), captures multiple Ping-Pong patterns by counting the reaccesses per cache line and blocks them with different obscuring actions (preload or lock). Although PPR mitigates Continuous Attacks, the added regulator directory (RDir) is vulnerable because it cannot record all cache lines simultaneously. Sophisticated attackers can evict the records of the attacked line from the RDir to avoid triggering defensive actions, thereby bypassing PPR. To improve robustness, we further propose PPR+, which dynamically changes the mapping of physical addresses to RDir locations by encryption and periodically changing keys. This randomness makes it difficult for attackers to evict target entries out of the RDir within a limited time. We show that PPR+ tolerates more than 100 years of attacks, induces negligible performance impacts (improves 0.13%), requires acceptable storage overhead (3.15%), and does not need any software support. Kai Wang 0061, Fengkai Yuan, Lutan Zhao, Rui Hou 0001, Zhenzhou Ji, Dan Meng 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | Conditional Speculation: An Effective Approach to Safeguard Out-of-Order Execution Against Spectre AttacksabstractSpeculative execution side-channel vulnerabilities such as Spectre reveal that conventional architecture designs lack security consideration. This paper proposes a software transparent defense mechanism, named as Conditional Speculation, against Spectre vulnerabilities found on traditional out-of-order microprocessors. It introduces the concept of security dependence to mark speculative memory instructions which could leak information with potential security risk. More specifically, security-dependent instructions are detected and marked with suspect speculation flags in the Issue Queue. All the instructions can be speculatively issued for execution in accordance with the classic out-of-order pipeline. For those instructions with suspect speculation flags, they are considered as safe instructions if their speculative execution will not refill new cache lines with unauthorized privilege data. Otherwise, they are considered as unsafe instructions and thus not allowed to execute speculatively. To reduce the performance impact from not executing unsafe instructions speculatively, we investigate two filtering mechanisms, Cachehit based Hazard Filter and Trusted Page Buffer based Hazard Filter to filter out false security hazards. Our design philosophy is to speculatively execute safe instructions to maintain the performance benefits of out-of-order execution while blocking the speculative execution of unsafe instructions for security consideration. We evaluate Conditional Speculation in terms of performance, security and area. The experimental results show that the hardware overhead is marginal and the performance overhead is minimal. Peinan Li, Lutan Zhao, Rui Hou 0001, Lixin Zhang 0002, Dan Meng 0002 |
HPCA | 2 |