Pengfei Qiu

dblp:170/8002 · DBLP profile ↗
← Back
22ranked-venue papers
9as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 7 first-author · 10 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploiting ARMeD Channels By Reverse Engineering ARM Memory Disambiguation Unit
abstract
ARM CPUs are widely used in both embedded systems and personal computers where security considerations are becoming important. Evidently, vulnerabilities on hardware components such as cache and translation look-aside buffer are well-documented. But there are much less studies on other components, especially those in the CPU backend, largely due to the unavailability of their design and implementation details. To address this gap, we present the first in-depth reverse engineering analysis of the Memory Disambiguation Unit (MDU) in the backend of ARM CPUs. Across four microarchitectures from ARM and Apple CPUs, we identify two different MDU designs, switch-based and counter-based. We then analyze the state machine, selection mechanism, and organization of these MDU designs. We further propose new side channels and covert channels, which we call ARMeD channels, that exploit ARM MDU to leak information. We demonstrate with three attacks using ARMeD channels: a cross-process covert channel, website fingerprinting, and a new implementation of the Spectre attack. Finally, we present a defense strategy against ARMeD Channels with less than 3% degradation on the MDU’s prediction accuracy.
Chang Liu 0117, Zhouyang Li, Haixia Wang 0001, Pengfei Qiu, Gang Qu 0001, Dongsheng Wang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 GhostCache: Timer- and Counter-Free Cache Attacks Exploiting Weak Coherence on RISC-V and ARM Chips
abstract
Microarchitectural side-channel attacks, which have become increasingly prevalent, often rely on high-resolution timers. Emerging processor architectures have sought to mitigate these vulnerabilities by restricting access to fine-grained timers. In this work, we verify the widespread existence of weak coherence in L1 cache on multiple RISC chips, exploit it to bypass this type of mitigation and propose GhostCache, which constructs timer-free and counter-free instruction cache attacks. It introduces two novel and widely applied attack primitives, Modify+Recall and Call+ModifyCall, which are applicable to both RISC-V and ARM architectures and affect 6 commercial and 3 open-source large RISC processors. To the best of our knowledge, we present the first demonstration of timer-free and counter-free cache attacks on RISC-V processors. We also identify undisclosed features, such as the next-three-line prefetching mechanism and direct forwarding of evicted instructions from data cache to instruction cache. Furthermore, we develop four types of covert channels, achieving up to 1.68 MB/s with a 0.01% error rate. For side-channel attacks, GhostCache enables three types of timer-free real-world attacks. The first is an end-to-end website fingerprinting attack, achieving 92.02% accuracy across 100 website classes. The second is a set of kernel leakage attacks, including the discovery of a new Spectre disclosure gadget via a function pointer to leak arbitrary kernel data at 92.91% accuracy. We also launched an attack to reconstruct cryptographic keys. Lastly, we propose potential countermeasures to address these vulnerabilities in both RISC-V and ARM architectures.
Yu Jin 0010, Minghong Sun, Dongsheng Wang 0002, Pengfei Qiu, Yinqian Zhang, Shuwen Deng
CCS4
2025 ChronosRV: Online Runtime Monitoring and Code Generation for Bounded Temporal Specifications in Low-Latency C++ Trading Systems
Pengfei Qiu
SETTA1
2025 Auto-GAN: GAN-Based Self-Supervised Collaborative Learning for Robust Spatio-Temporal Trajectory Classification in IoT
abstract
With the rapid proliferation of crowd mobility data produced by ubiquitous mobile devices equipped with spatial positioning modules, deep neural networks (DNNs) have become widely applied in spatio-temporal trajectory modeling. However, recent studies have shown that DNNs are vulnerable to adversarial examples with strong transferability, which are crafted by introducing small perturbations to original examples but can cause catastrophic mistakes. To mitigate this vulnerability and enhance model robustness, we propose a novel self-supervised collaborative learning framework named Auto-GAN that consists of a generator for automatically learning robust latent features and a discriminator for providing comprehensive guidance to the generator. By leveraging the collaboration between the generator and discriminator, our proposed method significantly improves the denoising performance. Moreover, we combine point-level and feature-level constraints into training processes between original example reconstruction and adversarial example denoising, thereby effectively suppressing the potential “error amplification effect". Extensive experiments conducted on two representative real-world mobility datasets show that our proposed method can significantly enhance the model’s robustness against various adversarial attacks, while preserving the model’s prediction accuracy on original examples.
Jia Jia 0007, Linghui Li 0001, Ximing Li 0005, Binsi Cai, Xu Zhang 0006, Pengfei Qiu
IEEE Internet Things J.7
2025 High-Quality Trajectory Generation via Domain-Knowledge Enhanced GANs
abstract
Simulating human mobility realistically and generating large-scale, high-quality trajectories are crucial for various location-based applications such as traffic management, epidemic spreading analysis, and location privacy protection. While the most popular model-free methods succeed by directly learning distribution of real-world data, they struggle to produce high-quality mobility data without leveraging the domain knowledge of human mobility. Moreover, such model-free methods primarily rely on auto-regressive paradigms, usually accompanied by error accumulation problem. To address the issues, we propose a model-free Domain-Knowledge Enhanced Generative Adversarial Network (DKE-GAN), which efficiently combines domain knowledge of urban context with model-free learning paradigm to generate high-quality mobility data. In addition, we incorporate reinforcement learning into the training process, thereby effectively alleviating the error accumulation. Furthermore, we introduce Trajectory Representation Learning (TRL) to convert noise-carrying raw trajectories into low-dimensional representation vectors for fully mining human mobility patterns. Extensive experiments conducted on two representative real-world mobility datasets demonstrate that our proposed method outperforms six state-of-the-art baselines, significantly achieving performance improvements in simulating human mobility.
Jia Jia 0007, Ximing Li 0005, Binsi Cai, Xu Kang 0001, Xu Zhang 0006, Pengfei Qiu
IEEE Internet Things J.7
2024 Whisper: Timing the Transient Execution to Leak Secrets and Break KASLR
abstract
The vulnerabilities of transient execution have been exploited in many side-channel attacks (SCA). We report Whisper, a novel transient execution timing (TET) side channel, which is based on the execution time difference of transient execution under different conditions. We develop TET version of SCAs including Meltdown, Zombieload, and Spectre-RSB that use Whisper as covert channel to leak information. We further propose TET-KASLR to break the kernel address space layout randomization (KASLR) mechanism under the protection of KPTI and FLARE. These attacks are simple to implement and can bypass the existing mitigation methods because the TET side channel relies on execution time that can be conveniently obtained by architectural level timing analysis. We demonstrate the correctness and effectiveness of these attacks on various x86-64 CPUs. The root cause of Whisper is analyzed with our toolset built on performance monitor unit (PMU) and potential defense against Whisper is also discussed.
Yu Jin 0010, Chunlu Wang, Pengfei Qiu, Chang Liu 0117, Hongpei Zheng, Yongqiang Lyu 0001, Xiaoyong Li 0003, Gang Qu 0001, Dongsheng Wang 0002
DAC3
2024 Uncovering and Exploiting AMD Speculative Memory Access Predictors for Fun and Profit
abstract
This paper presents a comprehensive investigation into the security vulnerabilities associated with speculative memory access on AMD processors. Firstly, employing novel reverse engineering techniques, our study uncovers two key predictors, namely the Predictive Store Forwarding Predictor (PSFP) and the Speculative Store Bypass Predictor (SSBP), along with elucidating their internal structures and state machine designs. Secondly, our research empirically confirms that these predictors can be deliberately manipulated and altered during transient execution, resulting in secret leakage across security domains. Leveraging these discoveries, we propose innovative attacks targeting these predictors, including an out-of-place variant of Spectre-STL and an entirely new form of Spectre attack named Spectre-CTL. Finally, we establish experimentally that enabling Speculative Store Bypass Disable alleviates the vulnerabilities. However, this comes at the expense of significant performance degradation.
Chang Liu 0117, Dongsheng Wang 0002, Yongqiang Lyu 0001, Pengfei Qiu, Yu Jin 0010, Zhuoyuan Lu, Yinqian Zhang, Gang Qu 0001
HPCA4
2024 Domain-Knowledge Enhanced GANs for High-Quality Trajectory Generation
Jia Jia 0007, Linghui Li 0001, Pengfei Qiu, Binsi Cai, Xu Kang 0001, Ximing Li 0005, Xiaoyong Li 0003
ICIC (9)3
2024 SCAFinder: Formal Verification of Cache Fine-Grained Features for Side Channel Detection
abstract
Recent research has unveiled numerous cache-timing side-channel attacks exploiting the side effects of fine-grained cache features, such as coherence protocol and prefetch, among others. Traditional modeling methods and verification techniques are insufficient for verifying caches with fine-grained features and detecting cache timing vulnerabilities. There is a necessity for comprehensive verification of such complex cache designs. This paper presents SCAFinder, a verification framework targeting the cache designs with fine-grained features; it identifies cache side-channel attacks through model checking techniques. Specifically, it proposes a modeling methodology for cache designs that enables us to abstract the cache’s behavior and latency characteristics. We implement a search algorithm for finding all counterexamples based on open-source model checking software. Subsequently, we add an attack scenario analysis module to discover attacks applicable to specific scenarios. We evaluate SCAFinder on Intel Skylake-X microarchitecture, demonstrating its capability to generate 7 new attack sequences exploiting coherence protocol and prefetch, and 12 new replacement policy-based side channels. As a case study, we successfully built a covert channel for one of the sequences on the real-world processor. To the best of our knowledge, we are the first to implement cross-core replacement policy-based attacks on non-inclusive caches.
Haixia Wang 0001, Pengfei Qiu, Yongqiang Lyu 0001, Hongpeng Wang 0002, Dongsheng Wang 0002
IEEE Trans. Inf. Forensics Secur.3
2024 Lightning: Leveraging DVFS-induced Transient Fault Injection to Attack Deep Learning Accelerator of GPUs
abstract
Graphics Processing Units (GPU) are widely used as deep learning accelerators because of its high performance and low power consumption. Additionally, it remains secure against hardware-induced transient fault injection attacks, a classic type of attacks that have been developed on other computing platforms. In this work, we demonstrate that well-trained machine learning models are robust against hardware fault injection attacks when the faults are generated randomly. However, we discover that these models have components, which we refer to as sensitive targets, that are vulnerable to faults. By exploiting this vulnerability, we propose the Lightning attack, which precisely strikes the model’s sensitive targets with hardware-induced transient faults based on the Dynamic Voltage and Frequency Scaling (DVFS). We design a sensitive targets search algorithm to find the most critical processing units of Deep Neural Network (DNN) models determining the inference results, and develop a genetic algorithm to automatically optimize the attack parameters for DVFS to induce faults. Experiments on three commodity Nvidia GPUs for four widely-used DNN models show that the proposed Lightning attack can reduce the inference accuracy by 69.1% on average for non-targeted attacks, and, more interestingly, achieve a success rate of 67.9% for targeted attacks.
Rihui Sun, Pengfei Qiu, Yongqiang Lyu 0001, Jian Dong 0010, Haixia Wang 0001, Dongsheng Wang 0002, Gang Qu 0001
ACM Trans. Design Autom. Electr. Syst.2
2023 PMU-Leaker: Performance Monitor Unit-Based Realization of Cache Side-Channel Attacks
abstract
Performance Monitor Unit (PMU) is a special hardware module in processors that contains a set of counters to record various architectural and micro-architectural events. In this paper, we propose PMU-Leaker, a novel realization of all existing cache side-channel attacks where accurate execution time measurements are replaced by information leaked through PMU. The efficacy of PMU-Leaker is demonstrated by (1) leaking the secret data stored in Intel Software Guard Extensions (SGX) with the transient execution vulnerabilities including Spectre and ZombieLoad and (2) extracting the encryption key of a victim AES performed in SGX. We perform thorough experiments on a DELL Inspiron 15-7560 laptop that has an Intel® Core™ i5-7200U processor with the Kaby Lake architecture and the results show that, among the 176 PMU counters, 24 of them are vulnerable and can be used to launch the PMU-Leaker attack.
Pengfei Qiu, Dongsheng Wang 0002, Yongqiang Lyu 0001, Chunlu Wang, Chang Liu 0117, Rihui Sun, Gang Qu 0001
ASP-DAC1
2023 Leaky MDU: ARM Memory Disambiguation Unit Uncovered and Vulnerabilities Exposed
abstract
Memory Disambiguation Unit (MDU) is widely used on modern processors to speculatively execute load instructions and improve pipeline performance. Given that the MDU design details on ARM processors are not available to the public, it is unclear whether there are any security vulnerabilities associated with its MDU. In this paper, we first reverse engineer the undocumented features of ARM MDU, then we discover three potential user-privilege attacks to leak secret data via MDU: cross-process attack that allows users to communicate through a convert channel, cross-domain attack that leaks kernel information and a new variant of inner-process and inter-processes Spectre attacks. These attacks pose serious security challenges as they can bypass both all the known countermeasures against cache side-channel attacks and those against transient execution attacks. Potential mitigation against the proposed MDU-based attacks are also discussed.
Chang Liu 0117, Yongqiang Lyu 0001, Haixia Wang 0001, Pengfei Qiu, Dapeng Ju, Gang Qu 0001, Dongsheng Wang 0002
DAC4
2023 Exploration and Exploitation of Hidden PMU Events
abstract
Performance Monitoring Unit (PMU) is a common hardware module in modern processors that monitors the processor's architectural and microarchitectural events (PMU events) for CPU performance analysis and optimization. Vendors publish PMU events in documents such as Intel's Software Development Manual (SDM) and ARM processor technical reference manuals. In this paper, we report our findings that these documented PMU events are only a very small portion of the PMU event space. We define hidden PMU events as those that can be triggered in the instruction's execution but are not documented by the vendors. The hidden PMU events may not be as useful as the documented ones for CPU performance analysis. However, they might introduce security vulnerabilities. We develop an automated tool to traverse all the possible PMU events during the execution of each valid instruction to locate the hidden PMU events. On six Intel processors with different micro-architectures, where there are about 307 documented PMU core events on average, our tool finds an average of 17,361 hidden PMU events. We further demonstrate the security implications in both defense and attack of these hidden PMU events. Our experimental results show that up to 6,613 hidden PMU events on the i7-6700 can be used to detect transient execution attacks and 1,192 hidden PMU events can be exploited for side-channel attacks.
Pengfei Qiu, Chunlu Wang, Yu Jin 0010, Xiaoyong Li 0003, Dongsheng Wang 0002, Gang Qu 0001
ICCAD2
2023 PMU-Spill: A New Side Channel for Transient Execution Attacks
abstract
Performance Monitor Unit (PMU) is an important hardware module in mainstream processors, which counts various architectural and microarchitectural events during the run-time of the processor. Theoretically, if an instruction is executed but doesn’t successfully retire (this is called transient execution), the events it triggers needn’t be recorded by PMU. However, in this study, we discover that current PMU implementations are capable of recording some events that are triggered in transient executions, which is a hardware vulnerability. Based on this vulnerability, we propose the PMU-Spill attack, a new kind of side channel attack that enables attackers to maliciously leak secret data in transient executions. We perform a thorough study of PMU counters on five Intel processors and find that they all have vulnerable PMU counters that will measure transient execution events (there are 162 vulnerable PMU counters among all the 383 PMU counters). We demonstrate on real hardware that 112 vulnerable PMU counters can be utilized in PMU-Spill attack to leak the secret data protected by Intel Software Guard Extensions (SGX). Besides, our experiments suggest that the throughput of PMU-Spill attack is up to 291.2 bytes per second (Bps) with an error rate of 2.45% on average. This discovery and the corresponding mitigation methods can be helpful for microarchitecture designers to reevaluate the security risks induced by the PMU module.
Pengfei Qiu, Chang Liu 0117, Dongsheng Wang 0002, Yongqiang Lyu 0001, Xiaoyong Li 0003, Chunlu Wang, Gang Qu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 DVFSspy: Using Dynamic Voltage and Frequency Scaling as a Covert Channel for Multiple Procedures
abstract
Dynamic Voltage and Frequency Scaling (DVFS) is a widely deployed low-power technology in modern systems. In this paper, we discover a vulnerability in the implementation of the DVFS technology that allows us to measure the processor's frequency in the userspace. By exploiting this vulnerability, we successfully implement a covert channel on the commercial Intel platform and demonstrate that the covert channel can reach a throughput of 28.41bps with an error rate of 0.53%. This work indicates that the processor's hardware information that is unintentionally leaked to the userspace by the privileged kernel modules may cause security risks.
Pengfei Qiu, Dongsheng Wang 0002, Yongqiang Lyu 0001, Gang Qu 0001
ASP-DAC1
2022 Auto-Encoding GAN for Reducing Mode Collapse and Enhancing Feature Representation
abstract
Generative Adversarial Nets (GAN) has been a popular research topic in processing of images, speech, texts, and videos, and many other fields.However, GAN still has some drawbacks such as unstable training and mode collapse.To address these challenges, this paper proposes an auto-encoding GAN, which is composed of a set of generators, a discriminator, an encoder and a decoder.A set of generators is responsible for learning different modes, accelerating the convergence of the model and preventing model collapse.The discriminator is used to distinguish between real samples and generated ones.In order to improve feature representation of the encoder and prevent multiple generators from covering a certain mode, an approach consisting of three phases is proposed accordingly.First, a clustering algorithm is presented to perceive the distribution of real and generated samples.Then, cluster center matching is utilized to keep consistency of the distribution of real and generated samples.Finally, the encoder and decoder are jointly optimized by the generated and real samples.Therefore, the encoder can map the generated and real samples to the embedding space so as to encode distinguishable features, and the decoder can distinguish from which generator the generated samples come and from which mode the real samples come.Experiments are conducted on image datasets to verify effectiveness of the auto-encoding GAN for reducing mode collapse and enhancing feature representation.
Xiaoxiang Lu, Yang Zou 0001, Xiaoqin Zeng, Xiangchen Wu, Pengfei Qiu
SEKE5
2021 VoltJockey: A New Dynamic Voltage Scaling-Based Fault Injection Attack on Intel SGX
abstract
Intel software guard extensions (SGX) increase the security of applications by enabling them to be performed in a highly trusted space (called enclave). Most state-of-the-art attacks on SGX focus on either mining the software vulnerabilities in the enclave or speculating the secret data with side channels. In this study, we report our recent work on breaking SGX by inducing voltage-oriented hardware faults. The novelty and importance of this attack are that it is completely controlled by software and does not require any security vulnerability in the software. Our proposed attack, called VoltJockey, exploits a vulnerability in the implementation of dynamic voltage and frequency scaling (DVFS) that achieves energy saving by dynamically adjusting the processor's operating voltage and thus clock frequency. However, if the operating voltage is lower than a certain critical level, the circuit's timing constraint will fail and hardware fault would be created. We propose to deliberately trigger such voltage-oriented hardware faults by a loadable kernel module that can set the processor's voltage through Intel's undocumented model-specific register (MSR). We first utilize the module to furnish the processor with a transient low voltage with controlled timing to inject a temporal fault into the target location of the program running in the enclave. Then, we perform a differential fault attack on the outputs before and after the injection of faults. For demonstration, we successfully deploy the proposed attack to extract the key of an AES executed in the enclave and lead an SGX-protected RSA to output our specified result.
Pengfei Qiu, Dongsheng Wang 0002, Yongqiang Lyu 0001, Ruidong Tian, Chunlu Wang, Gang Qu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 Mitigating Adversarial Attacks for Deep Neural Networks by Input Deformation and Augmentation
abstract
Typical Deep Neural Networks (DNN) are susceptible to adversarial attacks that add malicious perturbations to input to mislead the DNN model. Most of the state-of-theart countermeasures concentrate on the defensive distillation or parameter re-training, which require prior knowledge of the target DNN and/or the attacking methods and hence greatly limit their generality and usability. In this paper, we propose to defend against adversarial attacks by utilizing the input deformation and augmentation techniques that are currently widely utilized to enlarge the dataset during DNN's training phase. This is based on the observation that certain input deformation and augmentation methods will have little or no impact on DNN model's accuracy, but the adversarial attacks will fail when the maliciously induced perturbations are randomly deformed. We also use the ensemble of decisions to further improve DNN model's accuracy and the effectiveness of defending various attacks. Our proposed mitigation method is model independent (i.e. it does not require additional training, parameter finetuning, or any structure modifications of the target DNN model) and attack independent (i.e., it does not require any knowledge of the adversarial attacks). So it has excellent generality and usability. We conduct experiments on standard CIFAR-10 dataset and three representative adversarial attacks: Fast Gradient Sign Method, Carlini and Wagner, and Jacobian-based Saliency Map Attack. Results show that the average success rate of the attacks can be reduced from 96.5% to 28.7% while the DNN model accuracy is improved by about 2%.
Pengfei Qiu, Qian Wang 0022, Dongsheng Wang 0002, Yongqiang Lyu 0001, Zhaojun Lu, Gang Qu 0001
ASP-DAC1
2019 VoltJockey: Breaching TrustZone by Software-Controlled Voltage Manipulation over Multi-core Frequencies
abstract
ARM TrustZone builds a trusted execution environment based on the concept of hardware separation. It has been quite successful in defending against various software attacks and forcing attackers to explore vulnerabilities in interface designs and side channels. The recently reported CLKscrew attack breaks TrustZone through software by overclocking CPU to generate hardware faults. However, overclocking makes the processor run at a very high frequency, which is relatively easy to detect and prevent, for example by hardware frequency locking. In this paper, we propose an innovative software-controlled hardware fault-based attack, VoltJockey, on multi-core processors that adopt dynamic voltage and frequency scaling (DVFS) techniques for energy efficiency. Unlike CLKscrew, we manipulate the voltages rather than the frequencies via DVFS unit to generate hardware faults on the victim cores, which makes VoltJockey stealthier and harder to prevent than CLKscrew. We deliberately control the fault generation to facilitate differential fault analysis to break TrustZone. The entire attack process is based on software without any involvement of hardware. We implement VoltJockey on an ARM-based Krait processor from a commodity Android phone and demonstrate how to reveal the AES key from TrustZone and how to breach the RSA-based TrustZone authentication. These results suggest that VoltJockey has a comparable efficiency to side channels in obtaining TrustZone-guarded credentials, as well as the potential of bypassing the RSA-based verification to load untrusted applications into TrustZone. We also discuss both hardware-based and software-based countermeasures and their limitations.
Pengfei Qiu, Dongsheng Wang 0002, Yongqiang Lyu 0001, Gang Qu 0001
CCS1
2019 Photoplethysmogram-based Cognitive Load Assessment Using Multi-Feature Fusion Model
abstract
Cognitive load assessment is crucial for user studies and human--computer interaction designs. As a noninvasive and easy-to-use category of measures, current photoplethysmogram- (PPG) based assessment methods rely on single or small-scale predefined features to recognize responses induced by people’s cognitive load, which are not stable in assessment accuracy. In this study, we propose a machine-learning method by using 46 kinds of PPG features together to improve the measurement accuracy for cognitive load. We test the method on 16 participants through the classical n-back tasks (0-back, 1-back, and 2-back). The accuracy of the machine-learning method in differentiating different levels of cognitive loads induced by task difficulties can reach 100% in 0-back vs. 2-back tasks, which outperformed the traditional HRV-based and single-PPG-feature-based methods by 12--55%. When using “leave-one-participant-out” subject-independent cross validation, 87.5% binary classification accuracy was reached, which is at the state-of-the-art level. The proposed method can also support real-time cognitive load assessment by beat-to-beat classifications with better performance than the traditional single-feature-based real-time evaluation method.
Xiao Zhang 0008, Yongqiang Lyu 0001, Tong Qu, Pengfei Qiu, Xiaomin Luo, Shunjie Fan, Yuanchun Shi
ACM Trans. Appl. Percept.4
2018 Control Flow Integrity Based on Lightweight Encryption Architecture
abstract
Control-flow integrity (CFI) plays a very important role in defending against code reuse attacks by protecting the control flows of programs from being hijacked. However, previous CFI methods suffer from performance overheads, cost, or security issues. In this paper, we propose a new CFI based on a lightweight encryption architecture with advanced encryption standard (LEA-AES) to address the challenges above. The LEA exploits AES to encrypt and decrypt return addresses and instructions at indirect jump destinations, which protects function calls and indirect jumps from being reused by return-oriented programming (ROP) and jump-oriented programming (JOP) attacks. For ROP, the encryption and decryption of return addresses are performed when the call and ret instructions are executing; for JOP, the encryption of instructions are performed when programs are loading into memory and the decryption of instructions are performed right before they are executing. The LEA-AES does not need to revise instruction sets of CPU and its security is also guaranteed by the encryption mechanism in addition to its high performance. Experimental results showed that the run-time and loading time overheads of LEA-AES are both less than 4% and the memory overhead is 0.62%.
Pengfei Qiu, Yongqiang Lyu 0001, Jiliang Zhang 0002, Dongsheng Wang 0002, Gang Qu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2016 Physical unclonable functions-based linear encryption against code reuse attacks
abstract
Recently, code reuse attacks (CRAs) have emerged as a new class of ingenious security threatens. Attackers can utilize CRAs to hijack the control flow of programs to perform malicious actions without injecting any codes. Existing defenses against CRAs often incur high memory and performance overheads or require extending the existing processors' instruction set architectures (ISAs). To tackle these issues, we propose a hardware-based control flow integrity (CFI) that employs physical unclonable functions (PUF)-based linear encryption architecture (LEA) to protect against CRAs with negligible hardware extending and run time overheads. The proposed method can protect ret and indirect jmp instructions from return oriented programming (ROP) and jump oriented programming (JOP) without any additional software manipulations and extending ISAs. The pre-process will be conducted on codes once the executable binary is loaded into memory, and the real-time control flow verification based on LEA can be done while ret and jmp instructions are executed. Performance evaluations on benchmarks show that the proposed method only introduces 0.61% run-time overhead and 0.63% memory overhead on average.
Pengfei Qiu, Yongqiang Lyu 0001, Jiliang Zhang 0002, Xingwei Wang 0001, Di Zhai, Dongsheng Wang 0002, Gang Qu 0001
DAC1