Zhenkai Zhang 0002

dblp:117/0017-2 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0002-7934-7773ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 12 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploiting TLBs in Virtualized GPUs for Cross-VM Side-Channel Attacks
Hongyue Jin, Yanan Guo 0002, Zhenkai Zhang 0002
NDSS3
2026 GeForge: Hammering GDDR Memory to Forge GPU Page Tables for Fun and Profit
abstract
Over the years, Rowhammer has been leveraged to mount a wide range of attacks against system main memory. While a recent study has revealed that GPU memory is similarly vulnerable, the security implications remain largely under-explored. To advance this line of research, we introduce GeForge, an end-to-end Rowhammer attack that exploits bit flips induced in GPU memory to achieve system-level compromise. At its core, GeForge corrupts GPU page tables to seize control of address translation, enabling arbitrary access to the entire GPU memory. Moreover, by exploiting a special mapping feature in the GPU page table, GeForge extends its reach to directly access host memory. To make GeForge practical under default system settings, we develop novel techniques that eliminate restrictive assumptions in prior work. Our techniques include a method for aligning offline-profiled physical address mappings to runtime GPU allocations and a memory massaging strategy that steers target GPU page table structures into vulnerable locations via the stock driver allocator. In addition, we improve the hammering pattern to trigger many more bit flips than prior work. With these approaches, we successfully mount GeForge on widely deployed NVIDIA GPUs, including both workstationclass and consumer-grade ones. We show that GeForge allows an attacker to arbitrarily read and modify data across GPU contexts. More crucially, we demonstrate that GeForge can help the attacker escalate privileges to root on the host system.
Junpeng Wan, Yanan Guo 0002, Zhi Zhang 0001, Jing (Dave) Tian, Zhenkai Zhang 0002
SP6
2024 PowSpectre: Powering Up Speculation Attacks with TSX-based Replay
abstract
Trusted execution environment (TEE) offers data protection against malicious system software. However, the TEE (e.g., Intel SGX) threat model exacerbates information leakage as attackers can enhance and denoise the observations from hardware-based side channels through controlled victim execution (i.e., replay). The replay mechanism is especially critical for side channels from physical traces (e.g., power consumption) that not only vary instantaneously but also necessitate successive modulation for observability. In this paper, we identify and characterize the key limitations of existing replay techniques for speculation attacks. Our study unveils that architectural support for transactional memory (i.e., Intel TSX) can be leveraged as a highly efficient replay primitive for transient execution. Based on this observation, we design TMPlayer, an efficient and high-resolution TSX-based replay framework for enclave victims. Built on top of TMPlayer, we present PowSpectre- a novel replay-based transient execution attack using software-based power side channels (via RAPL) that can exfiltrate secretive enclave data accurately in the speculative domain. We evaluate PowSpectre using case studies on several representative SGX binaries. Our evaluation shows that PowSpectre can exfiltrate unintended secrets in enclaves with very high accuracy. We perform a gadget analysis in SGX libraries and identify widely-existing code patterns that are power differentiable for PowSpectre. Our work highlights the need to synergistically understand the impact of speculation security with the introduction of new hardware functionalities.
Md Hafizul Islam Chowdhuryy, Zhenkai Zhang 0002, Fan Yao 0001
AsiaCCS2
2024 Toward Understanding the Security of Plugins in Continuous Integration Services
abstract
Mainstream Continuous Integration (CI) platforms have provided the plugin functionality to accelerate the development of CI pipelines. Unfortunately, CI plugins, which are essentially reusable code snippets, also expose new attack surfaces as plugins might be developed by less trusted users. In this paper, we present an in-depth study to understand potential security risks in existing CI plugins. We conduct a comprehensive analysis of plugin implementations on four mainstream CI platforms (GitHub Actions, GitLab CI, CircleCI, and Azure Pipelines), and investigate several weak links in existing plugin distributions and isolation mechanisms. We investigate seven attack vectors that can enable attackers to hijack plugins and distribute malicious code without plugins users being aware, and further exploit hijacked plugins to manipulate the workflow execution. Additionally, we find that plugin dependency (a plugin references other plugins) might further amplify the attack impact of our disclosed attacks. To evaluate the potential impact, we conduct a large-scale measurement on GitHub and GitLab, covering a total of 1,328,912 repositories using the aforementioned CI platforms. Our measurement results show that a large number of repositories and existing plugins, including many widely used ones, are potentially vulnerable to the proposed attacks. We have duly reported the identified vulnerabilities and received positive responses.
Xiaofan Li 0009, Yacong Gu, Chu Qiao, Zhenkai Zhang 0002, Daiping Liu, Lingyun Ying, Hai-Xin Duan, Xing Gao 0001
CCS4
2024 WBP: Training-Time Backdoor Attacks Through Hardware-Based Weight Bit Poisoning
Kunbei Cai, Zhenkai Zhang 0002, Qian Lou, Fan Yao 0001
ECCV (65)2
2024 FreeEM: Uncovering Parallel Memory EMR Covert Communication in Volatile Environments
abstract
Memory Electromagnetic Radiation (EMR) allows attackers to manipulate the DRAM of infiltrated systems to leak sensitive secret information. Although most of the existing works have demonstrated its feasibility, practical concerns, such as the ideal electromagnetic environment and stationary attacking layout, make the covert channel attack less convincing, especially in vulnerable sites such as offices and data centers. This work removes the above impractical assumptions to uncover the potential of memory EMR by proposing the first parallel EMR covert communication protocol. Our design reshapes the current "1-to-1" covert communication mode to "n-to-1" mode via a novel pattern-based 2-dimensional symbol encoding scheme, allowing multiple victim computers to simultaneously perform data exfiltration to one attacker (the receiver) without mutual interference. Meanwhile, this novel scheme design also enables the very first mobile attacker, i.e., a smartphone connected to a software-defined radio (SDR) dongle, to capture parallel memory EMR signals in a volatile environment. Extensive experiments are conducted to verify the performance in a volatile environment with different parameter configurations, distances, motion modes, shielding materials, orientations, hardware configurations, and SDR platforms. Our experimental results demonstrate that FreeEM can support up to 4 parallel memory EMR transmissions to achieve an overall throughput of 625Kbps and a decoding accuracy of 96.88%. The maximum communication distance can reach up to 20 meters.
Sihan Yu, Jingjing Fu, Chenxu Jiang, ChunChih Lin, Zhenkai Zhang 0002, Long Cheng 0005, Ming Li 0006, Xiaonan Zhang 0001, Linke Guo
MobiSys5
2024 DeepVenom: Persistent DNN Backdoors Exploiting Transient Weight Perturbations in Memories
abstract
Backdoor attacks have raised significant concerns in machine learning (ML) systems. Mainstream ML backdoor attacks typically involve either poisoning the victim’s training samples or pre-training poisoned models for use by victim users. Meanwhile, recent advances in hardware-based threats reveal that ML model integrity at inference-time can be seriously tampered by inducing transient faults in model weights. However, the adversarial impacts of such hardware fault attacks at training time have not been well understood.In this paper, we present DeepVenom, the first end-to-end hardware-based DNN backdoor attack during victim model training. Particularly, DeepVenom can insert a targeted backdoor persistently at the victim model fine-tuning runtime through transient faults in model weight memory (via rowhammer). DeepVenom manifests in two main steps: i) an offline step that identifies weight perturbation transferable to the victim model using an ensemble-based local model bit search algorithm, and ii) an online stage that integrates advanced system-level techniques to efficiently massage weight tensors for precise rowhammer-based bit flips. DeepVenom further employs a novel iterative backdoor boosting mechanism that performs multiple rounds of weight perturbations to stabilize the backdoor. We implement an end-to-end DeepVenom attack in real systems with DDR3/DDR4 memories, and evaluate it using state-of-the-art Convolutional Neural Network and Vision Transformer models. The results show that DeepVenom can effectively generate backdoors in victim’s fine-tuned models with upto 99.8% attack success rate (97.8% on average) using as few as 11 total weight bit flips (maximum 49). The evaluation further demonstrates that DeepVenom is successful under varying victim fine-tuning hyperparameter settings, and is highly robust against catastrophic forgetting. Our work highlights the practicality of training-time backdoors through hardware-based weight perturbation, which represents a new dimension in adversarial machine learning.
Kunbei Cai, Md Hafizul Islam Chowdhuryy, Zhenkai Zhang 0002, Fan Yao 0001
SP3
2024 GPU Memory Exploitation for Fun and Profit
Yanan Guo 0002, Zhenkai Zhang 0002, Jun Yang 0002
USENIX Security Symposium2
2024 Invalidate+Compare: A Timer-Free GPU Cache Attack Primitive
Zhenkai Zhang 0002, Kunbei Cai, Yanan Guo 0002, Fan Yao 0001, Xing Gao 0001
USENIX Security Symposium1
2023 TunneLs for Bootlegging: Fully Reverse-Engineering GPU TLBs for Challenging Isolation Guarantees of NVIDIA MIG
abstract
Recent studies have revealed much detailed information about the translation lookaside buffers (TLBs) of modern CPUs, but we find that many properties of such components in modern GPUs still remain unknown or unclear. To fill this knowledge gap, we develop a new GPU TLB reverse-engineering method and apply it to a variety of consumer- and server-grade GPUs in Turing and Ampere generations. Aside from learning significantly more comprehensive and accurate GPU TLB properties, we discover a design flaw of NVIDIA Multi-Instance GPU (MIG) feature. MIG claims full partitioning of the entire GPU memory system for secure GPU sharing in cloud computing. However, we surprisingly find that MIG does not partition the last-level TLB, which is shared by all the compute units in a GPU. Exploiting this design flaw and learned TLB properties, we are able to construct a covert channel for data exfiltration across MIG-enforced isolation. To the best of our knowledge, this is the first attack on MIG. We evaluate the proposed attack on a commercial cloud platform, and we successfully achieve reliable data exfiltration from a victim tenant at a speed of up to 31 kbps with a very high accuracy around 99.8%. Even when the victim is using the GPU for deep neural network training, the transmission can still reach more than 25 kbps with a more than 99.5% accuracy. We propose and implement a mitigation approach that can effectively thwart data exfiltration through this covert channel. Additionally, we present a preliminary study on exploiting the access patterns of the last-level TLB to infer the identity of applications running in other MIG-created GPU instances.
Zhenkai Zhang 0002, Tyler N. Allen, Fan Yao 0001, Xing Gao 0001, Rong Ge 0002
CCS1
2023 BeKnight: Guarding Against Information Leakage in Speculatively Updated Branch Predictors
abstract
Information leakage through processor microarchi-tectural components exploiting speculative execution is raising significant security concerns. Modern commercial processors incorporate branch predictor designs where internal states of branch predictor structures are speculatively updated. Recent studies have shown that speculatively updated branch predictors allow side channel exploitation in the speculative domain, extending branch predictors to be another source of transmitting medium in transient execution attacks. While postponing updates of branch predictor states at a later time (e.g., during commit) can avoid exploitation in the speculation domain, it can result in belated correction of prediction outcomes (e.g., branch direction), leading to non-trivial degradation of prediction performance. In this paper, we present BeKnight, a secure branch predictor design that defeats speculative side channels targeting the branch direction prediction structure as the source of leakage. BeKnight aims to retain the performance advantage of early branch predictor updates (i.e., at resolution time) while ensuring no transient leakage. To achieve this, BeKnight conscientiously tracks the own-ership and speculative changes of potentially unsafe pattern history entries using a small Speculative Pattern Lookaside Buffer (SPLB). BeKnight efficiently audits the use of pattern history by allowing subsequent predictions in the same domain to benefit from early updates while annulling potential leakage through ensuring architecturally correct pattern is used on detection of a domain conflict. We evaluate the performance of BeKnight using 24 representative workloads from SPEC-2017. Notably, BeKnight achieves almost identical performance compared to the system with insecure but performant speculatively-updated predictors.
Md Hafizul Islam Chowdhuryy, Zhenkai Zhang 0002, Fan Yao 0001
ICCAD2
2022 CSDLEEG: Identifying Confused Students Based on EEG Using Multi-View Deep Learning
abstract
Distance learning has dramatically increased in recent years because of advanced technology. In addition, numerous universities had to offer courses in online mode in 2020 and 2021 because of the COVID-19 pandemic. However, there are more challenges in distance learning than in the traditional learning method (e.g., feedback and interaction). Recently, researchers started using simple EEG headsets to identify confused students during online courses based on machine learning approaches. However, they faced unpleasant accuracy using traditional machine learning algorithms or nondeep neural networks. In this paper, we present a data-driven approach based on a multi-view deep learning technique called CSDLEEG to identify confused students. We employ the students' demographic information and EEG signals to feed our novel neural networks. The results show that our proposed approach is superior to state-of-the-art methods for 98% accuracy and 98% F1-score.
Hashim Abu-gellban, Long Hoang Nguyen 0002, Zhenkai Zhang 0002, Essa Imhmed
COMPSAC4
2022 LockedDown: Exploiting Contention on Host-GPU PCIe Bus for Fun and Profit
abstract
The deployment of modern graphics processing units (GPUs) has grown rapidly in both traditional and cloud computing. Nevertheless, the potential security issues brought forward by this extensive deployment have not been thoroughly investigated. In this paper, we disclose a new exploitable side-channel vulnerability that ubiquitously exists in systems equipped with modern GPUs. This vulnerability is due to measurable contention caused on the host-GPU PCIe bus. To demonstrate the exploitability of this vulnerability, we conduct two case studies. In the first case study, we exploit the vulnerability to build a cross-VM covert channel that works on virtualized NVIDIA GPUs. To the best of our knowledge, this is the first work that explores covert channel attacks under the circumstances of virtualized GPUs. The covert channel can reach a speed up to 90 kbps with a considerably low error rate. In the second case study, we exploit the vulnerability to mount a website fingerprinting attack that can accurately infer which web pages are browsed by a user. The attack is evaluated against popular browsers like Chrome and Firefox on both Windows and Linux, and the results show that this fingerprinting method can achieve up to 95.2% accuracy. In addition, the attack is evaluated against Tor browser, and up to 90.6% accuracy can be achieved.
Mert Side, Fan Yao 0001, Zhenkai Zhang 0002
EuroS&P3
2022 A Vision Transformer Architecture for Open Set Recognition
abstract
Deep neural networks have demonstrated prominent capacities for image classification tasks in a closed set setting, where the test data come from the same distribution as the training data. However, in a more realistic open set scenario, traditional classifiers with incomplete knowledge cannot tackle test data that are not from the training classes. Open set recognition (OSR) aims to address this problem by both identifying unknown classes and distinguishing known classes simultaneously. In this paper, we propose a novel approach to OSR that is based on the vision transformer (ViT) technique. Specifically, our approach employs two separate training stages. First, a ViT model is trained to perform closed set classification. Then, an additional detection head is attached to the embedded features extracted by the ViT, trained to force the representations of known data to class-specific clusters compactly. Test examples are identified as known or unknown based on their distance to the cluster centers. To the best of our knowledge, this is the first time to leverage ViT for the purpose of OSR, and our extensive evaluation against several OSR benchmark datasets reveals that our approach significantly outperforms other baseline methods and obtains new state-of-the-art performance.
Feiyang Cai, Zhenkai Zhang 0002, Jie Liu 0001, Xenofon Koutsoukos
ICMLA2
2022 Graphics Peeping Unit: Exploiting EM Side-Channel Information of GPUs to Eavesdrop on Your Neighbors
abstract
As the popularity of graphics processing units (GPUs) grows rapidly in recent years, it becomes very critical to study and understand the security implications imposed by them. In this paper, we show that modern GPUs can “broadcast” sensitive information over the air to make a number of attacks practical. Specifically, we present a new electromagnetic (EM) side-channel vulnerability that we have discovered in many GPUs of both NVIDIA and AMD. We show that this vulnerability can be exploited to mount realistic attacks through two case studies, which are website fingerprinting and keystroke timing inference attacks. Our investigation recognizes the commonly used dynamic voltage and frequency scaling (DVFS) feature in GPU as the root cause of this vulnerability. Nevertheless, we also show that simply disabling DVFS may not be an effective countermeasure since it will introduce another highly exploitable EM side-channel vulnerability. To the best of our knowledge, this is the first work that studies realistic physical side-channel attacks on non-shared GPUs at a distance.
Zihao Zhan, Zhenkai Zhang 0002, Sisheng Liang, Fan Yao 0001, Xenofon Koutsoukos
SP2
2022 Moving target defense for the security and resilience of mixed time and event triggered cyber-physical systems
Bradley Potteiger, Abhishek Dubey, Feiyang Cai, Xenofon Koutsoukos, Zhenkai Zhang 0002
J. Syst. Archit.5
2021 Red Alert for Power Leakage: Exploiting Intel RAPL-Induced Side Channels
abstract
RAPL (Running Average Power Limit) is a hardware feature introduced by Intel to facilitate power management. Even though RAPL and its supporting software interfaces can benefit power management significantly, they are unfortunately designed without taking certain security issues into careful consideration. In this paper, we demonstrate that information leaked through RAPL-induced side channels can be exploited to mount realistic attacks. Specifically, we have constructed a new RAPL-based covert channel using a single AVX instruction, which can exfiltrate data across different boundaries (e.g., those established by containers in software or even CPUs in hardware); and, we have investigated the first RAPL-based website fingerprinting technique that can identify visited webpages with a high accuracy (up to 99% in the case of the regular network using a browser like Chrome or Safari, and up to 81% in the case of the anonymity network using Tor). These two studies form a preliminary examination into RAPL-imposed security implications. In addition, we discuss some possible countermeasures.
Zhenkai Zhang 0002, Sisheng Liang, Fan Yao 0001, Xing Gao 0001
AsiaCCS1
2021 OnlineDC: Leveraging Temporal Driving Behavior to Facilitate Driver Classification
abstract
Driver classification is used recently for vehicle anti-burglary and fake driver accounts based on driving behavior. Anti-burglary is a challenging problem as it leans on external devices to defend against vehicle theft. Several researchers analyzed the driving behavior to identify drivers, but they faced several challenges to produce a stable model for the cold start problem and for medium-long sequences. In addition, some approaches had an unpleasant performance when the action space increased (> 2 drivers). In this paper, we propose a novel approach named OnlineDC (Online Driver Classification), which leverages temporal driving behavior to identify a human subject behind the wheel. Our method utilizes the Gated Recurrent Unit (GRU) and the ResNet with the Squeeze-Excite blocks (SE) to analyze the long-short term patterns of driving behaviors. Moreover, we fostered the performance by building and applying the Feature Generation (FG) algorithm to extract spectral, temporal, and statistical features from the sensing data of vehicles. We conducted extensive experiments to show how our approach outperformed state-of-the-art baseline methods. The results also showed that our solution could resolve the cold-start problem for short patterns.
Hashim Abu-gellban, Long Hoang Nguyen 0002, Fang Jin, Zhenkai Zhang 0002
IEEE BigData5
2020 Security in Mixed Time and Event Triggered Cyber-Physical Systems using Moving Target Defense
abstract
Memory corruption attacks such as code injection, code reuse, and non-control data attacks have become widely popular for compromising safety-critical Cyber-Physical Systems (CPS). Moving target defense (MTD) techniques such as instruction set randomization (ISR), address space randomization (ASR), and data space randomization (DSR) can be used to protect systems against such attacks. CPS often use time-triggered architectures to guarantee predictable and reliable operation. MTD techniques can cause time delays with unpredictable behavior. To protect CPS against memory corruption attacks, MTD techniques can be implemented in a mixed time and event-triggered architecture that provides capabilities for maintaining safety and availability during an attack. This paper presents a mixed time and event-triggered MTD security approach based on the ARINC 653 architecture that provides predictable and reliable operation during normal operation and rapid detection and reconfiguration upon detection of attacks. We leverage a hardware-in-the-loop testbed and an advanced emergency braking system (AEBS) case study to show the effectiveness of our approach.
Bradley Potteiger, Feiyang Cai, Abhishek Dubey, Xenofon Koutsoukos, Zhenkai Zhang 0002
ISORC5
2020 Leveraging EM Side-Channel Information to Detect Rowhammer Attacks
abstract
The rowhammer bug belongs to software-induced hardware faults, and has been exploited to form a wide range of powerful rowhammer attacks. Yet, how to effectively detect such attacks remains a challenging problem. In this paper, we propose a novel approach named RADAR (Rowhammer Attack Detection via A Radio) that leverages certain electromagnetic (EM) signals to detect rowhammer attacks. In particular, we have found that there are recognizable hammering-correlated sideband patterns in the spectrum of the DRAM clock signal. As such patterns are inevitable physical side effects of hammering the DRAM, they can "expose" any potential rowhammer attacks including the extremely elusive ones hidden inside encrypted and isolated environments like Intel SGX enclaves. However, the patterns of interest may become unapparent due to the common use of spread-spectrum clocking (SSC) in computer systems. We propose a de-spreading method that can reassemble the hammering-correlated sideband patterns scattered by SSC. Using a common classification technique, we can achieve both effective and robust detection-based defense against rowhammer attacks, as evaluated on a RADAR prototype under various scenarios. In addition, our RADAR does not impose any performance overhead on the protected system. There has been little prior work that uses physical side-channel information to perform rowhammer defenses, and to the best of our knowledge, this is the first investigation on leveraging EM side-channel information for this purpose.
Zhenkai Zhang 0002, Zihao Zhan, Daniel Balasubramanian, Bo Li 0026, Péter Völgyesi, Xenofon Koutsoukos
SP1
2019 A model-based design approach for simulation and virtual prototyping of automotive control systems using port-Hamiltonian systems
Siyuan Dai, Zhenkai Zhang 0002, Xenofon Koutsoukos
Softw. Syst. Model.2
2017 Work-in-Progress: Cache-Aware Partitioned EDF Scheduling for Multi-core Real-Time Systems
abstract
As the number of cores and utilization of the system are increasing quickly, shared resources like caches are interfering tasks' execution behaviors more heavily. In order to achieve resource efficiency in both temporal and spatial domains for multi-core real-time systems, caches should be taken into consideration when performing partitions. In this paper, partitioned Earliest Deadline First (EDF) scheduling on a preemptive multi-core platform is considered. We propose a new system model that covers inter-task cache interference and describe some ongoing work in identifying proper partition schemes under such settings.
Zhishan Guo, Ying Zhang 0066, Lingxiang Wang, Zhenkai Zhang 0002
RTSS4
2016 Cache-related preemption delay analysis for multi-level inclusive caches
abstract
Cache-related preemption delay (CRPD) analysis is crucial when designing embedded control systems that employ preemptive scheduling. CRPD analysis for single-level caches has been studied extensively based on useful cache blocks (UCBs). As high-performance embedded processors are increasingly used, which are often equipped with multi-level caches, CRPD analysis for cache hierarchies also needs to be investigated. Recently, an approach has been proposed to estimate CRPD for multi-level non-inclusive caches. Since multi-level inclusive caches are also commonly used, especially in some multi-core processors, it becomes important to study how to analyze CRPD for inclusive cache hierarchies. However, as shown in this paper, new challenges appear due to the strict inclusion enforcement in the multi-level inclusive caches, which make the traditional UCB concept hard to use. In this paper, we propose a new concept of useful positive references (UPRs) to replace the UCB concept. Based on UPRs, we propose an approach to bound the additional cache misses due to a preemption in a two-level inclusive cache hierarchy. We present theoretical analysis to show the approach is safe, and we evaluate the proposed approach on a set of benchmarks to demonstrate its effectiveness. To the best of our knowledge, this is the first attempt to analyze CRPD for multi-level inclusive caches.
Zhenkai Zhang 0002, Xenofon Koutsoukos
EMSOFT1
2015 Improving the Precision of Abstract Interpretation Based Cache Persistence Analysis
abstract
When designing hard real-time embedded systems, it is required to estimate the worst-case execution time (WCET) of each task for schedulability analysis. Precise cache persistence analysis can significantly tighten the WCET estimation, especially when the program has many loops. Methods for persistence analysis should safely and precisely classify memory references as persistent. Existing safe approaches suffer from multiple sources of pessimism and may not provide precise results. In this paper, we first identify some sources of pessimism that two recent approaches based on younger set and may analysis may encounter. Then, we propose two methods to eliminate these sources of pessimism. The first method improves the update function of the may analysis-based approach; and the second method integrates the younger set-based and may analysis-based approaches together to further reduce pessimism. We also prove the two proposed methods are still safe. We evaluate the approaches on a set of benchmarks and observe the number of memory references classified as persistent is increased by the proposed methods. Moreover, we empirically compare the storage space and analysis time used by different methods.
Zhenkai Zhang 0002, Xenofon Koutsoukos
LCTES1
2015 Top-down and bottom-up multi-level cache analysis for WCET estimation
abstract
In many multi-core architectures, inclusive shared caches are used to reduce cache coherence complexity. However, the enforcement of the inclusion property can cause invalidation of memory blocks at higher cache levels. In order to ensure safety, analysis of cache hierarchies with inclusive caches for worst-case execution time (WCET) estimation is typically based on conservative decisions. Thus, the estimation may not be tight. In order to tighten the estimation, this paper proposes an approach that can more precisely analyze the behavior of a cache hierarchy maintaining the inclusion property. We illustrate the approach in the context of multi-level instruction caches. The approach first analyzes all the inclusive caches in the hierarchy in a bottom-up direction, and then analyzes the remaining non-inclusive caches in a top-down direction. In order to capture the inclusion victims and their effects, we also propose a concept of aging barrier and integrate it with the traditional must and persistence analyses to safely slow down their aging process so as to derive more precise analyses. We evaluate the proposed approach on a set of benchmarks and the evaluation reveals that the estimations are tightened.
Zhenkai Zhang 0002, Xenofon Koutsoukos
RTAS1
2015 Precise Multi-level Inclusive Cache Analysis for WCET Estimation
abstract
Multi-level inclusive caches are often used in multi-core processors to simplify the design of cache coherence protocol. However, the use of such cache hierarchies poses great challenges to tight worst-case execution time (WCET) estimation due to the possible invalidation behavior. Traditionally, multi-level inclusive caches are analyzed in a level-by-level manner, and at each level three analyses (i.e. must, may, and persistence) are performed separately. At a particular level, conservative decisions need to be made when the behaviors of other levels are not available, which hurts analysis precision. In this paper, we propose an approach which analyzes a multi-level inclusive cache by integrating the three analyses for all levels together. The approach is based on the abstract interpretation of a concrete operational semantics defined for multi-level inclusive caches. We evaluate the proposed approach and also compare it with two state-of-the-art approaches. From the experimental results, we can observe the proposed approach can significantly improve the analysis precision under relatively small cache size configurations.
Zhenkai Zhang 0002, Xenofon Koutsoukos
RTSS1