EDBT 2026 Demo / reviewers in the wild / expert
Yanan Guo 0002
dblp:120/5486-2
· DBLP profile ↗
24ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0003-0034-0358ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 4 first-author · 17 since 2021Security and privacy · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ASPA: Reassigning DDR5 Parity BandwidthabstractRecent memory advancements, such as DDR5, HBM3, and emerging memory-centric accelerators, primarily focus on increasing bandwidth capacity, yet often omit the significance of effective bandwidth utilization, i.e., bandwidth efficiency. Motivated by the suboptimal channel allocation in DDR5, where parity accounts for 25% bandwidth overhead, we propose ASPA, an efficiency-oriented solution that reallocates parity bandwidth to boost regular data transfer. The objective of ASPA is to enhance bandwidth efficiency without compromising reliability while maintaining low hardware overhead. In particular, we leverage existing CRC (Cyclic Redundancy Check) units in DRAM chips to opportunistically generate second-tier parity for the existing ECC (Error Correction Code) parity. For bulk-sized memory accesses, only the small second-tier parity is transmitted. This reduces the parity bandwidth consumption, allowing data chips to reuse the freed parity bandwidth for data transfer. Our observation indicates that the protection capability is sufficient when the second-tier parity (with a 64-bit size) is used exclusively for error detection. Furthermore, by exploiting underutilized resources in high-performance memory systems, ASPA is implemented with negligible hardware overhead. Qiufeng Li, Yanan Guo 0002, Weidong Cao 0001, Xin Xin 0008 |
HPCA | 3 |
| 2026 | Coarse-Grained Duplication First, Fine-Grained Deduplication Later: Duplication-Centric Multi-GPU Memory Management
Xiangyue Huang 0001, Yanan Guo 0002, Yuanchao Xu 0001 |
ISCA | 2 |
| 2026 | LIBRA: A High-Accuracy, Cost-Aware, and Coordinated Multi-GPU Page Prefetcher
Xiangyue Huang 0001, Yanan Guo 0002, Yuanchao Xu 0001 |
ISCA | 2 |
| 2026 | HBM-CASO: A Coordinated Approach to HBM System-Level and On-Die ECC
Ruizhi Zhu, Yanan Guo 0002, Huize Li, Weidong Cao 0001, Qian Lou, Xin Xin 0008 |
ISCA | 2 |
| 2026 | Exploiting TLBs in Virtualized GPUs for Cross-VM Side-Channel Attacks
Hongyue Jin, Yanan Guo 0002, Zhenkai Zhang 0002 |
NDSS | 2 |
| 2026 | GeForge: Hammering GDDR Memory to Forge GPU Page Tables for Fun and ProfitabstractOver the years, Rowhammer has been leveraged to mount a wide range of attacks against system main memory. While a recent study has revealed that GPU memory is similarly vulnerable, the security implications remain largely under-explored. To advance this line of research, we introduce GeForge, an end-to-end Rowhammer attack that exploits bit flips induced in GPU memory to achieve system-level compromise. At its core, GeForge corrupts GPU page tables to seize control of address translation, enabling arbitrary access to the entire GPU memory. Moreover, by exploiting a special mapping feature in the GPU page table, GeForge extends its reach to directly access host memory. To make GeForge practical under default system settings, we develop novel techniques that eliminate restrictive assumptions in prior work. Our techniques include a method for aligning offline-profiled physical address mappings to runtime GPU allocations and a memory massaging strategy that steers target GPU page table structures into vulnerable locations via the stock driver allocator. In addition, we improve the hammering pattern to trigger many more bit flips than prior work. With these approaches, we successfully mount GeForge on widely deployed NVIDIA GPUs, including both workstationclass and consumer-grade ones. We show that GeForge allows an attacker to arbitrarily read and modify data across GPU contexts. More crucially, we demonstrate that GeForge can help the attacker escalate privileges to root on the host system. Junpeng Wan, Yanan Guo 0002, Zhi Zhang 0001, Jing (Dave) Tian, Zhenkai Zhang 0002 |
SP | 2 |
| 2025 | SCREME: A Scalable Framework for Resilient Memory DesignabstractThe continuing advancement of memory technology has not only fueled a surge in performance, but also substantially exacerbated reliability challenges. Traditional solutions have primarily focused on improving the efficiency of protection schemes, i.e., Error Correction Codes (ECC), under the assumption that allocating additional memory space for ECC parity is always costly and therefore unsustainable as parity scales. We break the stereotype by proposing an orthogonal approach that provides additional, cost-effective memory space for resilient memory design. In particular, we recognize that ECC chips (used for parity storage) do not necessarily require the same performance level as regular data chips. This offers two-fold benefits: First, the bandwidth originally provisioned for a regularperformance ECC chip can instead be used to accommodate multiple low-performance chips. Second, the cost of ECC chips can be effectively reduced, as lower performance often correlates with lower expense. In addition, we observe that server-class memory chips are often provisioned with ample, yet underutilized I/O resources. This suggests an opportunity to repurpose these resources for flexible on-DIMM interconnections. Based on the above two insights, we finally propose SCREME, a scalable memory framework that leverages cost-effective yet slower chips - byproducts of rapid technology evolution - to meet the growing reliability demands driven by this evolution. Mimi Xie, Yanan Guo 0002, Huize Li, Xin Xin 0008 |
PACT | 3 |
| 2025 | Chekhov's Gun: Uncovering Hidden Risks in macOS Application-Sandboxed PID-Domain ServicesabstractmacOS delegates many high-privilege operations to dedicated PID-domain services, which applications can register and communicate with through inter-process communication (IPC). This architecture improves userland stability and security but also introduces attractive attack surfaces for adversaries. In this paper, we systematically analyze PID-domain services and uncover an overlooked attack vector: PID-domain services that are restricted to an Application Sandbox identical to the calling application can still be exploited due to subtle entitlement differences. Minghao Lin, Jiaxun Zhu, Tingting Yin, Zechao Cai, Guanxing Wen, Yanan Guo 0002, Mengyuan Li 0004 |
CCS | 6 |
| 2025 | Security and Performance Implications of GPU Cache Eviction Priority HintsabstractNVIDIA provides cache eviction priority hints such as evict_first and evict_last on recent GPUs.These hints allow users to specify the eviction priority that should be used for individual cache lines to improve cache utilization.However, NVIDIA does not disclose the microarchitectural details of these hints or cache eviction behaviors when using them, which makes their security and performance implications unclear.In this paper, we first reverse engineer the detailed comprehensive behaviors of these eviction priority hints.Then based on our findings, we analyze their impact on system security and performance.First, we found that these priority hints introduce new security problems.Specifically, we develop a new covert channel using the evict_first priority hint, which is more efficient than existing GPU covert channels.We also demonstrate a performance degradation attack using the evict_last priority hint, which is more stealthy compared to the known methods.Second, from the performance perspective, we show that marking a cache line as evict_last does not always keep it in the cache.In fact, if more than 12/16 (or 3/16, depending on the driver version) of the L2 cache size worth of data are marked as evict_last, cache thrashing can occur, which leads to performance degradation for real GPU workloads. Qizhong Wang, Xiangyue Huang 0001, Yanan Guo 0002, Yuanchao Xu 0001 |
MICRO | 3 |
| 2024 | QRCC: Evaluating Large Quantum Circuits on Small Quantum Computers through Integrated Qubit Reuse and Circuit CuttingabstractQuantum computing has recently emerged as a promising computing paradigm for many application domains. However, the size of quantum circuits that can be run with high fidelity is constrained by the limited quantity and quality of physical qubits. Recently proposed schemes, such as wire cutting and qubit reuse, mitigate the problem but produce sub-optimal results as they address the problem individually. In addition, gate cutting, an alternative circuit-cutting strategy that is suitable for circuits computing expectation values, has not been fully explored in the field. Aditya Pawar, Yingheng Li, Zewei Mo, Yanan Guo 0002, Xulong Tang, Youtao Zhang, Jun Yang 0002 |
ASPLOS (4) | 4 |
| 2024 | RTT-UAF: Reuse Time Tracking for Use-After-Free DetectionabstractMemory safety continues to be a critical challenge in modern computing, with approximately 70% of reported vulnerabilities annually attributed to memory-related issues. Among these issues, Use-After-Free (UAF) vulnerabilities or bugs, where a program accesses memory through a dangling pointer, pose significant threats. Existing UAF detection methods, such as Key-And-Lock (KAL) mechanisms, incur notable performance overhead due to explicitly propagating the key and lock address (metadata). We identify that approximately 67% of KAL’s performance overhead is introduced by this metadata propagation. This paper introduces RTT, which significantly reduces performance overhead by reducing the number of memory accesses related to metadata propagation. RTT achieves an average performance overhead of 170%, substantially lower than existing KAL methods, and incurs only an 8% memory overhead on SPEC CPU 2017. The results of experimental evaluations on real-world UAF bugs further demonstrate that RTT’s UAF bug detection rate is equivalent to other KAL methods. Yubo Du, Yanan Guo 0002, Youtao Zhang, Jun Yang 0002 |
ICS | 2 |
| 2024 | GPU Memory Exploitation for Fun and Profit
Yanan Guo 0002, Zhenkai Zhang 0002, Jun Yang 0002 |
USENIX Security Symposium | 1 |
| 2024 | Invalidate+Compare: A Timer-Free GPU Cache Attack Primitive
Zhenkai Zhang 0002, Kunbei Cai, Yanan Guo 0002, Fan Yao 0001, Xing Gao 0001 |
USENIX Security Symposium | 3 |
| 2023 | Orchestrating Measurement-Based Quantum Computation over Photonic Quantum ProcessorsabstractQuantum computing has rapidly evolved in recent years and has established its supremacy in many application domains. While matter-based qubit platforms such as superconducting qubits have received the most attention so far, there is a rising interest in photonic qubits lately, which show advantages in parallelism, speed, and scalability. Photonic qubits are best served by the paradigm of measurement-based quantum computation (MBQC). To deliver the promise of measurement-based photonic quantum computing (MBPQC), the photon cluster state depth and photon utilization are two of the most important metrics. However, little attention has been paid to optimizing the depth and utilization when mapping quantum circuits to the photon clusters. In this paper, we propose a compiler framework that achieves automatic and dynamic depth and utilization optimizations. Our approach consists of an MBPQC mapping mechanism that maps optimized measurement patterns on a cluster state and a cluster state pruning strategy that removes all possible redundancies without impacting the circuit functions. Experimental results on five quantum benchmark with three different qubit numbers indicate our approach achieves an average of 63.4% cluster depth reduction and 22.8% photon utilization improvements. Yingheng Li, Aditya Pawar, Mohadeseh Azari, Yanan Guo 0002, Youtao Zhang, Jun Yang 0002, Kaushik Parasuram Seshadreesan, Xulong Tang |
DAC | 4 |
| 2023 | Understanding and Defending Patched-based Adversarial Attacks for Vision TransformerabstractVision Transformer (ViT) is an attention-based model architecture that has demonstrated superior performance on many computer vision tasks. However, its security properties, in particular, the robustness against adversarial attacks, are yet to be thoroughly studied. Recent works have shown that ViT is vulnerable to attention-based adversarial patch attacks, which cover 1-3% area of the input image using adversarial patches and degrades the model accuracy to 0%. This work provides a generic study targeting the attention-based patch attack. First, we experimentally observe that adversarial patches only activate in a few layers and become lazy during attention updating. According to experiments, we study the theory of how a small adversarial patch perturbates the whole model. Based on understanding adversarial patch attacks, we propose a simple but efficient defense that correctly detects more than 95% of adversarial patches. Yanan Guo 0002, Youtao Zhang, Jun Yang 0002 |
ICML | 2 |
| 2023 | Uncore Encore: Covert Channels Exploiting Uncore Frequency ScalingabstractModern processors dynamically adjust clock frequencies and voltages to reduce energy consumption. Recent Intel processors separate the uncore frequency from the core frequency, using Uncore Frequency Scaling (UFS) to adapt the uncore frequency to various workloads. While UFS improves power efficiency, it also introduces security vulnerabilities. In this paper, we study the feasibility of covert channels exploiting UFS. First, we conduct a series of experiments to understand the details of UFS, such as the factors that can cause uncore frequency variations. Then, based on the results, we build the first UFS-based covert channel, UF-variation, which works both across-cores and across-processors. Finally, we analyze the robustness of UF-variation under known defense mechanisms against uncore covert channels, and show that UF-variation remains functional even with those defenses in place. Yanan Guo 0002, Dingyuan Cao 0002, Xin Xin 0008, Youtao Zhang, Jun Yang 0002 |
MICRO | 1 |
| 2023 | IDYLL: Enhancing Page Translation in Multi-GPUs via Light Weight PTE InvalidationsabstractMulti-GPU systems have emerged as a desirable platform to deliver high computing capabilities and large memory capacity to accommodate large dataset sizes. However, naively employing multi-GPU incurs non-scalable performance. One major reason is that execution efficiency suffers expensive address translations in multi-GPU systems. The data-sharing nature of GPU applications requires page migration between GPUs to mitigate non-uniform memory access overheads. Unfortunately, frequent page migration incurs substantial page table invalidation overheads to ensure translation coherence. A comprehensive investigation of multi-GPU address translation efficiency identifies two significant bottlenecks caused by page table invalidation requests: (i) increased latency for demand TLB miss requests and (ii) increased waiting latency for performing page migrations. Based on observations, we propose IDYLL, which reduces the number of page table invalidations by maintaining an “in-PTE" directory and reduces invalidation latency by batching multiple invalidation requests to exploit spatial locality. We show that IDYLL improves overall performance by 69.9% on average. Bingyao Li 0001, Yanan Guo 0002, Aamer Jaleel, Jun Yang 0002, Xulong Tang |
MICRO | 2 |
| 2023 | Generating Robust DNN With Resistance to Bit-Flip Based Adversarial Weight AttackabstractRowhammer Attack, a new DRAM-based attack, was developed exploiting weak cells to alter their content. Such attacks can be launched at the user level without requiring access permission to the victim memory cells. Leveraging such attacks, a new bit-flip-based adversarial weights attack (BFA) was developed targeting deep neural network models. When BFA attackers acquire a DNN model, they manipulate the existing DNN adversarial attack into locating vulnerable bits in the target DNN model. By flipping a subset of them using Rowhammer, they can crash that model within 30 trails. In this paper, we propose a lightweight and easy-to-deploy defense mechanism in the bit-level, Randomized Rotated and Nonlinear Encoding (RREC), which generates both robustness and fault-tolerant against BFA. Since flipping the most significant bit (MSB) in quantized data is too dangerous, we introduce randomized Rotation to obfuscate the bit order of model data and efficiently hide truly vulnerable bits with less vulnerable ones. Further, RREC reduces the average bit-flipped distance by more than 3x from the nonlinear encoding. It decreases the bit-flip distance among the majority of bits (including those vulnerable bits). Theoretically, RREC minimized the impact of a single bit BFA to 1/24 compared with baseline. Experimentally, RREC tolerates more than 17x flipped bits versus baseline model and 4.8x and 5.7x more bits compared with the existing BFA defenses (4B QAT and WR) with 0.01x to 0.08x of runtime latency. Moreover, we evaluate RREC against a newly emerged attack, Targeted-BFA, and it improves the defense rate from$5\%$to$95\%$. Yanan Guo 0002, Yueqiang Cheng, Youtao Zhang, Jun Yang 0002 |
IEEE Trans. Computers | 2 |
| 2022 | Q-GPU: A Recipe of Optimizations for Quantum Circuit Simulation Using GPUsabstractIn recent years, quantum computing has undergone significant developments and has established its supremacy in many application domains. While quantum hardware is accessible to the public through the cloud environment, a robust and efficient quantum circuit simulator is necessary to investigate the constraints and foster quantum computing development, such as quantum algorithm development and quantum device architecture exploration. In this paper, we observe that most of the publicly available quantum circuit simulators (e.g., QISKit from IBM, QDK from Microsoft, and Qsim-Cirq from Google) suffer from slow simulation and poor scalability when the number of qubits increases. To this end, we systematically investigate the deficiencies in quantum circuit simulation (QCS) and propose Q-GPU, a framework that leverages GPUs with comprehensive optimizations to allow efficient and scalable QCS. Specifically, Q-GPU features i) proactive state amplitude transfer, ii) zero state amplitude pruning, iii) delayed qubit involvement, and iv) lossless nonzero state amplitude compression. Experimental results across nine representative quantum circuits indicate that Q-GPU significantly reduces the execution time of the state-of-the-art GPU-based QCS by 71.89% (3.55× speedup). Q-GPU also outperforms the state-of-the-art OpenMP CPU implementation, the Google Qsim-Cirq simulator, and the Microsoft QDK simulator by 1.49×, 2.02×, and 10.82×, respectively. Yilun Zhao 0002, Yanan Guo 0002, Amanda Dumi, Devin M. Mulvey, Shiv Upadhyay, Youtao Zhang, Kenneth D. Jordan, Jun Yang 0002, Xulong Tang |
HPCA | 2 |
| 2022 | Leaky Way: A Conflict-Based Cache Covert Channel Bypassing Set AssociativityabstractModern $\times$86 processors feature many prefetch instructions that developers can use to enhance performance. However, with some prefetch instructions, users can more directly manipulate cache states which may result in powerful cache covert channel and side channel attacks. In this work, we reverse-engineer the detailed cache behavior of PREFETCHNTA on various Intel processors. Based on the results, we first propose a new conflict-based cache covert channel named NTP+NTP. Prior conflict-based channels often require priming the cache set in order to cause cache conflicts. In contrast, in NTP+NTP, the data of the sender and receiver can compete for one specific way in the cache set, achieving cache conflicts without cache set priming for the first time. As a result, NTP+NTP has higher bandwidth than prior conflict-based channels such as Prime+Probe. The channel capacity of NTP+NTP is 302 KB/s. Second, we found that PREFETCHNTA can also be used to boost the performance of existing side channel attacks that utilize cache replacement states, making those attacks much more efficient than before. Yanan Guo 0002, Xin Xin 0008, Youtao Zhang, Jun Yang 0002 |
MICRO | 1 |
| 2022 | Adversarial Prefetch: New Cross-Core Cache Side Channel AttacksabstractModern x86 processors have many prefetch instructions that can be used by programmers to boost performance. However, these instructions may also cause security problems. In particular, we found that on Intel processors, there are two security flaws in the implementation of PREFETCHW, an instruction for accelerating future writes. First, this instruction can execute on data with read-only permission. Second, the execution time of this instruction leaks the current coherence state of the target data. Based on these two design issues, we build two cross-core private cache attacks that work with both inclusive and non-inclusive LLCs, named Prefetch+Reload and Prefetch+Prefetch. We demonstrate the significance of our attacks in different scenarios. First, in the covert channel case, Prefetch+Reload and Prefetch+Prefetch achieve 782 KB/s and 822 KB/s channel capacities, when using only one shared cache line between the sender and receiver, the largest-to-date single-line capacities for CPU cache covert channels. Further, in the side channel case, our attacks can monitor the access pattern of the victim on the same processor, with almost zero error rate. We show that they can be used to leak private information of real-world applications such as cryptographic keys. Finally, our attacks can be used in transient execution attacks in order to leak more secrets within the transient window than prior work. From the experimental results, our attacks allow leaking about 2 times as many secret bytes, compared to Flush+Reload, which is widely used in transient execution attacks. Yanan Guo 0002, Andrew Zigerelli, Youtao Zhang, Jun Yang 0002 |
SP | 1 |
| 2021 | IVcache: Defending Cache Side Channel Attacks via Invisible AccessesabstractThe sharing of last-level cache (LLC) among different CPU cores makes cache vulnerable to side channel attacks. An attacker can get private information about co-running applications (victims) by monitoring their accesses in LLC. Cache side channel attacks can be mitigated by partitioning cache between the victim and attacker. However, previous partition works either make weak assumptions about the attacker's strength or force their security mechanisms and thus overhead to every user on the system, regardless of their security requirement. Yanan Guo 0002, Andrew Zigerelli, Youtao Zhang, Jun Yang 0002 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2021 | ModelShield: A Generic and Portable Framework Extension for Defending Bit-Flip based Adversarial Weight AttacksabstractBit-flip attack (BFA) has become one of the most serious threats to Deep Neural Network (DNN) security. By utilizing Rowhammer to flip the bits of DNN weights stored in memory, the attacker can turn a functional DNN into a random output generator. In this work, we propose ModelShield, a defense mechanism against BFA, based on protecting the integrity of weights using hash verification. ModelShield performs real-time integrity verification on DNN weights. Since this can slow down a DNN inference by up to 7×, we further propose two optimizations for ModelShield. We implement ModelShield as a lightweight software extension that can be easily installed into popular DNN frameworks. We test both the security and performance of ModelShield, and the results show that it can effectively defend BFA with less than 2% performance overhead. Yanan Guo 0002, Yueqiang Cheng, Youtao Zhang, Jun Yang 0002 |
ICCD | 1 |
| 2021 | SAM: Accelerating Strided Memory AccessesabstractStrided memory accesses are an important type of operations for In-Memory Databases (IMDB) applications. Strided memory accesses often demand data at word granularity with fixed strides. Hence, they tend to produce sub-optimal performance on DRAM memory (the de facto standard memory in modern computer systems) that accesses data at cacheline granularity. Recently proposed optimizations either introduce significant reliability degradation or are limited to non-volatile crossbar memory structures. Xin Xin 0008, Yanan Guo 0002, Youtao Zhang, Jun Yang 0002 |
MICRO | 2 |