VLDB 2026 Research / reviewers in the wild / expert
Yuezhi Che
dblp:258/6101
· DBLP profile ↗
9ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-4351-7993ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Characterizing and Mitigating Context Redundancy in LLM Agent Workflows
Rongyu Luo, Yuezhi Che |
APPT | 3 |
| 2026 | Malope: Memory-Aware and Locality-Preserved Graph Neural Network Training
Junkun Shen, Yuezhi Che, Xiaobo Zhou 0002, Dazhao Cheng |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | User-Aware Critical Thread Identification and Proactive CPU Governing on Mobile DevicesabstractModern mobile devices now support increasingly complex, interactive applications that demand both low-latency frame rendering and energy efficiency. However, current Android CPU governors and schedulers react only to coarse CPU-utilization and ignore dynamic inter-thread dependencies and the computation demand of frame rendering, therefore, often fail to provide timely CPU resources, leading to missed frame deadlines or unnecessary energy consumption. To bridge this semantic gap, we propose Argus, a lightweight, kernel-integrated framework that dynamically identifies UI-critical threads based on real-time inter-thread interactions, and boosts their priority to ensure responsive user experiences. Furthermore, it incorporates a frame-aware CPU governor that adjusts frequency decisions based on frame deadlines and critical thread load. Comprehensive evaluations on popular mobile applications demonstrate that Argus can reduce the number of frame drops by up to 76.2 % and improve energy efficiency by up to 17.2 % compared to existing CPU governors, with negligible runtime overhead. Jiahao Qiu, Yuezhi Che, Dazhao Cheng |
ICPADS | 2 |
| 2025 | COSMOS: RL-Enhanced Locality-Aware Counter Cache Optimization for Secure MemoryabstractSecure memory systems employing AES-CTR encryption face significant performance challenges due to high counter (CTR) cache miss rates, especially in applications with irregular memory access patterns.These high miss rates increase memory traffic and latency, as each CTR cache miss triggers additional DRAM accesses.To address these bottlenecks and adapt to diverse access patterns, we propose COSMOS (Counter Optimized Secure Memory Operation Scheme), a novel solution leveraging reinforcement learning to reduce long memory access latency.COSMOS integrates two RL-based specialized predictors: one for data location prediction and another for CTR locality prediction, each with a well-defined state space, action space, and reward function.The RL-based data location predictor determines whether data reside on-chip or offchip after an L1 cache miss, enabling early CTR access for off-chip predictions with minimal changes to the existing cache hierarchy.The RL-based CTR locality predictor identifies CTRs with high locality, supporting a locality-centric CTR cache (LCR-CTR) to improve cache efficiency and reduce miss rates.COSMOS improves performance over MorphCtr by 25% in for irregular memory access applications, with minimal hardware overhead. Xiaoyang Lu, Yuezhi Che, Ziang Tian, Dazhao Cheng, Xian-He Sun, Michael T. Niemier, Xiaobo Sharon Hu |
MICRO | 3 |
| 2024 | Opca: Enabling Optimistic Concurrent Access for Multiple Users in Oblivious Data StorageabstractThe challenges of data privacy and security posed by data outsourcing are becoming increasingly prevalent. Oblivious RAM (ORAM)-based oblivious data storage guarantees data confidentiality through data encryption and access pattern obfuscation. However, it suffers from performance degradation and low throughput. To address these issues, the concurrency of ORAM in a multi-user scenario has been explored. We investigate several existing concurrent oblivious data storage solutions and discover that a trusted proxy is used to serve concurrent accesses between users and storage, with processing locks involved in the proxy to ensure correctness and prevent conflicts. The proxy-based system is inherently prone to pessimistic concurrency control, and as the number of users grows, a proxy might become a performance bottleneck, causing significant delays. In this study, we propose Opca, a novel oblivious data storage framework that enables optimistic concurrent access. Opca refines the proxy design by temporally storing multiple versions of modified data with labeled timestamps, committing only the latest version to the storage during a separate processing period. Opca is implemented and evaluated in different real-world storage backends with a scalable number of users, and its performance is compared to alternative schemes. Opca outperforms the state-of-the-art concurrent oblivious storage system TaoStore, which relies on a similar system setting. Our results show that Opca can improve 3.77x throughput and reduce 73.5% response time. Yuezhi Che, Dazhao Cheng, Xiao Wang 0012, Rujia Wang |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | DNNCloak: Secure DNN Models Against Memory Side-channel Based Reverse Engineering AttacksabstractAs deep neural networks (DNN) expand their attention into various domains and the high cost of training a model, the structure of a DNN model has become a valuable intellectual property and needs to be protected. However, reversing DNN models by exploiting side-channel leakage has been demonstrated in various ways. Even if the model is encrypted and the processing hardware units are trusted, the attacker can still extract the model’s structure and critical parameters through side channels, potentially posing significant commercial risks. In this paper, we begin by analyzing representative memory side-channel attacks on DNN models and identifying the primary causes of leakage. We also find that the full encryption used to protect model parameters could add extensive overhead. Based on our observations, we propose DNNCloak, a lightweight and secure framework aiming at mitigating reverse engineering attacks on common DNN architectures. DNNCloak includes a set of obfuscation schemes that increase the difficulty of reverse-engineering the DNN structure. Additionally, DNNCloak reduces the overhead of full weights encryption with an efficient matrix permutation scheme, resulting in reduced memory access time and enhanced security against retraining attacks on the model parameters. At last, we show how DNNCloak can defend DNN models from side-channel attacks effectively, with minimal performance overhead. Yuezhi Che, Rujia Wang |
ICCD | 1 |
| 2021 | Streamline Ring ORAM Accesses through Spatial and Temporal OptimizationabstractMemory access patterns could leak temporal and spatial information in a sensitive program; therefore, obfuscated memory access patterns are desired from the security perspective. Oblivious RAM (ORAM) has been the favored candidate to eliminate the access pattern leakage through randomly remapping data blocks around the physical memory space. Meanwhile, accessing memory with ORAM protocols results in significant memory bandwidth overhead. For each memory request, after going through the ORAM obfuscation, the main memory needs to service tens of actual memory accesses, and only one real access out of them is useful for the program execution. Besides, to ensure the memory bus access patterns are indistinguishable, extra dummy blocks need to be stored and transmitted, which cause memory space waste and poor performance. In this work, we introduce a new framework, String ORAM, that accelerates the Ring ORAM accesses with Spatial and Temporal optimization schemes. First, we identify that dummy blocks could significantly waste memory space and propose a compact ORAM organization that leverages the real blocks in memory to obfuscate the memory access pattern. Then, we identify the inefficiency of current transaction-based Ring ORAM scheduling on DRAM devices and propose an effective scheduling technique that can overlap the time spent on row buffer misses while ensuring correctness and security. With a minimal modification on the hardware and software, and negligible impact on security, the framework reduces 30.05% execution time and up to 40% memory space overhead compared to the state-of-the-art bandwidth-efficient Ring ORAM. Dingyuan Cao 0002, Mingzhe Zhang 0005, Xiaochun Ye, Dongrui Fan, Yuezhi Che, Rujia Wang |
HPCA | 6 |
| 2020 | Multi-Range Supported Oblivious RAM for Efficient Block Data RetrievalabstractData locality exists everywhere in the memory hierarchy. Most applications show temporal and spatial locality, and computer system and architecture designers utilize this property to improve the system performance with better data layout, prefetching, and scheduling. The locality property can be represented by memory access patterns, which records the time and frequency of accessed addresses. From the security perspective, if an attacker can trace the access pattern, sensitive information inside of the application could be observed and leaked. Oblivious RAM is one of the most effective solutions to mitigate the access pattern leakage on the system, which adds redundant data blocks in space and time. With ORAM protection, the intrinsic data locality is broken by the randomly stored data. Therefore, the application cannot gain any performance benefits from locality if the ORAM protocol is used. In this work, we would like to study the potential to support multi-range accesses with new storage and access efficient ORAM construction. Our proposed designs include two major schemes: Lite-rORAM, which minimize the storage overhead of existing rORAM; and Hybrid-rORAM, which support multiple ranges accesses with minimum storage overhead. We achieve the goal to preserve the locality for consecutive data blocks with different ranges in the application while obfuscates the access pattern as well. We tested our proposed schemes with different workloads on local and remote backends. The experimental results show that, in the best case, our proposed ORAM construction can reduce the data block retrieval time to 0.24x of the baseline Path ORAM, with 87.5% storage overhead reduction compared to rORAM. Yuezhi Che, Rujia Wang |
HPCA | 1 |
| 2019 | Imbalance-Aware Scheduler for Fast and Secure Ring ORAM Data RetrievalabstractData encryption can enhance data privacy but can not prevent privacy leakage completely. Side channels such as memory access pattern can also leak critical information in the program, while encryption cannot help when the access addresses are exposed to the attackers. Oblivious RAM (ORAM) was proposed to eliminate the access pattern leakage, and it can be incorporated with the memory controller to obfuscate and reshuffle memory accesses. Tree-based ORAM, such as Path ORAM and Ring ORAM, is a cost-effective ORAM organization, and it obfuscates the memory access pattern by reading and remapping data blocks along the path. Ring ORAM can achieve a better online data retrieval performance with a read protocol which only reads selective blocks along the path, compared to Path ORAM. However, the read path operation in Ring ORAM does not take the implementation of multi-channel memory system into consideration, which yields a potential bandwidth waste and a slower response time. Therefore, in this work, we investigate the root cause of the bandwidth waste and propose a fast and secure scheduler which balances the read path operation for Ring ORAM. Our scheduling schemes can improve the read operation response time as well as the overall system performance. The experimental results show that the read path latency for data retrieval can be reduced by 33% compared to the baseline Ring ORAM implementation. Yuezhi Che, Yuan Hong 0001, Rujia Wang |
ICCD | 1 |