VLDB 2026 Research / reviewers in the wild / expert
Chenlu Miao
dblp:288/4087
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
0009-0005-8233-1528ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Security and privacy · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory AccessesabstractIrregular memory accesses pose challenges for effective and efficient data prefetching. While temporal prefetchers have recently shown promise for irregular memory access patterns, their effectiveness fundamentally depends on temporal address recurrence and large metadata storage. When memory addresses exhibit weak or no recurrence, as in indirect memory accesses, temporal prefetchers achieve limited performance gains while incurring substantial storage overhead. This paper proposes Instruction-Correlation Prefetching (ICP), a new hardware prefetching mechanism that exploits instruction-level correlations rather than memory-address correlations to handle irregular memory accesses. ICP observes that although memory addresses may not repeat, the instructions generating them often recur with stable data-dependency relationships. By learning these persistent instruction correlations, ICP speculatively computes and prefetches future irregular accesses using the execution results of their correlated predecessors. Across irregular SPEC CPU and GAP benchmarks, ICP outperforms the state-of-the-art temporal prefetcher Triangel by 14.0% and the indirect prefetcher DMP by 6.0%, while requiring only 2.1 KB of hardware storage, over three orders of magnitude smaller than temporal prefetchers. Mengming Li, Chenlu Miao, Buqing Xu, Qijun Zhang, Xiangfeng Sun, Ceyu Xu, Yuan Xie 0001, Shang Liu 0006, Zhiyao Xie |
ISCA | 2 |
| 2025 | EARTH: Efficient Architecture for RISC-V Vector Memory AccessabstractVector processors frequently suffer from inefficient memory accesses, particularly for strided and segment memory access patterns. While coalescing strided accesses is a natural solution, effectively gathering or scattering elements at fixed strides remains a significant challenge. Naive approaches typically rely on high-overhead crossbars that remap any byte in memory or registers to any position in registers or memory, leading to physical design issues. Meanwhile, segment operations require row-column transpositions, which are often handled using either element-level in-place transposition (degrading performance) or large buffer-based bulk transposition (incurring high area overhead). Both options are undesirable, highlighting a need for more efficient solutions. In this paper, we present EARTH, a novel vector memory access architecture designed to overcome these challenges through shifting-based optimizations. For strided accesses, EARTH integrates specialized shift networks for gathering and scattering strided elements. After coalescing multiple accesses into one request within the same cache line, data can be routed between memory and registers through the shifting network with minimal overhead. For segment operations, EARTH employs a shifted register bank that enables direct column-wise access, eliminating the need for dedicated segment buffers while providing highperformance, in-place bulk transposition at acceptable overhead. We implemented the entire EARTH design on FPGA with Chisel HDL based on an open-source RISC-V vector unit Saturn. Our evaluation demonstrates that EARTH enhances performance for strided memory accesses proportionally to their prevalence in workloads, achieving $\mathbf{4 x}-\mathbf{8 x}$ speedups in benchmarks dominated by strided operations. The architecture also delivers area-efficient segment handling. Compared to conventional designs, EARTH reducing hardware area by 9% and power consumption by 41%. By optimizing these necessary memory access patterns, EARTH significantly advances both the performance and efficiency of vector processors. Hongyi Guan, Yichuan Gao, Chenlu Miao, Mingfeng Lin, Huayue Liang |
PACT | 3 |
| 2024 | Symphony: Path Validation at Scale
Anxiao He, Jiandong Fu, Kai Bu, Ruiqi Zhou, Chenlu Miao, Kui Ren 0001 |
NDSS | 5 |
| 2024 | TreasureCache: Hiding Cache Evictions Against Side-Channel AttacksabstractCache side-channel attacks remain a stubborn source of cross-core secret leakage. Such attacks exploit the timing difference between cache hits and misses. Most defenses thus choose to prevent cache evictions. Given that two possible types of evictions—flush-based and conflict-based—use different architectural features, these defenses have to integrate hybrid defense strategies, incur OS modification, and sacrifice performance to completely throttle cache side-channel attacks. In this paper, we present TreasureCache against cache side-channel attacks without modifying OS or sacrificing performance. Instead of preventing cache evictions with various costs, we advocate to allow cache evictions as is and hide exploitable evictions in our specialized small eviction-hidden buffer. The buffer guarantees a fast hit time comparative to LLC hits. This instantly closes the timing gap between accessing exploitable blocks when they are in and out of the LLC. Moreover, with the help of our buffer, we no longer have to disable flush instructions or shared memory. A lightweight constant-time flush instruction can help TreasureCache to prevent both flush-based and conflict-based side-channel attacks. We validate TreasureCache security and performance through extensive experiments. With a hardware overhead of less than 0.5%, TreasureCache reduces the secret-leakage resolution by about 1,000 times without introducing any performance slowdown. Mengming Li, Kai Bu, Chenlu Miao, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2023 | Hummingbird: Dynamic Path Validation With Hidden Equal-Probability SamplingabstractPath validation has already been incrementally deployed in the Internet architecture. It secures packet forwarding by enabling end hosts to negotiate specific forwarding paths and enforcing on-path routers to prove their forwarding behaviors along these paths. Most existing path validation solutions target static paths, paying less attention to fully dynamic paths that support flexible routing. In this paper, we present Hummingbird as the first validation solution over fully dynamic paths. It features a hidden equal-probability sampling technique. Gaining efficiency via routers probabilistically sampling packets to validate, we craft the sampling probability such that each router validates a similar amount of packets given an unknown path length. We further hide the state of whether a packet has been sampled and validated using a lightweight, non-cryptographic scheme. This prevents attackers from differentiating and selectively mis-forwarding packets. We validate security and efficiency of Hummingbird through both theoretical proof and experimental evaluation. Anxiao He, Xiang Li 0001, Jiandong Fu, Kai Bu, Chenlu Miao, Kui Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2022 | unXpec: Breaking Undo-based Safe SpeculationabstractSpeculative execution attacks exploiting speculative execution to leak secrets have aroused significant concerns in both industry and academia. They mainly exploit covert or side channels over microarchitectural states left by mis-speculated and squashed instructions (i.e., transient instructions). Most such attacks target cache states. Existing cache-based defenses against speculative execution attacks fall into two categories, Invisible and Undo. Most Invisible defenses buffer execution metadata of speculative instructions and place them into the cache only if the speculatively executed instructions become determined. Motivated by the fact that mis-speculations are rare cases, Undo defenses allow speculative instructions to modify cache states. Upon a mis-speculation, they rollback cache states to the ones prior to the execution of transient instructions. However, Invisible defenses have been recently found insecure by the speculative interference attack. This calls for a deep security inspection of Undo defenses against speculative execution attacks.In this paper, we present unXpec as the first attack against Undo-based safe speculation. It exploits the secret-dependent timing channel exhibited through the rollback operations of Undo defenses. Specifically, the rollback process requires both invalidating cache lines brought into the cache by transient instructions and restoring evicted cache lines from the cache by transiently loaded data. This opens up a channel that encodes secret via the timing difference between when rollback involves much invalidation and restoration or not. We further leverage eviction sets to enforce more restoration operations. This yields a longer rollback time and thus a larger secret-dependent timing difference. We demonstrate the timing channel over the open-source CleanupSpec, a representative Undo solution. A single transient load can trigger a secret-dependent timing difference of 22 cycles (without eviction sets) of 32 cycles (with eviction sets), which is sufficiently exploitable for constructing a covert channel for speculative execution attacks. We run unXpec on the gem5 simulator with CleanupSpec enabled. The results show that unXpec can leak secrets at a high rate of 140 Kbps with an accuracy over 90%. Simply enforcing constant-time rollback to mitigate unXpec may induce an over 70% performance overhead. Mengming Li, Chenlu Miao, Yilong Yang 0006, Kai Bu |
HPCA | 2 |
| 2022 | SwiftDir: Secure Cache Coherence without OverprotectionabstractCache coherence states have recently been exploited to leak secrets through timing-channel attacks. The root cause lies in the fact that shared data in state Exclusive (E) and state Shared (S) are served from different cache layers. The state-of-the-art countermeasure—S-MESI—serves both E- and S-state shared data from the last-level cache (LLC) by explicitly synchronizing the Modified (M) state across private caches and the LLC. This has to sacrifice the silent upgrade feature that MESI introduces for speedup. Moreover, it enforces protection to not only exploitable shared data but also unshared data. This further slows down performance, especially for write-after-read intensive applications. In this paper, we propose SwiftDir to efficiently secure cache coherence against cover-channel attacks without overprotection. SwiftDir fundamentally narrows down the protection scope to write-protected data. Such exploitable shared data can be uniquely identified with the write-protection permission in the memory management unit (MMU) and do not necessarily transit to state M. We validate this idea through tracing system calls of shared libraries on Linux. We then investigate all three commercial cache architectures (i.e., PIPT, VIPT, and VIVT) and find it feasible to hitchhike the address translation process to transmit the write-protection information from the MMU to the coherence controller. Then SwiftDir enforces protection over only write-protected data by serving all requests toward them directly from the LLC with a constant latency. This not only simplifies how MESI handles write-protected data but also avoids how S-MESI overprotects them. Meanwhile, SwiftDir still preserves silent upgrade for efficient handling of unshared data. Extensive experiments demonstrate that our SwiftDir can secure cache coherence while outperforming not only secure SMESI but also unprotected MESI. Chenlu Miao, Kai Bu, Mengming Li, Shaowu Mao, Jianwei Jia |
MICRO | 1 |