Yuchen Zhou 0005

dblp:39/10084-5 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0005-2731-8988ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Intermittence-Aware Cache Compression
abstract
Cache compressions are proven effective in improving the performance of caches in conventional processors. They compress data into a smaller size, allowing caches to accommodate more blocks. This helps reduce cache misses and expensive memory accesses, ultimately improving performance. However, conventional cache compression is less effective for energy harvesting systems (EHSs) which experience frequent power failure, as many compressed blocks end up not being used before their loss upon power outage. This wastes hard-won energy, which would otherwise be used for making more program progress. To address this issue, this paper introduces Kagura, an adaptive cache compression extension with frequent power failure in mind. Specifically, Kagura disables cache compression when it finds out that many cached blocks are unlikely to be reused before the next power outage. That way, Kagura avoids the energy waste on useless compressions/decompressions, and the resulting speedup is on par with the ideal intermittenceaware cache compressor. Experimental results show that when combined with an existing cache compressor, Kagura reduces the total energy consumption by an average of$\mathbf{4. 5 3 \%}$(up to 16.21 %) and improves the performance by an average of 4.74 % (up to$\mathbf{1 7. 8 7 \%}$) compared to the baseline EHS without cache compression.
Gan Fang, Jianping Zeng 0001, Yuchen Zhou 0005, Changhee Jung
HPCA3
2026 Anchoring Whole-System Persistence and Resilience in CXL
abstract
This paper presents ANCHOR, a novel CXL-based memory hierarchy that achieves both persistence and resilience without recompilation or core modifications. For performant whole-system persistence (WSP), ANCHOR introduces a dual-path store architecture in which every committed store is sent to both the L1 cache and a CXL-attached memory-semantic SSD. To reduce the SSD write traffic, ANCHOR employs an STT-RAM write-combining buffer that coalesces stores before they reach the SSD. In particular, the SSD-resident committed store data serves as a ground-truth anchor for error repair, where ANCHOR repurposes cache and DRAM ECCs as detection codes and repairs detected errors using the corresponding SSD-resident clean copy. This approach effectively upgrades cache-side end-to-end reliability from intrinsic 1-bit correction to 3-bit correction, while also extending DRAM repair coverage beyond the limits of standard ChipKill—all at no additional ECC cost. Evaluation across 42 benchmarks shows that ANCHOR incurs only a 4.36% average run-time overhead while ensuring both WSP and system-wide resilience.
Yuchen Zhou 0005, Jianping Zeng 0001, Changhee Jung
ICS1
2024 LightWSP: Whole-System Persistence on the Cheap
abstract
Whole-system persistence (WSP) has recently attracted more interest thanks to its transparency and performance benefits over partial-system persistence where users are not only burdened by complex persistent programming but also incapable of using DRAM as LLC. Nevertheless, existing WSP work either introduces high hardware cost or causes non-trivial performance overhead. To this end, this paper presents LightWSP, a compiler/architecture co-design scheme that can achieve WSP in a lightweight yet performant manner. LightWSP compiler partitions program into a series of recoverable regions (epochs) with their live-out registers checkpointed, while LightWSP hardware persists the stores of the regions-whose boundary serves as a power failure recovery point-enforcing crash consistency; LightWSP leverages the battery-backed write pending queue (WPQ) of a memory controller as a redo buffer, i.e., all stores are first buffered in WPQ and then persisted together in non-volatile memory (NVM) at each region end. In this way, no matter when power failure happens, NVM is never corrupted by the stores of the power-interrupted region, facilitating correct recovery. In particular, LightWSP supports multiple memory controllers on the cheap without costly speculation/misspeculation handling mechanisms used by prior work. The experimental results with 38 applications show that LightWSP incurs only an average of 9.0% run-time overhead. This is on par with the state-of-the-art work, that complicates the core microarchitecture significantly with its intrusive design for memory controller speculation, yet the hardware cost of LightWSP is near zero (0.5B per core).
Yuchen Zhou 0005, Jianping Zeng 0001, Changhee Jung
MICRO1
2023 SweepCache: Intermittence-Aware Cache on the Cheap
abstract
This paper presents SweepCache, a new compiler/architecture co-design scheme that can equip energy harvesting systems with a volatile cache in a performant yet lightweight way. Unlike prior just-in-time checkpointing designs that persists volatile data just before power failure and thus dedicates additional energy, SweepCache partitions program into a series of recoverable regions and persists stores at region granularity to fully utilize harvested energy for computation. In particular, SweepCache introduces persist buffer—as a redo buffer resident in nonvolatile memory (NVM)—to keep the main memory consistent across power failure while persisting region’s stores in a failure-atomic manner. Specifically, for writebacks during region execution, SweepCache saves their cachelines to the persist buffer. At each region end, SweepCache first flushes dirty cachelines to the buffer, allowing the next region to start with a clean cache, and then moves all buffered cachelines to the corresponding NVM locations. In this way, no matter when power failure occurs, the buffer contents or their memory locations always remain intact, which serves as a basis for correct recovery. To hide the persistence delay, SweepCache speculatively starts a region right after the prior region finishes its execution—as if its stores were already persisted—with the two regions having their own persist buffer, i.e., dual-buffering. This region-level parallelism helps SweepCache to achieve the full potential of a high-performance data cache. The experimental results show that compared to the original cache-free nonvolatile processor, SweepCache delivers speedups of 14.60x and 14.86x—outperforming the state-of-the-art work by 3.47x and 3.49x—for two representative energy harvesting power traces, respectively.
Yuchen Zhou 0005, Jianping Zeng 0001, Jungi Jeong, Jongouk Choi, Changhee Jung
MICRO1