Zengshi Wang

dblp:368/5317 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0007-5575-8440ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Khepri: Crystallizing TAGE for Memory Efficient Prewarm in Serverless Computing
abstract
As an increasingly popular cloud computing model, serverless computing suffers from performance degradation caused by microarchitectural cold start. Previous studies identify the front-end as the bottleneck and explore solutions such as instruction prefetching and restoring Branch Target Buffer. However, they fail to prewarm the Conditional Branch Predictor (CBP), because the large size of its core component, the TAgged GEometric history length predictor (TAGE), makes it impractical to be saved and restored.This paper observes the predictive sparsity of TAGE, where only a small subset of entries can dominate the predictor’s coverage and accuracy. We introduce Khepri, a memory efficient CBP prewarming mechanism that uses a TAGE Crystallization algorithm to identify these dominant entries. Khepri records them in main memory and restores them to prewarm TAGE at the next invocation. Khepri achieves a 1.57× speedup over the baseline and outperforms the state-of-the-art technique by 14%, requiring only 1.54KB in main memory on average.
Zengshi Wang, Zhuoyuan Yang, Kanheng Jiang, Jun Han 0003
DATE1
2026 DARE: An Irregularity-Tolerant Matrix Processing Unit with a Densifying ISA and Filtered Runahead Execution
Zengshi Wang, Jun Han 0003
DATE3
2026 Thoth: Uncovering Data-Dependent Memory Access Patterns via Annotation-Directed Load Sampling
abstract
Sparse data structures are ubiquitous in graph analytics, machine learning, and high-performance computing. Algorithms operating on these structures typically exhibit highly irregular data-dependent memory access (DDMA) patterns, leading to frequent cache misses and degraded memory performance. Prior work on hardware prefetching to mitigate DDMA-induced misses falls into two categories: address-based methods that sample correlated sequences of load data and addresses from cache miss streams, and instruction-based methods that record instruction-level dependency chains. Although both learn single relations effectively, they struggle with multi-level range relations prevalent in DDMA-intensive workloads, leaving substantial prefetching opportunities unexploited. In address-based schemes, misses from deeper-level consumers are often miscorrelated with the producer. Moreover, out-of-order execution and the range relations themselves perturb sampling, yielding mismatched load instances. In instruction-based schemes, chain-structured representations and suboptimal learning strategies prevent the construction of complete dependency chains for these relations. To overcome these limitations, we present Thoth, a hardware prefetcher that operates at the granularity of explicit producer-consumer load pairs rather than constructing dependency chains. Thoth detects such pairs via register-level dependency tracking. It adopts an annotation-directed load sampling strategy that annotates matched producer-consumer load instances and samples only those annotated instances, thereby robustly uncovering DDMA patterns—including multi-level range relations—while avoiding mismatches. To maintain annotation correctness across pipeline flushes, Thoth employs precise load annotation, which leverages reorder identifiers to resume or terminate annotation precisely. On a suite of DDMA-intensive benchmarks, Thoth delivers a 51.1% speedup over a no-prefetching baseline and outperforms two state-of-the-art DDMA prefetchers by 14.7% and 8.2%, respectively.
Kanheng Jiang, Yongxin Lyu, Zengshi Wang, Jun Han 0003
ACM Trans. Archit. Code Optim.4
2025 VLSUMaP: A High-Performance Matrix Processor with Virtually Expanded LSU Boosting HBM Bandwidth Utilization
Xinjie Kong, Zikang Zhou, Zhuoyuan Yang, Zengshi Wang, Jun Han 0003
ACM Great Lakes Symposium on VLSI7
2024 Chimera: A co-simulation framework combining with gem5 and FPGA platform for efficient verification
abstract
To accommodate the requirements of increasingly diverse applications, the scale of System-on-Chip (SoC) has extended beyond single-chip, evolving towards chiplets. This expansion is accompanied by a significant increase in the number of functional IPs. Each new IP undergoes a lengthy development process to meet expected performance and ensure compatibility with existing SoCs. The conventional development process is generally divided into several distinct stages, including architecture exploration, RTL development, standalone testing, and integration verification. As IP complexity escalates and coupling with the software stack becomes intricate, the standalone testing phase fails to guarantee complete IP functionality. This deficiency leads to extended development periods and necessitates numerous iterative design revisions during the integration verification phase, hindering the efficiency of IP integration into SoC. In this paper, we introduce Chimera, the first co-simulation framework that combines ESL simulators and the FPGA platform. The framework streamlines the conventional development process, providing a direct approach to architecture exploration and integration verification of a Register-Transfer Level (RTL) IP within an SoC running on the gem 5 simulator. We propose three optimizations including an optimistic synchronization mechanism, batch processing, and dataflow merging to enhance the efficiency of Chimera. Our experimental results demonstrate that Chimera can sustain co-simulation speeds comparable to the speed of the standalone gem5 simulator within typical system configurations. Furthermore, Chimera outperforms the state-of-the-art co-simulation framework, achieving up to a $20 \times$ performance improvement.
Zengshi Wang, Jun Han 0003
FPL2