Zhuoyuan Yang

dblp:298/9322 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0007-7328-957XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Khepri: Crystallizing TAGE for Memory Efficient Prewarm in Serverless Computing
abstract
As an increasingly popular cloud computing model, serverless computing suffers from performance degradation caused by microarchitectural cold start. Previous studies identify the front-end as the bottleneck and explore solutions such as instruction prefetching and restoring Branch Target Buffer. However, they fail to prewarm the Conditional Branch Predictor (CBP), because the large size of its core component, the TAgged GEometric history length predictor (TAGE), makes it impractical to be saved and restored.This paper observes the predictive sparsity of TAGE, where only a small subset of entries can dominate the predictor’s coverage and accuracy. We introduce Khepri, a memory efficient CBP prewarming mechanism that uses a TAGE Crystallization algorithm to identify these dominant entries. Khepri records them in main memory and restores them to prewarm TAGE at the next invocation. Khepri achieves a 1.57× speedup over the baseline and outperforms the state-of-the-art technique by 14%, requiring only 1.54KB in main memory on average.
Zengshi Wang, Zhuoyuan Yang, Kanheng Jiang, Jun Han 0003
DATE3
2026 ChatArch: A Knowledge-driven Graph-of-thought LLM Framework for Processor Architecture Optimization
abstract
Processors serve as the cornerstone of modern computing systems. Although processor design encompasses multiple VLSI levels, the architectural design plays a critical role in determining performance, power consumption, and area efficiency. To address the growing pressure to shorten chip time-to-market, there is an increasing demand for rapid iteration methods in processor architecture development. To achieve efficient and effortless architecture design optimization, we develop ChatArch, a knowledge-driven graph-of-thought multi-LLM-agent framework for processor architecture optimization. Based on processor architecture expertise, we decompose the processor architecture design space and construct an LLM agent graph-of-thought framework to characterize and iteratively optimize these subspaces. Also, by systematically consolidating domain-specific knowledge and empirical design principles validated by experts, we establish a comprehensive RISC-V processor design knowledge repository. Moreover, a knowledge-driven multi-agent framework is developed to enable efficient microarchitecture optimization. Finally, the optimized microarchitecture modules aggregate to form system-level designs. This methodology achieves automated iterative optimization of microarchitectures targeting PPA objectives while generating corresponding behavioral models. The experiments demonstrate that our method effectively designs behavioral processor models, with LLM-generated architectures achieving a validation success rate of over 97.39%, surpassing the performance of other LLMs, including GPT-4o. ChatArch consistently meets requirements, delivering up to 9.97% times better PPA and 32–68x efficiency gains compared with traditional black-box optimization methods.
Zhuochu Yang, Zhuoyuan Yang, Li Shang 0001, Fan Yang 0001
ACM Trans. Design Autom. Electr. Syst.3
2025 AcclMT: A Highly Resource-Efficient and Flexible Poseidon Hash-Based Merkle Tree Architecture
abstract
Merkle Tree is a fundamental cryptographic primitive in Zero-Knowledge Proof (ZKP) protocols, sharing significant computational workloads with the Number Theoretic Transform (NTT) in zkSTARK schemes. Merkle Tree is a tree structure where nodes are primarily generated through hash computations. Among them, Poseidon Hash, as a ZK-friendly hash function, has emerged as one of the most widely adopted choices. Therefore, hardware acceleration of building Merkle Tree based on Poseidon Hash can significantly enhance the performance of ZKP protocols. We propose AcclMT, a highly resourceefficient and flexible Poseidon Hash-based Merkle Tree architecture. Our design employs hardware-software co-design and optimizes the hashing data flow, resulting in an area-efficient Poseidon Hash engine that improves modular multiplication resource utilization. Furthermore, AcclMT uses these engines alongside hierarchical on-chip cache and optimized task scheduling for building large Merkle Trees. It also supports flexible parameter configurations for various requirements. Experimental results show that our proposed Poseidon Hash engine achieves a $14.3 \times$ speedup compared to the latest FPGA-based work. By improving resource utilization, it also reduces area usage by 14.8% compared to unoptimized design. AcclMT achieves up to $1665 \times$ speedup over software implementations in building Merkle tree, with average utilization of 95.9% and 99.2% for the two hash engines.
Changxu Liu, Hao Zhou 0015, Zhuoyuan Yang, Yinlong Li, Shiyong Wu, Fan Yang 0001
DAC6
2025 VLSUMaP: A High-Performance Matrix Processor with Virtually Expanded LSU Boosting HBM Bandwidth Utilization
Xinjie Kong, Zikang Zhou, Zhuoyuan Yang, Zengshi Wang, Jun Han 0003
ACM Great Lakes Symposium on VLSI6
2023 SVP: Safe and Efficient Speculative Execution Mechanism through Value Prediction
abstract
Speculative execution attacks such as Spectre and Meltdown exploit the wrong execution patch to leak private data. In current state-of-the-art defense strategies, executions of all memory accesses that use speculatively-loaded addresses are blocked, resulting in high overhead. Our key observation is that these blocked memory accesses can be executed without operand-dependent hardware resource usage through value prediction. Therefore, we propose a novel hardware defense framework, named Speculative Value Prediction (SVP), to safely and efficiently execute the potentially unsafe memory accesses earlier. We build SVP on the cycle-accurate Gem5 simulator and its performance improvement is positively correlated with the coverage of value predictors. Experiments show that when using the value predictor with 30%/60%/100% coverage, SVP outperforms the state-of-the-art defense mechanism STT in the Spectre model by 21.5%/50.3%/107.7% respectively, and in the Futuristic model by 28.7%/55.4%/105.7% respectively.
Xinyu Qin, Zhuoyuan Yang, Weiliang He, Yifan Liu 0017, Jun Han 0003
ACM Great Lakes Symposium on VLSI3