Zhiheng Yue

dblp:287/1266 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0003-4084-3478ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 5 first-author · 15 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 HR-DCIM: High-Reliability Floating-Point Digital CIM Architecture With Unified Low-Cost Iterative Error Correction
abstract
Digital computing-in-memory (CIM) is a promising computing paradigm for the neural network (NN) acceleration. However, during the actual deployment process of digital CIM chips, we find that existing digital CIM designs face severe computing reliability issues, which are crucial for real product development but remain underexplored. Therefore, this work pioneers a systematic computing reliability analysis for digital CIM across off-memory and in-memory levels. We find that both the off-memory floating-point (FP) exponent alignment and the in-memory random cell bit-flip errors impair digital CIM's computing reliability, causing significant truncation and bit-flip accuracy loss. Critically, existing reliability solutions are incompatible with the unique multi-row accumulation structure of digital CIM, which either severely damage digital CIM's performance or result in prohibitive overhead. To address above challenges, we propose HR-DCIM: a highreliability FP digital CIM architecture featuring unified lowcost iterative error correction. Specifically, for the off-memory reliability, we propose an exponent-mantissa joint-alignment mechanism to repurpose inherent invalid bits of aligned mantissas as compensation bits to reduce alignment truncation loss, without damaging digital CIM's performance. Then, for the in-memory reliability, we propose a remainder aliasing-based unified multiply-accumulation (MAC) error correction mechanism to correct possible MAC errors caused by various cell error cases with low-cost iteration. Experimental results show that the proposed techniques enable digital CIM to maintain high performance and efficiency across various operating voltage conditions without significant accuracy loss.
Yiqi Wang 0005, Zhiheng Yue, Zihan Wu 0006, Huiming Han, Shaojun Wei, Yang Hu 0001, Fengbin Tu, Shouyi Yin
HPCA3
2026 ReThermal: Co-Design of Thermal-Aware Static and Dynamic Scheduling for LLM Training on Liquid-Cooled Wafer-Scale Chips
abstract
With the increasing demand for high computational power in Large Language Models, wafer-scale chips have emerged as a solution, providing the necessary integration and computing capability to meet these needs. However, their ultralarge area and extreme heat dissipation introduce critical thermal management challenges under liquid-cooling environments. In addressing this issue, we identify two key opportunities and three major challenges: the behavior-thermal black box, the waferscale simulation bottleneck, and the runtime heat-schedule drift. To tackle these challenges, we propose ReThermal, a holistic scheduling framework that integrates three innovations. First, we introduce behavior-driven thermal modeling to capture workload-induced compute, communication, and heat coupling patterns at the system level. Second, we develop a DNNaccelerated wafer-scale thermal simulator that enables fast and accurate temperature prediction, significantly reducing simulation time. Third, we implement an adaptive thermal-aware scheduling strategy that coordinates compile-time and runtime decisions to dynamically optimize task placement. Evaluations show that ReThermal reduces peak temperature by up to 8.0° C and improves throughput by up to 39.23 %, providing a scalable and effective thermal control solution for future liquid-cooled wafer-scale systems.
Chengran Li, Huizheng Wang, Zhiheng Yue, Shenfei Jiang, Jinyi Deng, Yang Hu 0001, Shouyi Yin
HPCA5
2026 PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion
abstract
Attention-based models have revolutionized AI, but the quadratic cost of self-attention incurs severe computational and memory overhead. Sparse attention methods alleviate this by skipping low-relevance token pairs. However, current approaches lack practicality due to the heavy expense of added sparsity predictor, which severely drops their hardware efficiency. This paper advances the state-of-the-art (SOTA) by proposing a bit-serial enable stage-fusion (BSF) mechanism, which eliminates the need for a separate predictor. However, it faces key challenges: 1) Inaccurate bit-sliced sparsity speculation leads to incorrect pruning; 2) Hardware under-utilization due to finegrained and imbalanced bit-level workloads. 3) Tiling difficulty caused by the row-wise dependency in sparsity pruning criteria. We propose PADE, a predictor-free algorithm-hardware codesign for dynamic sparse attention acceleration. PADE features three key innovations: 1) Bit-wise uncertainty interval-enabled guard filtering (BUI-GF) strategy to accurately identify trivial tokens during each bit round; 2) Bidirectional sparsity-based out-of-order execution (BS-OOE) to improve hardware utilization; 3) Interleaving-based sparsity-tiled attention (ISTA) to reduce both I/O and computational complexity. These techniques, combined with custom accelerator designs, enable practical sparsity acceleration without relying on an added sparsity predictor. Extensive experiments on 22 benchmarks show that PADE achieves$7.43 \times$speed up and$31.1 \times$higher energy efficiency than Nvidia H100 GPU. Compared to SOTA accelerators, PADE achieves$5.1 \times, 4.3 \times$and$3.4 \times$energy saving than Sanger, DOTA and SOFA.
Huizheng Wang, Zichuan Wang, Zhiheng Yue, Yang Wang 0089, Chao Li 0009, Yang Hu 0001, Shouyi Yin
HPCA4
2026 Designing Spatial Architectures for Sparse Attention: STAR Accelerator via Cross-Stage Tiling
abstract
Large language models (LLMs) rely on self–attention for contextual understanding, demanding high-throughput inference and large–scale token parallelism (LTPP). Existing dynamic sparsity accelerators falter under LTPP scenarios due to stage-isolated optimizations. Revisiting the end-to-end sparsity acceleration flow, we identify an overlooked opportunity: crossstage coordination can substantially reduce redundant computation and memory access. We propose STAR, a cross-stage computetation and memory–efficient algorithm–hardware co-design tailored for Transformer inference under LTPP. STAR introduces a leading-zero-based sparsity prediction using log-domain add only operations to minimize prediction overhead. It further employs distributed sorting and a sorted updating FlashAttention mechanism, guided by a coordinated tiling strategy that enables fine-grained stage interaction for improved memory efficiency and latency. These optimizations are supported by a dedicated STAR accelerator architecture, achieving up to 9.2× speedup and 71.2× energy efficiency over A100, and surpassing SOTA accelerators by up to 16.1× energy and 27.1× area efficiency gains. Further, we deploy STAR onto a multi-core spatial architecture, optimizing dataflow and execution orchestration for ultra-long sequence processing. Architectural evaluation shows that, compared to the baseline design, Spatial-STAR achieves a 20.1× throughput improvement.
Huizheng Wang, Taiquan Wei, Zichuan Wang, Xinru Tang, Zhiheng Yue, Shaojun Wei, Yang Hu 0001, Shouyi Yin
IEEE Trans. Computers6
2026 An Energy-Efficient Transformer Fine-Tuning Processor for Personalized Edge Applications
abstract
Transformer models have achieved remarkable success in various domains. Given concerns about user privacy, there is an urgent need for on-device fine-tuning of Transformer models at the edge. Transformer fine-tuning faces three key challenges: 1)$O(n^{3})$re-computations during BP/WG save only$O(n^{2})$storage, limiting batch size for fine-tuning speedup. 2)Weakly related tokens account for 87.9% of computations but contribute only 8.7% to accuracy. 3)89.1% of multiplications in matrix multiplications (MM) involve dual near-zero operands, leading to a$1.9\times $increase in logic toggling energy due to frequent exponent/mantissa variations near zero. This paper proposes a Transformer-based processor supporting energy-efficient fine-tuning with three key features to tackle the above challenges. 1)An exponent-stationary re-computing scheduler (ESRS) reduces 44.2% of the storage requirement for each batch. 2)An aggressive linear fitting unit (ALFU) saves 47.4% of the computations in each iteration. 3)A logarithmic domain processing element (LDPE) decreases 36.3% of energy for MM in fine-tuning. Fabricated with 22nm technology, the proposed processor has an area of 6.4 mm2. The proposed Transformer processor achieves a peak energy efficiency of 54.94 TFLOPS/W. It reduces fine-tuning energy by$4.27\times $and offers$3.57\times $speedup for GPT-2.
Yang Wang 0089, Yubin Qin, Wende Xu, Zhiheng Yue, Huiming Han, Shaojun Wei, Yang Hu 0001, Shouyi Yin
IEEE Trans. Circuits Syst. I Regul. Pap.6
2025 MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
Huizheng Wang, Zichuan Wang, Zhiheng Yue, Yousheng Long, Taiquan Wei, Jianxun Yang, Yang Wang 0089, Chao Li 0009, Shaojun Wei, Yang Hu 0001, Shouyi Yin
MICRO3
2025 3D-PATH: A Hierarchy LUT Processing-in-memory Accelerator with Thermal-aware Hybrid Bonding Integration
abstract
LUT-based processing-in-memory (PIM) architectures enable generalpurpose in-situ computing by retrieving precomputed results.However, they suffer from limited computing precision, redundancy, and high latency of off-table access.To address these challenges, we present 3D-PATH, a novel PIM architecture that employs 3D hybrid bonding to integrate a DRAM-LUT, enhancing system capacity and reducing access latency.To further optimize efficiency, 3D-PATH introduces a hierarchical fast-LUT design that reduces storage redundancy and accelerates computation.Additionally, 3D-PATH extends computing precision by efficiently supporting floating-point operations via representation transformation and parallel interleaving banks.While hybrid bonding offers significant benefits, it induces heat dissipation challenges.To address this, we implement thermal-aware hardware that ensures the DRAM Die temperature maintains below the threshold of 85°C.Evaluations on arithmetic and AI workloads demonstrate that 3D-PATH achieves up to 12.68× higher throughput than GPUs and 2.27-7.54×over prior LUT-PIMs, while delivering a 12.24× improvement in floating-point energy efficiency over GPU and 2.13× over a 3D baseline.
Zhiheng Yue, Yang Wang 0089, Chao Li 0009, Shaojun Wei, Yang Hu 0001, Shouyi Yin
MICRO1
2025 Dyn-Bitpool: A 28 nm 27 TOPS/W Two-Sided Sparse CIM Accelerator Featuring a Balanced Workload Scheme and High CIM Macro Utilization
abstract
Deep neural networks (DNNs) have brought about a transformative impact across various sectors. However, the proliferation of DNNs has led to a surge in computational intensity and data traffic, thereby imposing substantial demands on the power capacity and battery life of computing systems. Computing-in-memory (CIM) is considered a promising architecture to resolve or mitigate the memory wall challenge by integrating computational elements within memory arrays. Yet prior studies on CIM have seldom capitalized on sparsity in both activations and weights simultaneously. Furthermore, the exploitation of two-sided sparsity—sparsity in both activations and weights—presents new challenges, such as imbalanced workload and low hardware substrate utilization. To harness the full potential of two-sided sparsity for acceleration, we present Dyn-Bitpool, an accelerator that introduces innovations on two fronts: 1) a balanced workload scheme, “pool first and cross lane sharing”, which maximizes performance gains enabled by the bit-level sparsity in activations; and 2) a dynamic topology for CIM arrays to effectively address the low CIM macro utilization issue caused by the value-level sparsity in weights. These collective advancements yield an average speedup of 1.91x and 2.67x for Dyn-Bitpool on eight prevalent neural networks, outperforming two cutting-edge CIM-based accelerators.
Xujiang Xiang, Zhiheng Yue, Shaojun Wei, Yang Hu 0001, Shouyi Yin
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 Dyn-Bitpool: A Two-sided Sparse CIM Accelerator Featuring a Balanced Workload Scheme and High CIM Macro Utilization
abstract
Computing-in-memory (CIM), a promising computing paradigm, has demonstrated great energy-efficiency by integrating computing units into memory. However, previous research on CIM has rarely utilized sparsity in activation and weight concurrently. Moreover, new challenges arise when harnessing sparsity in both activation and weight (two-sided sparsity), such as unbalanced workload and low hardware substrate utilization.
Xujiang Xiang, Zhiheng Yue, Liuxin Lv, Shaojun Wei, Yang Hu 0001, Shouyi Yin
DAC2
2024 CAP: A General Purpose Computation-in-memory with Content Addressable Processing Paradigm
abstract
Demands for efficient computing under memory wall have led to computation-in-memory (CIM) accelerators that leverage memory structure to perform in-situ computing. The content addressable memory (CAM) processing is a CIM paradigm that accomplishes general purpose functions, via sequences of search and update operations on CAM. However, the conventional CAM-based CIM is customized for vector-vector operation only and requires long search-update iterations for computing.
Zhiheng Yue, Shaojun Wei, Yang Hu 0001, Shouyi Yin
DAC1
2024 Exploiting Similarity Opportunities of Emerging Vision AI Models on Hybrid Bonding Architecture
abstract
While extensive research has focused on optimizing performance and efficiency in vision-based AI accelerators, an unexplored phenomenon, Clustering Similarity Effect, presents a significant opportunity for further improvement. This effect reveals that clusters of neighboring data points exhibit similar values, enabling the potential to skip redundant computations.To fully capitalize on the potential of the Clustering Similarity Effect (CSE), this work integrates hybrid bonding DRAM technology. We conduct a comprehensive analysis of the associated design considerations and integration overhead. Leveraging these insights, we propose a novel CSE-aware architecture specifically tailored for hybrid bonding memory. This architecture facilitates similarity detection and adapts to the inherent data characteristics associated with CSE.Compared with state-of-the-art 2D/2.5D AI accelerators, the hybrid bonding baseline demonstrates an average energy efficiency improvement of $2.89 \times \sim 14.28 \times$ and an area efficiency improvement of $2.67 \times \sim 7.68 \times$. Incorporating the similarity optimizations further enhances energy efficiency and area efficiency improvement to $5.69 \times \sim 28.13 \times$ and $3.82 \times \sim 10.98 \times$, respectively.
Zhiheng Yue, Huizheng Wang, Jiahao Fang, Jinyi Deng, Guangyang Lu, Fengbin Tu, Yubin Qin, Yang Wang 0089, Chao Li 0009, Huiming Han, Shaojun Wei, Yang Hu 0001, Shouyi Yin
ISCA1
2024 SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated Tiling
abstract
Benefiting from the self-attention mechanism, Transformer models have attained impressive contextual comprehension capabilities for lengthy texts. The requirements of high-throughput inference arise as the large language models (LLMs) become increasingly prevalent, which calls for large-scale token parallel processing (LTPP). However, existing dynamic sparse accelerators struggle to effectively handle LTPP, as they solely focus on separate stage optimization, and with most efforts confined to computational enhancements. By re-examining the end-to-end flow of dynamic sparse acceleration, we pinpoint an ever-overlooked opportunity that the LTPP can exploit the intrinsic coordination among stages to avoid excessive memory access and redundant computation. Motivated by our observation, we present SOFA, a cross-stage compute-memory efficient algorithm-hardware co-design, which is tailored to tackle the challenges posed by LTPP of Transformer inference effectively. We first propose a novel leading zero computing paradigm, which predicts attention sparsity by using log-based add-only operations to avoid the significant overhead of prediction. Then, a distributed sorting and a sorted updating FlashAttention mechanism are proposed with cross-stage coordinated tiling principle, which enables fine-grained and lightweight coordination among stages, helping optimize memory access and latency. Further, we propose a SOFA accelerator to support these optimizations efficiently. Extensive experiments on 20 benchmarks show that SOFA achieves$9.5\times$speed up and$71.5\times$higher energy efficiency than Nvidia A100 GPU. Compared to eight SOTA accelerators, SOFA achieves an average$15.8\times$energy efficiency,$10.3\times$area efficiency and$9.3\times$speed up, respectively.
Huizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue, Yubin Qin, Sihan Guan, Qinze Yang, Yang Wang 0089, Chao Li 0009, Yang Hu 0001, Shouyi Yin
MICRO4
2023 STAR: An STGCN ARchitecture for Skeleton-Based Human Action Recognition
abstract
Skeleton-based human action cognition (HAR) has drawn increasing attention recently. As an emerging approach for skeleton-based HAR tasks, Spatial-Temporal Graph Convolution Network (STGCN) achieves remarkable performance by fully exploiting the skeleton topology information via graph convolution. Unfortunately, existing GCN accelerators lose efficiency when processing STGCN models due to two limitations. (1) At the dataflow level, the hardware parallelism of GCN accelerators cannot match the computation parallelism of STGCN models, leading to computing resource under-utilization. (2) At the computation level, GCN accelerators fail to exploit the inherent temporal redundancy in STGCN models. To overcome the limitations, this paper proposes STAR, an STGCN architecture for skeleton-based human action recognition. STAR is designed based on the characteristics of different computation phases in STGCN. For limitation (1), a spatial-temporal dimension consistent (STDC) dataflow is proposed to fully exploit the data reuse opportunities in all the different dimensions of STGCN. For limitation (2), we propose a node-wise exponent sharing scheme and a temporal-structured redundancy elimination mechanism, to exploit the inherent temporal redundancy specially introduced by STGCN. To further address the under-utilization induced by redundancy elimination, we design a dynamic data scheduler to manage the feature data storage and schedule the features and weights for valid computation in real time. STAR achieves$4.48\times $,$5.98\times $,$2.54\times $, and$103.88\times $energy savings on average over the HyGCN, AWB-GCN, TPU, and Jetson TX2 GPU.
Fengbin Tu, Mengqi Niu, Zhiheng Yue, Leibo Liu, Shaojun Wei, Yang Hu 0001, Shouyi Yin
IEEE Trans. Circuits Syst. I Regul. Pap.4
2022 MC-CIM: a reconfigurable computation-in-memory for efficient stereo matching cost computation
abstract
This paper proposes the design of a computation-in-memory for stereo matching cost computation. The matching cost computation incurs large energy and latency overhead because of frequent memory access. To overcome previous design limitations, this work, named MC-CIM, performs matching cost computation without incurring memory access and introduces several key features. (1) Lightweight balanced computing unit is integrated within cell array to reduce memory access and improve system throughput. (2) Self-optimized circuit design enables to alter arithmetic operation for matching algorithm in various scenario. (3) Flexible data mapping method and reconfigurable digital peripheral explore maximum parallelism on different algorithm and bit-precision. The proposed design is implemented in 28nm technology and achieves average performance of 277 TOPs/W.
Zhiheng Yue, Leibo Liu, Shaojun Wei, Shouyi Yin
DAC1
2022 BR-CIM: An Efficient Binary Representation Computation-In-Memory Design
abstract
Deep neural network (DNN) has recently attracted tremendous attention in various fields. But the computing operation requirement and the memory bottleneck limit the energy efficiency of hardware implementations. Binary quantization is proposed to relieve the pressure of hardware design. And the Computing-In-Memory (CIM) is regarded as a promising method to resolve the memory wall challenge. However, the binary computing paradigm is mismatched with the CIM scheme, which incurs complex circuits and peripheral to realize binary operation in previous works. To overcome previous issues, this work presents Binary Representation Computation-In-Memory (BR-CIM) with several key features. (1) A lightweight computation unit is realized within the 6T SRAM array to accelerate binary computing and enlarge signal margin; (2) The reconfigurable computing scheme and mapping method support extendable bit precision to satisfy the accuracy requirement of various applications; (3) Simultaneous computing and weight loading is supported by column circuitry, which shortens the data loading latency; Several experiments are conducted to estimate algorithm accuracy, the computing latency, and power consumption. The energy efficiency reaches up to 1280 TOPs/W for binary representation. And the algorithm accuracy achieves 97.82%/76.4% on MNIST/CIFAR-100 dataset.
Zhiheng Yue, Yubin Qin, Leibo Liu, Shaojun Wei, Shouyi Yin
IEEE Trans. Circuits Syst. I Regul. Pap.1