Huizheng Wang

dblp:224/5514 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0002-9763-8208ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 9 first-author · 18 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 MOCAP: Wafer-Scale-Chip-Oriented Memory-Orchestrated Chunked Pipelining Framework for Prefill-Only LLM Inference
Zichuan Wang, Huizheng Wang, Yuheng Xiao, Haonan Zuo, Taiquan Wei, Jinyi Deng, Yang Hu 0001, Shouyi Yin
APPT2
2026 LAPA: Log-Domain Prediction-Driven Dynamic Sparsity Accelerator for Transformer Model
Huizheng Wang, Shaojun Wei, Yang Hu 0001, Shouyi Yin
ASP-DAC1
2026 BitStopper: An Efficient Transformer Attention Accelerator via Stage-fusion and Early Termination
Huizheng Wang, Shaojun Wei, Yang Hu 0001, Shouyi Yin
ASP-DAC1
2026 ReThermal: Co-Design of Thermal-Aware Static and Dynamic Scheduling for LLM Training on Liquid-Cooled Wafer-Scale Chips
abstract
With the increasing demand for high computational power in Large Language Models, wafer-scale chips have emerged as a solution, providing the necessary integration and computing capability to meet these needs. However, their ultralarge area and extreme heat dissipation introduce critical thermal management challenges under liquid-cooling environments. In addressing this issue, we identify two key opportunities and three major challenges: the behavior-thermal black box, the waferscale simulation bottleneck, and the runtime heat-schedule drift. To tackle these challenges, we propose ReThermal, a holistic scheduling framework that integrates three innovations. First, we introduce behavior-driven thermal modeling to capture workload-induced compute, communication, and heat coupling patterns at the system level. Second, we develop a DNNaccelerated wafer-scale thermal simulator that enables fast and accurate temperature prediction, significantly reducing simulation time. Third, we implement an adaptive thermal-aware scheduling strategy that coordinates compile-time and runtime decisions to dynamically optimize task placement. Evaluations show that ReThermal reduces peak temperature by up to 8.0° C and improves throughput by up to 39.23 %, providing a scalable and effective thermal control solution for future liquid-cooled wafer-scale systems.
Chengran Li, Huizheng Wang, Zhiheng Yue, Shenfei Jiang, Jinyi Deng, Yang Hu 0001, Shouyi Yin
HPCA2
2026 MoEntwine: Unleashing the Potential of Wafer-Scale Chips for Large-Scale Expert Parallel Inference
abstract
As large language models (LLMs) continue to scale up, mixture-of-experts (MoE) has become a common technology in SOTA models. MoE models rely on expert parallelism (EP) to alleviate memory bottleneck, which introduces all-to-all communication to dispatch and combine tokens across devices. However, in widely-adopted GPU clusters, high-overhead crossnode communication makes all-to-all expensive, hindering the adoption of EP. Recently, wafer-scale chips (WSCs) have emerged as a platform integrating numerous devices on a wafer-sized interposer. WSCs provide a unified high-performance network connecting all devices, presenting a promising potential for hosting MoE models. Yet, their network is restricted to a mesh topology, causing imbalanced communication pressure and performance loss. Moreover, the lack of on-wafer disk leads to high-overhead expert migration on the critical path. To fully unleash this potential, we first propose Entwined Ring Mapping (ER-Mapping), which co-designs the mapping of attention and MoE layers to balance communication pressure and achieve better performance. We find that under ER-Mapping, the distribution of cold and hot links in the attention and MoE layers is complementary. Therefore, to hide the migration overhead, we propose the Non-invasive Balancer (NI-Balancer), which splits a complete expert migration into multiple steps and alternately utilizes the cold links of both layers. Evaluation shows ER-Mapping achieves communication reduction up to 62 %. NIBalancer further delivers 54 % and 22 % improvements in MoE computation and communication, respectively. Compared with the SOTA NVL72 supernode, the WSC platform delivers an average 39 % higher per-device MoE performance owing to its scalability to larger EP.
Xinru Tang, Jingxiang Hou, Dingcheng Jiang, Taiquan Wei, Jinyi Deng, Huizheng Wang, Qize Yang, Haoran Shang, Chao Li 0009, Yang Hu 0001, Shouyi Yin
HPCA7
2026 WATOS: Efficient LLM Training Strategies and Architecture Co-Exploration for Wafer-Scale Chip
abstract
Training large language models (LLMs) imposes extreme demands on computation, memory capacity, and interconnect bandwidth, driven by their ever-increasing parameter scales and intensive data movement. Wafer-scale integration offers a promising solution by densely integrating multiple single-die chips with high-speed die-to-die (D2D) interconnects. However, the limited wafer area necessitates trade-offs among compute, memory, and communication resources. Fully harnessing the potential of wafer-scale integration while mitigating its architectural constraints is essential for maximizing LLM training performance. This imposes significant challenges for the co-optimization of architecture and training strategies. Unfortunately, existing approaches all fall short in addressing these challenges. To bridge the gap, we propose WATOS, a co-exploration framework for LLM training strategy and wafer-scale architecture. We first define a highly configurable hardware template designed to explore optimal architectural parameters for waferscale chips. Based on it, we capitalize on the high D2D bandwidth and fine-grained operation advantages inherent to wafer-scale chips to explore optimal parallelism and resource allocation strategies, effectively addressing the memory underutilization issues during LLM training. Compared to the state-of-the-art (SOTA) LLM training framework Megatron and Cerebras' weight streaming wafer training strategy, WATOS can achieve an average overall throughput improvement of$2.74 \times$and$1.53 \times$across various LLM models, respectively. In addition, we leverage WATOS to reveal intriguing insights about wafer-scale architecture design with the training of LLM workloads.
Huizheng Wang, Zichuan Wang, Jingxiang Hou, Taiquan Wei, Chao Li 0009, Yang Hu 0001, Shouyi Yin
HPCA1
2026 TEMP: A Memory Efficient Physical-Aware Tensor Partition-Mapping Framework on Wafer-Scale Chips
abstract
Large language models (LLMs) demand significant memory and computation resources. Wafer-scale chips (WSCs) provide high computation power and die-to-die (D2D) bandwidth but face a unique trade-off between on-chip memory and compute resources due to limited wafer area. Therefore, tensor parallelism strategies for wafer should leverage communication advantages while maintaining memory efficiency to maximize WSC performance. However, existing approaches fail to address these challenges. To address these challenges, we propose the tensor stream partition paradigm (TSPP), which reveals an opportunity to leverage WSCs' abundant communication bandwidth to alleviate stringent on-chip memory constraints. However, the 2D mesh topology of WSCs lacks long-distance and flexible interconnects, leading to three challenges: 1) severe tail latency, 2) prohibitive D2D traffic contention, and 3) intractable search time for optimal design. We present TEMP, a framework for LLM training on WSCs that leverages topology-aware tensor-stream partition, trafficconscious mapping, and dual-level wafer solving to overcome hardware constraints and parallelism challenges. These integrated approaches optimize memory efficiency and throughput, unlocking TSPP's full potential on WSCs. Evaluations show TEMP achieves$1.7 \times$average throughput improvement over state-of-the-art LLM training systems across various models.
Huizheng Wang, Taiquan Wei, Zichuan Wang, Dingcheng Jiang, Qize Yang, Jingxiang Hou, Chao Li 0009, Jinyi Deng, Yang Hu 0001, Shouyi Yin
HPCA1
2026 PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion
abstract
Attention-based models have revolutionized AI, but the quadratic cost of self-attention incurs severe computational and memory overhead. Sparse attention methods alleviate this by skipping low-relevance token pairs. However, current approaches lack practicality due to the heavy expense of added sparsity predictor, which severely drops their hardware efficiency. This paper advances the state-of-the-art (SOTA) by proposing a bit-serial enable stage-fusion (BSF) mechanism, which eliminates the need for a separate predictor. However, it faces key challenges: 1) Inaccurate bit-sliced sparsity speculation leads to incorrect pruning; 2) Hardware under-utilization due to finegrained and imbalanced bit-level workloads. 3) Tiling difficulty caused by the row-wise dependency in sparsity pruning criteria. We propose PADE, a predictor-free algorithm-hardware codesign for dynamic sparse attention acceleration. PADE features three key innovations: 1) Bit-wise uncertainty interval-enabled guard filtering (BUI-GF) strategy to accurately identify trivial tokens during each bit round; 2) Bidirectional sparsity-based out-of-order execution (BS-OOE) to improve hardware utilization; 3) Interleaving-based sparsity-tiled attention (ISTA) to reduce both I/O and computational complexity. These techniques, combined with custom accelerator designs, enable practical sparsity acceleration without relying on an added sparsity predictor. Extensive experiments on 22 benchmarks show that PADE achieves$7.43 \times$speed up and$31.1 \times$higher energy efficiency than Nvidia H100 GPU. Compared to SOTA accelerators, PADE achieves$5.1 \times, 4.3 \times$and$3.4 \times$energy saving than Sanger, DOTA and SOFA.
Huizheng Wang, Zichuan Wang, Zhiheng Yue, Yang Wang 0089, Chao Li 0009, Yang Hu 0001, Shouyi Yin
HPCA1
2026 Designing Spatial Architectures for Sparse Attention: STAR Accelerator via Cross-Stage Tiling
abstract
Large language models (LLMs) rely on self–attention for contextual understanding, demanding high-throughput inference and large–scale token parallelism (LTPP). Existing dynamic sparsity accelerators falter under LTPP scenarios due to stage-isolated optimizations. Revisiting the end-to-end sparsity acceleration flow, we identify an overlooked opportunity: crossstage coordination can substantially reduce redundant computation and memory access. We propose STAR, a cross-stage computetation and memory–efficient algorithm–hardware co-design tailored for Transformer inference under LTPP. STAR introduces a leading-zero-based sparsity prediction using log-domain add only operations to minimize prediction overhead. It further employs distributed sorting and a sorted updating FlashAttention mechanism, guided by a coordinated tiling strategy that enables fine-grained stage interaction for improved memory efficiency and latency. These optimizations are supported by a dedicated STAR accelerator architecture, achieving up to 9.2× speedup and 71.2× energy efficiency over A100, and surpassing SOTA accelerators by up to 16.1× energy and 27.1× area efficiency gains. Further, we deploy STAR onto a multi-core spatial architecture, optimizing dataflow and execution orchestration for ultra-long sequence processing. Architectural evaluation shows that, compared to the baseline design, Spatial-STAR achieves a 20.1× throughput improvement.
Huizheng Wang, Taiquan Wei, Zichuan Wang, Xinru Tang, Zhiheng Yue, Shaojun Wei, Yang Hu 0001, Shouyi Yin
IEEE Trans. Computers1
2026 CrossKV: Accelerating Large Language Model Inference via Cross-Stage Dynamic Co-Optimization for KV Cache
abstract
The autoregressive nature of Large Language Models (LLMs) has enabled remarkable performance in language generation, making them a cornerstone in natural language processing. As context windows lengthen, the per-token key–value (KV) cache grows linearly with sequence length and turns the generation stage into a memory-bound operation. However, existing works focus primarily on an isolated stage of KV cache reduction, and most efforts incur significant hardware overheads. This results in a lack of cross-stage co-optimization and leaves significant reduction potential untapped. To address these challenges, we propose CrossKV, a software-hardware co-design architecture that accelerates LLM inference to fully achieve the potential of KV caching optimization. Motivated by our observation, we identify the three stages in KV caching and present a cross-stage co-optimization algorithm, including: a derivative-enhanced dynamic progressive pruning method, a DCT-driven key-vector low-rank compressing, and a dynamic clustered hybrid run-length encoding to reduce the KV caching while maintaining high accuracy. Then, an efficient architecture featuring cross-stage dynamic co-optimization with negligible hardware overhead is proposed to fully harness the algorithm of CrossKV. Our evaluations of CrossKV over 9 popular LLMs models and various long-context tasks demonstrate an average of$6.03\times $and$4.10\times $improvement for energy efficiency and speedup compared to existing SoTA Accelerators architecture. Compared to the Nvidia A100 GPU, CrossKV achieves an average$23.43\times $energy efficiency and$4.64\times $speedup, respectively.
Shenyu Wang, Huizheng Wang, Peng Wang 0220, Xiao Liu 0001, Zhihua Wang 0001, Yang Hu 0001, Hanjun Jiang
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 Spatial-Aware Orchestration of LLM Attention on Waferscale Chips
Taiquan Wei, Huizheng Wang, Zichuan Wang, Shouyi Yin, Yang Hu 0001
APPT2
2025 PD Constraint-aware Physical/Logical Topology Co-Design for Network on Wafer
abstract
As cluster scales for LLM training expand, waferscale chips, characterized by the high integration density and bandwidth, emerge as a promising approach to enhancing training performance.The role of Network on Wafer (NoW) is becoming increasingly significant, which puts an emphasis on two facts: physical and logical topology.However, existing networks fail to co-design both aspects.Additionally, physical topology typically focuses on optimizing communication or computation separately, neglecting opportunities to improve overall training performance.In this paper, we propose a physical design (PD) constraint-aware joint optimization strategy, developing mesh-switch physical topology and a dual-granularity logical topology.Mesh-switch leverages the high integration density of mesh and the efficient communication performance of fat tree, optimizing the allocation of on-chip
Qize Yang, Taiquan Wei, Sihan Guan, Chengran Li, Haoran Shang, Jinyi Deng, Huizheng Wang, Chao Li 0009, Yan Zhang 0163, Shouyi Yin, Yang Hu 0001
ISCA7
2025 MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
Huizheng Wang, Zichuan Wang, Zhiheng Yue, Yousheng Long, Taiquan Wei, Jianxun Yang, Yang Wang 0089, Chao Li 0009, Shaojun Wei, Yang Hu 0001, Shouyi Yin
MICRO1
2024 Exploiting Similarity Opportunities of Emerging Vision AI Models on Hybrid Bonding Architecture
abstract
While extensive research has focused on optimizing performance and efficiency in vision-based AI accelerators, an unexplored phenomenon, Clustering Similarity Effect, presents a significant opportunity for further improvement. This effect reveals that clusters of neighboring data points exhibit similar values, enabling the potential to skip redundant computations.To fully capitalize on the potential of the Clustering Similarity Effect (CSE), this work integrates hybrid bonding DRAM technology. We conduct a comprehensive analysis of the associated design considerations and integration overhead. Leveraging these insights, we propose a novel CSE-aware architecture specifically tailored for hybrid bonding memory. This architecture facilitates similarity detection and adapts to the inherent data characteristics associated with CSE.Compared with state-of-the-art 2D/2.5D AI accelerators, the hybrid bonding baseline demonstrates an average energy efficiency improvement of $2.89 \times \sim 14.28 \times$ and an area efficiency improvement of $2.67 \times \sim 7.68 \times$. Incorporating the similarity optimizations further enhances energy efficiency and area efficiency improvement to $5.69 \times \sim 28.13 \times$ and $3.82 \times \sim 10.98 \times$, respectively.
Zhiheng Yue, Huizheng Wang, Jiahao Fang, Jinyi Deng, Guangyang Lu, Fengbin Tu, Yubin Qin, Yang Wang 0089, Chao Li 0009, Huiming Han, Shaojun Wei, Yang Hu 0001, Shouyi Yin
ISCA2
2024 Code Length Compatible Belief Propagation Polar Decoder Based on Folding and Unfolding
abstract
This paper presents a code-length compatible architecture for belief propagation (BP) polar decoders. This decoder incorporates folding and unfolding techniques with control signals, allowing it to decode codes with varying code lengths. By modifying the architecture originally designed for code length N, the proposed decoder can handle codes of length 2iN, where i ∈ Z+using folding, and i ∈ Z−using unfolding. To reduce the critical path and implementation complexity, a new routing design is proposed. Moreover, we introduce a memory architecture utilizing shift registers instead of RAM to increase the throughput. We demonstrate gate-level implementations to illustrate the design’s architecture. Finally, we analyze the throughput, area, and power consumption of the decoders. Compared with traditional single-column designs with N = 1024, the proposed decoder architecture can achieve up to 136% hardware efficiency while consuming 1.6% less area.
Muhao Li, Huizheng Wang, Yifei Shen 0003, Xiaosi Tan, Chuan Zhang 0001
ISCAS2
2024 SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated Tiling
abstract
Benefiting from the self-attention mechanism, Transformer models have attained impressive contextual comprehension capabilities for lengthy texts. The requirements of high-throughput inference arise as the large language models (LLMs) become increasingly prevalent, which calls for large-scale token parallel processing (LTPP). However, existing dynamic sparse accelerators struggle to effectively handle LTPP, as they solely focus on separate stage optimization, and with most efforts confined to computational enhancements. By re-examining the end-to-end flow of dynamic sparse acceleration, we pinpoint an ever-overlooked opportunity that the LTPP can exploit the intrinsic coordination among stages to avoid excessive memory access and redundant computation. Motivated by our observation, we present SOFA, a cross-stage compute-memory efficient algorithm-hardware co-design, which is tailored to tackle the challenges posed by LTPP of Transformer inference effectively. We first propose a novel leading zero computing paradigm, which predicts attention sparsity by using log-based add-only operations to avoid the significant overhead of prediction. Then, a distributed sorting and a sorted updating FlashAttention mechanism are proposed with cross-stage coordinated tiling principle, which enables fine-grained and lightweight coordination among stages, helping optimize memory access and latency. Further, we propose a SOFA accelerator to support these optimizations efficiently. Extensive experiments on 20 benchmarks show that SOFA achieves$9.5\times$speed up and$71.5\times$higher energy efficiency than Nvidia A100 GPU. Compared to eight SOTA accelerators, SOFA achieves an average$15.8\times$energy efficiency,$10.3\times$area efficiency and$9.3\times$speed up, respectively.
Huizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue, Yubin Qin, Sihan Guan, Qinze Yang, Yang Wang 0089, Chao Li 0009, Yang Hu 0001, Shouyi Yin
MICRO1
2023 A DSP-Purposed REconfigurable Acceleration Machine (DREAM) for High Energy Efficiency MIMO Signal Processing
abstract
The wireless baseband processing algorithms are still developing and show a great diversity. The development of ASIC implementations cannot quickly adapt to the evolution of algorithms and standards. Meanwhile, the general-purpose processors cannot meet the real-time requirements in some scenarios. This paper proposes a DSP-purposed REconfigurable Acceleration Machine (DREAM) core for wireless baseband digital signal processing, which has a good trade-off between flexibility and performance. First, we abstract a set of shared operators with a moderate granularity from a variety of wireless MIMO signal processing algorithms. Then, we propose a two-step configuration process to reduce the size of the required reconfiguration bits. Besides, we design a conflict-free address generator to transfer data between the on-chip scratchpad memory and reconfiguration processing elements with high efficiency and high throughput. Finally, the prototype DREAM core has been implemented in TSMC CMOS 28 nm, and its area and power consumption have been analyzed. The chip has great flexibility in supporting a variety of wireless MIMO processing algorithms and a wide range of MIMO scales. The proposed DREAM core can achieve the normalized area efficiency and the normalized energy efficiency of$0.67~Gbps/MGE$and$15.05~Gbps/W$, which are$1.56\times $and$4.18\times $those of state-of-the-art reconfigurable implementations when running the WeJi-based MIMO detection algorithm.
Kai Chen 0034, Wenqing Song, Guoqiang He, Sirui Shen, Huizheng Wang, Chuan Zhang 0001, Li Li 0003
IEEE Trans. Circuits Syst. I Regul. Pap.6
2023 An Efficient Approximate Expectation Propagation Detector With Block-Diagonal Neumann-Series
abstract
Expectation propagation (EP) achieves near-optimal performance for large-scale multiple-input multiple-output (L-MIMO) detection, however, at the expense of unaffordable matrix inversions. To tackle the issue, several low-complexity EP detectors have been proposed. However, they all fail to exploit the properties of channel matrices, thus resulting in unsatisfactory performance in non-ideal scenarios. To this end, in this paper, a block-diagonal Neumann-series-based expectation propagation approximation (BD-NS-EPA) algorithm is proposed, which is applicable for both ideal uncorrelated channels and the correlated channels with multiple-antenna user equipment system. First, a block-diagonal-based Neumann iteration is employed, which skillfully exerts the main information of the channels while reducing computational cost. An adjustable sorting message updating scheme then is introduced to reduce the update of redundant nodes during iterations. Numerical results show that, for$128\times 32$MIMO with the non-ideal channel, the proposed algorithm exhibits 0.3 dB away from the original EP when bit error-rate (BER)$=10^{-3}$, at the cost of mere 3% normalized complexity. The implementation results on SMIC 65-nm CMOS technology suggest that the proposed detector can achieve 1.252 Gbps/W and 0.275 Mbps/kGE hardware efficiency, further demonstrating that the proposed detectors can achieve a good trade-off between error-rate performance and hardware efficiency.
Huizheng Wang, Bingyang Cheng, Xiaosi Tan, Xiaohu You 0001, Chuan Zhang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2020 Efficient stochastic successive cancellation list decoder for polar codes
Xiao Liang 0005, Huizheng Wang, Yifei Shen 0003, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001
Sci. China Inf. Sci.2
2018 NEUTag's Classification System for Zhihu Questions Tagging Task
Yuejia Xiang, Huizheng Wang, Duo Ji, Zheyang Zhang
NLPCC (1)2