Zixuan Zhu 0001

dblp:218/7058-1 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0000-2925-7975ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PacViT: An Efficient ViT Accelerator with Native Dynamic Pruning and Sliding Cache Attention
Jun Gong 0003, Zixuan Zhu 0001, Yongxin Zhu 0001
ISCAS2
2025 Bit-Sparsity Aware Acceleration With Compact CSD Code on Generic Matrix Multiplication
abstract
The ever-increasing demand for matrix multiplication in artificial intelligence (AI) and generic computing emphasizes the necessity of efficient computing power accommodating both floating-point (FP) and quantized integer (QINT). While state-of-the-art bit-sparsity-aware acceleration techniques have demonstrated impressive performance and efficiency in neural networks through software-driven methods such as pruning and quantization, these approaches are not always feasible in typical generic computing scenarios. In this paper, we propose Bit-Cigma, a hardware-centric architecture that leverages bit-sparsity to accelerate generic matrix multiplication. Bit-Cigma features (1) CCSD encoding, an optimized on-chip sparsification technique based on canonical signed digit (CSD) representation; (2) segmented dot product, a multi-stage exponent matching technique for long FP vectors; and (3) the versatility to efficiently process both FP and QINT data types. CCSD encoding halves the cost of CSD encoding while achieving optimal bit-sparsity, and segmented dot product improves both accuracy and throughput. Bit-Cigma cores are implemented using 65 nm technology at 1 GHz, demonstrating substantial gains in performance and efficiency for both FP and QINT configurations. Compared to state-of-the-art Bitlet, Bit-Cigma achieves 3.2$\boldsymbol{\times}$performance, 6.1$\boldsymbol{\times}$area efficiency, and 15.3$\boldsymbol{\times}$energy efficiency when processing FP32 data while ensuring zero computing error.
Zixuan Zhu 0001, Chundong Wang 0001, Zunkai Huang, Yongxin Zhu 0001
IEEE Trans. Computers1
2021 Energy-Efficient Spin-Orbit Torque MRAM Operations for Neural Network Processor
abstract
Emerging energy-efficient neural network processor is a promising hardware design to accelerate neural network algorithms with high performance and low power consumption. Typically, static random-access memory (SRAM) is employed to develop large buffers using in the processor. The bit cell of SRAM contains six transistors, leading to low density and large leakage current. In particular, several AI processors need multiple port and transfer-based SRAMs, which decrease the density and increase the power consumption. Recently, emerging spin-orbit torque magnetic random-access memory (SOT-MRAM) becomes a possible solution to replace the SRAM as working memory. However, more operations should be supported by the SOT- MRAM to provide sufficient functions, such as multiple-port memory, transpose memory, data-streaming operations. In this paper, we develop the working memory of neural network processor with SOT-MRAM to build the design library including the transpose operations, multiple-port memory, and data-streaming based buffer arrays. Equiped with those operations provided by SOT-MRAM, we can build high performance and energy-efficient neural network processors.
Liang Chang 0002, Zixuan Zhu 0001, Zhen Zhu 0005, Siqi Yang 0002, Weihang Li, Jun Zhou 0017
ISCAS2
2021 Distilling Bit-level Sparsity Parallelism for General Purpose Deep Learning Acceleration
abstract
Along with the rapid evolution of deep neural networks, the ever-increasing complexity imposes formidable computation intensity to the hardware accelerator. In this paper, we propose a novel computing philosophy called “bit interleaving” and the associate accelerator design called “Bitlet” to maximally exploit the bit-level sparsity. Apart from existing bit-serial/parallel accelerators, Bitlet leverages the abundant “sparsity parallelism” in the parameters to enforce the inference acceleration. Bitlet is versatile by supporting diverse precisions on a single platform, including floating-point 32 and fixed-point from 1b to 24b. The versatility enables Bitlet feasible for both efficient inference and training. Empirical studies on 12 domain-specific deep learning applications highlight the following results: (1) up to 81 × /21 × energy efficiency improvement for training/inference over recent high performance GPUs; (2) up to 15 × /8 × higher speedup/efficiency over state-of-the-art fixed-point accelerators; (3) 1.5mm2 area and scalable power consumption from 570mW (float32) to 432mW (16b) and 365mW (8b) @28nm TSMC; (4) highly configurable justified by ablation and sensitivity studies.
Liang Chang 0002, Zixuan Zhu 0001, Shengjian Lu, Yanhuan Liu, Mingzhe Zhang 0005
MICRO4
2021 Energy-efficient computing-in-memory architecture for AI processor: device, circuit, architecture perspective
Liang Chang 0002, Zhaomin Zhang, Jianbiao Xiao, Zhen Zhu 0005, Weihang Li, Zixuan Zhu 0001, Siqi Yang 0002, Jun Zhou 0017
Sci. China Inf. Sci.8