EDBT 2026 Demo / reviewers in the wild / expert
Zihan Zou
dblp:372/2562
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0009-0002-1510-1341ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 3D-TANoC: Thermal-Aware 3D LLM Accelerator with Hierarchical NoC and Operator-Aware Dataflow Mapping Strategy
Xingyu Xu 0008, Dengke Liu, Zihan Zou, Xilong Kang, Hui Kou, Hao Cai 0001, Bo Liu 0019 |
ISCAS | 4 |
| 2026 | Low Bit-Width LLM Acceleration via Symmetric Lookup Format and Compute-in-Decoding Paradigm
Zihan Zou, Jiaming Lin, Xinming Yan, Shikuang Chen, Chen Zhang 0025, Xilong Kang, Hao Cai 0001, Bo Liu 0019 |
IEEE Trans. Computers | 1 |
| 2025 | TWDP: A Vision Transformer Accelerator with Token-Weight Dual-Pruning Strategy for Edge Device DeploymentabstractVision Transformers (ViTs) have attracted significant attention due to their superior accuracy compared to convolutional neural networks (CNNs) in various computer vision tasks. However, their substantial computational load and significant memory footprint lead to excessive delay and considerable data storage overhead, posing challenges for resource-limited edge device deployment. To address these issues, we present TWDP, a vision transformer accelerator employing a Token-Weight Dual-Pruning strategy to enhance the efficiency of the inference process. Firstly, we propose a parameter-free self-adaptive token pruning method to skip redundant computations in an image-dependent manner. Secondly, we apply a Hessian-aware layer-wise N:M weight pruning approach to minimize storage overhead, memory access, and computational power consumption. Additionally, to manage the complex computing patterns in ViTs, an overlapping dataflow is utilized to further reduce temporal storage and inference latency. Implemented and evaluated under an industrial 28nm technology, the proposed TWDP framework reduces 66.1% weight storage requirements and achieves an energy efficiency of 2070.9 FPS/W. Compared to state-of-the-art architectures, TWDP obtains a 1.6× energy efficiency improvement with negligible accuracy loss, demonstrating the superiority of TWDP in edge device deployment scenarios. Guang Yang 0036, Xinming Yan, Hui Kou, Zihan Zou, Qingwen Wei, Hao Cai 0001, Bo Liu 0019 |
ASP-DAC | 4 |
| 2025 | OutlierCIM: Outlier-Aware Digital CIM-Based LLM Accelerator with Hybrid-Strategy Quantization and Unified FP-INT ComputationabstractActivation outliers in Large Language Models (LLMs), which exhibit large magnitudes but small quantities, significantly affect model performance and pose challenges for the acceleration of LLMs. To address this bottleneck, researchers have proposed several co-design frameworks with outlier-aware algorithms and dedicated hardware. However, they face challenges balancing model accuracy with hardware efficiency when accelerating LLMs in a low bit-width manner. To this end, we propose OutlierCIM, the first algorithm and hardware codesign framework for the compute-in-memory (CIM) accelerator with outlier-aware quantization algorithm. The key contributions of OutlierCIM are 1) an outlier-clustered tiling strategy that regulates memory access and reduces inefficient workloads which are both introduced by outliers, 2) a hybrid-strategy quantization and a reconfigurable double-bit CIM macro array that overcome the low storage utilization and high latency of outlier-based LLM quantization, and 3) a quantization factor post-processing strategy and a dedicated quantizer that efficiently unify the multiplication and accumulation of outlier-caused FP-INT workloads. Implemented in a 28 nm CMOS technology, OutlierCIM occupies an area of $2.25 \mathrm{~mm}^{2}$. When evaluated at comprehensive benchmarks, OutlierCIM achieves up to $4.54 \times$ energy efficiency improvement and $3.91 \times$ speedup compared to the state-of-the-art outlier-aware accelerators. Zihan Zou, Shikuang Chen, Chen Zhang 0001, Xin Si, Hao Cai 0001, Bo Liu 0019 |
DAC | 1 |
| 2025 | SparCIM: A Heterogeneous CIM-Based Accelerator for Large Language Models with Contextual and Unstructured Bit SparsityabstractTransformer-based Large Language Models (LLMs) exhibit high computation and memory demands, especially on resource-constrained edge devices. Sparsification techniques like contextual sparsity have shown potential in reducing these demands by selectively pruning model components, but their integration with efficient hardware remains challenging. Firstly, sparsity prediction introduces extra energy and latency cost. Secondly, prefilling and decoding stages in LLM inference cause imbalanced workload. Thirdly, Computing-in-Memory (CIM) suffers from low utilization due to unstructured bit sparsity. To address the above challenges, this paper proposes SparCIM, a heterogeneous CIM accelerator optimized for decoder-only LLMs with contextual and unstructured bit sparsity, leveraging both Analog-Domain and Digital-Domain CIM architectures. The main contributions are: (1) a Heterogeneous CIM Architecture (HCA) which utilizes Analog CIM for sparse prediction of attention heads and Feed Forward Network parameters, while Digital CIM processes essential computations based on contextual sparsity, enhancing energy efficiency; (2) a Dynamic Workload Allocator (DWA) which configures DCIM memory and computing modes dynamically to reduce external memory access and token generation latency; (3) a Bit-level Butterfly Sparsity Alignment Unit (BSAU) which leverages bit-level sparsity, improving CIM utilization and throughput. The proposed accelerator is implemented with an industry 28-nm technology, achieving 1.27-5.98× higher energy efficiency compared to state-of-the-art transformer accelerators with minimal accuracy loss. Xingyu Xu 0008, Zihan Zou, Xin Si, Bo Liu 0019 |
ICCAD | 5 |
| 2025 | S-DMA: Sparse Diffusion Models Acceleration via Spatiality-Aware Prediction and Dimension-Adaptive DataflowabstractDiffusion Models (DMs) have demonstrated remarkable performance in a variety of image generation tasks.However, their complex architectures and intensive computations result in significant overhead and latency, posing challenges for hardware deployment.To address these issues, researchers have explored the sparsity in DMs to reduce computational workloads, including semantic sparsity in image generation and spatial sparsity in local editing.Unfortunately, existing sparsity prediction methods face critical limitations in deployment: 1) additional prediction overheads offset the benefits of sparsity; 2) convolution and general matrix multiplication (GEMM) exhibit distinct sparsity patterns, which current co-design frameworks struggle to process.In this paper, we introduce S-DMA, a software-hardware co-design framework that unifies efficient sparsity prediction while supporting various sparse operators.First, we propose a spatiality-aware similarity computation method that leverages the local similarity of images, reducing the computational complexity of sparsity prediction from O(𝑁 2 ) to O(N ).Second, we implement NAND-based similarity for sparsity prediction, which minimizes the computational overheads and ensures adaptability to different sparsity schemes.Finally, a dedicated hardware architecture is designed to efficiently leverage the algorithm optimizations.A NAND-based sparsity prediction processing unit is designed to adaptively handle the sparsity patterns.Additionally, a sparsity-aware reduction network and a dimension-adaptive Bo Liu is the corresponding author. Zihan Zou, Xinming Yan, Guang Yang 0036, Hao Cai 0001, Bo Liu 0019 |
MICRO | 1 |
| 2025 | PF²A-ViT: Parameter-Free and Feature-Aware Dynamic Token Pruning Accelerator With Complementary Quantization-Encoding for Vision TransformerabstractVision Transformers (ViTs) have achieved outstanding performance in visual applications. However, ViT’s increasing parameters and computation overhead limit its deployment on hardware. Previous ViT accelerators focus on optimizing the core attention mechanism of ViTs due to the high overheads of language transformer-based neural networks. Nevertheless, linear layers are the actual bottleneck of ViT inference, accounting for larger than 90% FLOPs on numerous ViTs because of the short and fixed token length of ViTs. To this end, we propose PF2A-ViT, an algorithm and accelerator co-design framework, comprehensively accelerating both core attention and linear layers. At the algorithm level, a parameter-free and feature-aware dynamic token pruning (PF2ATP) strategy is proposed to reduce the dimension of feature maps and dynamically adjust the pruning ratio according to the complexity of the feature without complex execution of subnetworks. Meanwhile, a mixed-precision quantization strategy combines Hessian trace and parameter-aware signal-to-quantization-noise ratio to boost the deployment efficiency of PF2ATP. In addition, a cluster-regroup-based weight encoding strategy is proposed to compensate for the bit-wise redundant information of the quantization strategy. At the hardware level, a token pruning module based on bitonic sorters is designed to fully leverage PF2ATP. Simultaneously, a 3D-PE array with reconfigurable 4/8-bit processing elements (PE) is designed to implement the mixed-precision quantization strategy and equipped with greedy bit-wise compensation decoders to exploit the encoding strategy. Extensive experiments on multiple ViTs demonstrate the achievements of PF2A-ViT: (1) Maximally realize 3.89× speedup, 5.54× energy efficiency compared to state-of-the-art ViT accelerators. (2) Occupying a 2.25 mm area and consuming 76 mW power in 28-nm technology. Zihan Zou, Xinming Yan, Chen Zhang 0025, Shikuang Chen, Guang Yang 0036, Han Yan 0014, Hao Cai 0001, Bo Liu 0019 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | FDCA: Fine-grained Digital-CIM based CNN Accelerator with Hybrid Quantization and Weight-Stationary DataflowabstractDigital-Compute-in-memory (DCIM) has demonstrated significant energy and area efficiency in convolutional neural network (CNN) accelerators, particularly for high precision applications. However, to mitigate parasitic effects on word and bit lines, most DCIMs employ fine-grained multiply-accumulate operations, which introduces new challenges and opportunities but has not been widely explored. This paper proposes FDCA: a fine-grained digital-CIM based CNN accelerator with hybrid quantization and weight-stationary dataflow, in which the key contributions are :1) a hybrid quantization approach for CNNs leveraging hessian trace and approximation is utilized. This method incorporates the ratio of computation time and storage time into quantization, achieving high efficiency while maintaining accuracy; 2) a Cartesian Genetic Programming based approximate shift and accumulate with error compensation is proposed, where an approximate adder tree is generated to compensate for errors introduced by DCIM; 3) an optimized weight-stationary dataflow is used to improve the utilization of CIM and eliminate dataflow stalls. The experimental results demonstrate that under 28-nm process, when running VGG16 and ResNet50 on CIFAR100, the proposed FDCA achieves 17.1TOPS/W and 18.79TOPS/W with only a slight decrease in accuracy by 0.71% and 0.98%, respectively. Compared to previous works, this work achieves 1.76× and 1.28× better in energy efficiency with less accuracy loss. Bo Liu 0019, Qingwen Wei, Yang Zhang 0132, Xingyu Xu 0008, Zihan Zou, Xinxiang Huang, Xin Si, Hao Cai 0001 |
DAC | 5 |
| 2024 | Live Demonstration: A Target-Separable BWN Inspired Speech Recognition Processor with Low-power Precision-adaptive Approximate ComputingabstractIn this live demonstration, a speech recognition system supporting both keyword spotting (KWS) and speaker verification (SV) is presented. The live demonstration is composed of three parts: the speech recognition system, a display screen, and a PC. The system can be spotted by the speaker (pre-trained in the PC), and the recognition results of the KWS and the SV can be shown on the display screen under different background noises. Chenjie Xia, Xuanhao Zhang, Zihan Zou, Hao Cai 0001, Bo Liu 0019 |
ISCAS | 3 |
| 2024 | Layer-Sensitive Neural Processing Architecture for Error-Tolerant ApplicationsabstractNeural network (NN) operation has high requirements for storage resources and parallel computing, which bring huge challenges to the deployment of NNs in Internet-of-Things (IoT) devices. Consequently, this work proposed a low-power NN architecture, comprising an energy-efficient NN processor and a Cortex-M3 host processor to achieve state-of-the-art (SOTA) end-to-end inference at the edge. The innovations of this article are as follows: 1) to minimize the bit width of the weight while keeping the loss of accuracy within a small range, cross-layer error tolerance has been analyzed, and mixed precision quantization has been adopted for cross-layer mapping; 2) dynamic reconfigurable tensor processing unit (DR-TPU) with approximate computing has been proposed, which brings$1.45\times $computing energy reduction within 0.46% accurate loss in ResNet-50; and 3) a customized input feature map (IFM) reuse and over-writeback strategy has been adopted, eliminating the recurrent fetching from the on-chip and off-chip memories. The times of on-chip storage access can be reduced by 25%–60%, and the capacity of on-chip memory can be reduced to half of the original. The processor has been implemented at 28-nm CMOS technology. Combining the above work, the proposed architecture can achieve a 53.1% reduction of power and 17.2-TOPS/W energy efficiency. Zeju Li, Qinfan Wang, Zihan Zou, Qiao Shen 0001, Na Xie, Hao Cai 0001, Hao Zhang 0111, Bo Liu 0019 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |