EDBT 2026 Demo / reviewers in the wild / expert
Shien Zhu
dblp:265/5738
· DBLP profile ↗
12ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-2094-7643ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 6 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSR: Sparse Segment Reduction for Ternary GEMM AccelerationabstractLarge Language Models (LLMs) require substantial computational resources, limiting their deployment on resource-constrained hardware. Ternary LLMs mitigate these demands through weight quantization via ternary values, achieving significant compression often with 50 − 90% sparsity. However, existing approaches have limitations: methods optimized for ternary weights, such as BitNet, redundant segment reduction (RSR), and its improved version RSR++, do not exploit sparsity structures, while conventional sparse formats neglect ternary characteristics, foregoing dual optimization opportunities.In this paper, we introduce Sparse Segment Reduction (SSR), a ternary matrix multiplication method designed to accelerate the inference of ternary LLMs and general Ternary Weight Networks (TWNs). SSR has a dedicated optimized ternary data format and an algorithm that systematically exploits sparsity patterns through computation trees that scale with the sparsity. SSR provides theoretical gains with asymptotically faster inference than RSR++ for sparsity above 50%, while practical evaluations reveal performance improvements across all sparsity levels. Evaluation results show that SSR achieves 2.1-11.3× speedup over RSR++ on ternary GEMM with 45-95% sparsity. Furthermore, SSR achieves 3.5-6.3× end-to-end speedup and 4.9% of memory saving over RSR++ on the Llama-3 1B model inference. Adeline Pittet, Shien Zhu, Valérie Verdan, Gustavo Alonso |
DATE | 2 |
| 2026 | Efficient Addition-Based Sparse GEMM for Fast Ternary Large Language Model Inference on Edge DevicesabstractLarge Language Models (LLMs) are the new dominant application but suffer from high memory and computational cost. Ternary LLMs have been proposed for easier deployment on edge platforms as 2-bit ternary weights {-1, 0, +1} can reduce the model size by 16× compared to 32-bit Floating-Point (FP32) representations. In addition, ternary General Matrix Multiplication (GEMM) can reduce the computational complexity by performing addition and subtraction operations with non-zero weights only. However, existing Central Processing Unit (CPU) and Graphic Processing Unit (GPU) do not support native 2-bit operations, and existing libraries like PyTorch and Compute Unified Device Architecture (CUDA) do not have dedicated computing kernels for ternary weights. Moreover, existing sparse formats like Compressed Sparse Column are not optimized for ternary values, causing extra storage and decompression overhead. In this article, we accelerate ternary LLMs on edge devices through efficient data formats and specialized computing kernels. We propose an efficient ternary sparse data format storing only the indices of non-zero values and simplifying the decompression at runtime. We also design a novel ternary GEMM algorithm that performs sparse addition on activations instead of multiplication to reduce the computation complexity. It achieves a 4× theoretical speedup over dense GEMM with 50% sparsity in weights. We have implemented these algorithms and optimized computing kernels on both x86 CPU and Nvidia GPUs. Evaluation results show that they achieve 1.3-3.9× speedup over Eigen Sparse GEMM, 3.3-6.9× speedup over PyTorch Sparse GEMM, and around 5.5× speedup over cuSPARSE. The GPU implementation can serve Llama-3 3B and 8B models on an RTX-3080Ti with 22 and 7 tokens/s, while the full-precision versions run out of memory. Shien Zhu, Guanshujie Fu, Mila Kjoseva, Gustavo Alonso |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2026 | TABv2: A Faster Ternary and Binary Neural Network Inference Library on the EdgeabstractModern deep neural networks and large language models (LLMs) are trained and inferred using low-bitwidth representations like FP8 and FP4. Ternary and binary quantization can further reduce the computation cost by extreme 1/2-bit data representation and bitwise operations. As ternary and binary neural networks (TNNs, BNNs, and the mixed-precision ternary-activation binary-weight and binary-activation ternary-weight networks (TBNs and BTNs)) achieve different trade-off points between the speed, model size, and accuracy, they are suitable for certain applications on the edge. However, existing works mainly focus on accelerating BNN and DoReFa-Net-style bit-serial operations, leaving TNNs and mixed-precision ones under-optimized. Although some related works such as ternary and binary (TAB) provide reference implementation for ternary and binary networks, they encounter high data access overhead during quantization and suffer from slow scalar popcount on advanced vector extension 2 (AVX2) central processing unit (CPU) which have no single-instruction multiple-data (SIMD) popcount instructions. In this article, we propose TABv2 to achieve faster inference for TNNs, BNNs, TBNs, and BTNs on CPU and graphic processing unit (GPU). First, we optimize the quantization and image-to-row (img2row) by operator fusion and better address calculation to improve the data locality. We also propose warp-cooperative quantization on GPU by utilizing data sync intrinsics. Second, we replace the scalar popcount on AVX2 with equivalent SIMD operations and reduce the complexity of the SIMD popcount utilizing ternary encoding, which can reduce approximately 15% total instruction in ternary bitwise general matrix-matrix multiplication (GEMM). Third, we propose new bitwise GEMM algorithms for TNNs and TBNs by utilizing GPU-native bitwise matrix multiplication intrinsics on tensor cores for higher efficiency. We further apply a four-degree pipeline for bitwise GEMM on GPU to hide the memory access latency. Finally, we implement these methods in C++ and combine them into an open-source library. Evaluation results show that we achieve layer-level speedup of up to 2.7× on AVX2 CPU, 2.3× on ARM CPU, and 8.7× on Nvidia GPU over the TAB baseline for TNNs, TBNs, BTNs, and BNNs. Moreover, we achieve 1.3× - 1.9× end-to-end speedup and 1.2× - 1.8× energy efficiency on CPUs and 1.3× - 3.5× end-to-end speedup and 1.2× - 3.0× energy efficiency on GPU compared to TAB on ResNet, Darknet, and visual geometry group (VGG) models. Guanshujie Fu, Olivier Fischer, Shien Zhu, Gustavo Alonso |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | Faster Ternary and Binary Neural Network Inference on CPU by Reducing Popcount OverheadabstractQuantization is a widely adopted method of reducing resource consumption of neural network models while maintaining good model accuracy. Ternary and Binary Neural Networks (TNNs and BNNs) can be implemented by lightweight bitwise operations and are thus very suitable for edge platforms. Existing efforts mainly optimize the bitwise computation algorithms for BNN inference. However, TNNs and mixed-precision Ternary-Binary Neural Networks (TBNs and BTNs) still lack optimized computing libraries on AVX2 and ARM CPUs. Their data preparation walks through the data multiple times, resulting in low data locality. Moreover, the popcount accounts for up to 28% of total operations in bitwise matrix multiplication, but the popcount has throughput of only 1 and no SIMD instructions in AVX2, becoming the central performance bottleneck.In this paper, we propose a faster inference method for TNNs, TBNs, and BTNs on AVX2 and ARM CPUs. First, we optimize the data preparation stage by fusing the quantization, bit-packing, and image-to-row into one loop to improve the data locality. Second, we propose an efficient bitwise matrix multiplication algorithm for AVX2 by replacing the low-throughput popcount instructions with high-throughput SIMD instructions and applying new data encoding. This algorithm reduces the total instruction count by 15% and brings 2.2× theoretical speedup. Third, we implement a fast C++ inferen ce library for TNNs, TBNs, and BTNs with standard optimizations like blocking and loop unrolling. Benchmarking results show that our new matrix multiplication algorithm is up to 2.1× faster than the related work TAB on AVX2 CPUs. We further achieve layer-level speedup of up to 2.7× on AVX2 and 2.3× on ARM over the baseline for TNNs, TBNs, and BTNs. Moreover, we achieve 1.3-1.9× end-to-end speedup and 1.2-1.8× energy efficiency compared to TAB on Resnet, Darknet, and VGG models. Olivier Fischer, Shien Zhu, Gustavo Alonso |
ISLPED | 2 |
| 2025 | Exploring Large Language Models for Hierarchical Hardware Circuit and Testbench GenerationabstractDesigning and verifying hardware circuits using a Hardware Description Language (HDL) is an essential but time-consuming part of hardware design. Generating the desired correct circuit and testbench code usually requires a significant engineering effort. Recently, Large Language Models (LLMs) have claimed to have strong code generation capabilities to reduce such engineering costs. Existing work has provided quantitative evaluations using LLMs for single-module, simple circuit generation. However, it is still unclear whether modern LLMs are useful in production workflows, e.g., generating correct hierarchical circuits with testbenches. And if they are capable, what are the best prompt engineering practices for hardware design? In this article, we evaluate LLMs for HDL generation by exploring a 3-dimensional design space: commercial and open-source language models, single-module and hierarchical circuits, and prompting methods with varying complexity. We propose a 3-step design space exploration methodology to answer the two aforementioned questions. First, we explore the best prompt engineering practices across generating simple, middle, and hard single-module circuits with testbenches on CodeLLama-34B. We also define two fine-grained checklists to evaluate the circuit and testbench quality from a user’s perspective. Second, we benchmark 11 LLMs with prompt adaptation on 4 single-module circuits that CodeLLama-34B has trouble with to further find models that may be useful in a production workflow. Third, we apply the learned prompt practices on four top-level models to generate simple 2 to 4-module and more complex multi-module hierarchical circuits and testbenches. As a result, we find that some of the latest LLMs can generate correct simple hierarchical circuits and testbenches with given proper prompts, but still struggle with complex hierarchical circuits. We further provide useful guidelines from an end-user’s perspective on leveraging LLMs for hardware design. Samuel Gomes Lopes, Shien Zhu, Gustavo Alonso |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2023 | An Efficient Sparse LSTM Accelerator on Embedded FPGAs with Bandwidth-Oriented PruningabstractLong short-term memory (LSTM) networks have been widely used in natural language processing applications. Although over 80% weights can be pruned to reduce the memory requirement with little accuracy loss, the pruned model still cannot be buffered on-chip for small embedded FPGAs. Considering that weights are stored in the off-chip DDR, the performance of LSTM is bounded by the available memory bandwidth. However, current pruning strategies did not consider bandwidth utilization and thus lead to bad performance in this situation. In this work, we propose an efficient sparse LSTM accelerator on embedded FPGAs with bandwidth-oriented pruning. The key idea is that data sequences can be compressed if items can be represented by a linear function of their indices in the sequences. Inspired by this idea, we first propose a column-wise pruning strategy that removes all the column indices and around 75% row indices of the remaining weights. Based on the strategy, we design a dedicated compressed format to fill the bandwidth. Further, we propose a fully pipelined hardware accelerator, which achieves the workload balance and shortens the critical path. Finally, we train the LSTM model using the TIMIT dataset and implement the accelerator on the Xilinx PYNQ-Z1 platform. The experimental result shows that our design achieves around 0.3% accuracy improvement, a 2.18x performance speedup, and a 1.96x power efficiency compared to the state-of-the-art work. Shiqing Li, Shien Zhu, Tao Luo 0014, Weichen Liu 0001 |
FPL | 2 |
| 2023 | iMAT: Energy-Efficient In-Memory Acceleration for Ternary Neural Networks With Sparse Dot ProductabstractTernary Neural Networks (TNNs) achieve an excellent trade-off between model size, speed, and accuracy, quantizing weights and activations into ternary values {+1, 0, -1}. The ternary multiplication operations in TNNs equal light-weight bitwise operations, favorably in In-Memory Computing (IMC) platforms. Therefore, many IMC-based TNN accelerators have been proposed. They build dedicated ternary multiplication cells or utilize efficient bitwise operations on IMC architectures. However, existing ternary value accumulation schemes on IMC architectures are inefficient. They extend the sign bit of integer operands or conduct two-round accumulation with specially designed encoding, bringing long latency and extra memory write overhead. Moreover, existing IMC-based TNN accelerators overlook TNNs' sparsity and conduct operations on zero weights, resulting in unnecessary power consumption and latency. In this paper, we propose iMAT to accelerate TNNs with operator-, architecture- and layer-level optimizations. First, we propose a single-round Ternary Variable-Bitwidth Accumulation scheme, which efficiently extends the addition result sign bit without extra memory write overhead. Second, we propose an in-memory accelerator with enhanced sensing circuits for the accumulation scheme and a Sparse Dot Product Unit to exploit TNNs' weight sparsity, utilizing zero weights to skip unnecessary operations. Further, we propose Fused Scaling Functions which combine the scaling, activation, normalization, and quantization layers to reduce the hardware complexity without affecting the model accuracy. Simulation results show that compared with dense in-memory TNN accelerators, our iMAT achieves up to 2.7× speedup and 3.7×energy efficiency on ternary ResNet-18. Shien Zhu, Shuo Huai, Guochu Xiong, Weichen Liu 0001 |
ISLPED | 1 |
| 2023 | FAT: An In-Memory Accelerator With Fast Addition for Ternary Weight Neural NetworksabstractConvolutional neural networks (CNNs) demonstrate excellent performance in various applications but have high computational complexity. Quantization is applied to reduce the latency and storage cost of CNNs. Among the quantization methods, binary and ternary weight networks (BWNs and TWNs) have a unique advantage over 8 and 4-bit quantization. They replace the multiplication operations in CNNs with additions, which are favored on in-memory-computing (IMC) devices. IMC acceleration for BWNs has been widely studied. However, though TWNs have higher accuracy and better sparsity than BWNs, IMC acceleration for TWNs has limited research. TWNs on the existing IMC devices are inefficient because the sparsity is not well utilized, and the addition operation is not efficient. In this article, we propose FAT as a novel IMC accelerator for TWNs. First, we propose a sparse addition control unit, which utilizes the sparsity of TWNs to skip the null operations on zero weights. Second, we propose a fast addition scheme based on the memory sense amplifier (SA) to avoid the time overhead of both carry propagation and writing back the carry to memory cells. Third, we further propose a combined-stationary data mapping to reduce the data movement of activations and weights and increase the parallelism across memory columns. Simulation results show that for addition operations at the SA level, FAT achieves$2.00\times $speedup,$1.22\times $power efficiency, and$1.22\times $area efficiency compared with a state-of-the-art IMC accelerator ParaPIM. FAT achieves$10.02\times $speedup and$12.19\times $energy efficiency compared with ParaPIM on networks with 80% average sparsity. Shien Zhu, Luan H. K. Duong, Hui Chen 0016, Di Liu 0002, Weichen Liu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | iMAD: An In-Memory Accelerator for AdderNet with Efficient 8-bit Addition and Subtraction OperationsabstractAdder Neural Network (AdderNet) is a new type of Convolutional Neural Networks (CNNs) that replaces the computational-intensive multiplications in convolution layers with lightweight additions and subtractions. As a result, AdderNet preserves high accuracy with adder convolution kernels and achieves high speed and power efficiency. In-Memory Computing (IMC) is known as the next-generation artificial-intelligence computing paradigm that has been widely adopted for accelerating binary and ternary CNNs. As AdderNet has much higher accuracy than binary and ternary CNNs, accelerating AdderNet using IMC can obtain both performance and accuracy benefits. However, existing IMC devices have no dedicated subtraction function, and adding subtraction logic may bring larger area, higher power, and degraded addition performance. Shien Zhu, Shiqing Li, Weichen Liu 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2022 | TAB: Unified and Optimized Ternary, Binary, and Mixed-precision Neural Network Inference on the EdgeabstractTernary Neural Networks (TNNs) and mixed-precision Ternary Binary Networks (TBNs) have demonstrated higher accuracy compared to Binary Neural Networks (BNNs) while providing fast, low-power, and memory-efficient inference. Related works have improved the accuracy of TNNs and TBNs, but overlooked their optimizations on CPU and GPU platforms. First, there is no unified encoding for the binary and ternary values in TNNs and TBNs. Second, existing works store the 2-bit quantized data sequentially in 32/64-bit integers, resulting in bit-extraction overhead. Last, adopting standard 2-bit multiplications for ternary values leads to a complex computation pipeline, and efficient mixed-precision multiplication between ternary and binary values is unavailable. In this article, we propose TAB as a unified and optimized inference method for ternary, binary, and mixed-precision neural networks. TAB includes unified value representation, efficient data storage scheme and novel bitwise dot product pipelines on CPU/GPU platforms. We adopt signed integers for consistent value representation across binary and ternary values. We introduce a bitwidth-last data format that stores the first and second bits of the ternary values separately to remove the bit extraction overhead. We design the ternary and binary bitwise dot product pipelines based on Gated-XOR using up to 40% fewer operations than State-Of-The-Art (SOTA) methods. Theoretical speedup analysis shows that our proposed TAB-TNN is 2.3× fast as the SOTA ternary method RTN, 9.8× fast as 8-bit integer quantization (INT8), and 39.4× fast as 32-bit full-precision convolution (FP32). Experiment results on CPU and GPU platforms show that our TAB-TNN has achieved up to 34.6× speedup and 16× storage size reduction compared with FP32 layers. TBN, Binary-activation Ternary-weight Network (BTN), and BNN in TAB are up to 40.7×, 56.2×, and 72.2× as fast as FP32. TAB-TNN is up to 70.1% faster and 12.8% more power-efficient than RTN on Darknet-19 while keeping the same accuracy. TAB is open source as a PyTorch Extension 1 for easy integration with existing CNN models. Shien Zhu, Luan H. K. Duong, Weichen Liu 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2021 | Parallel Multipath Transmission for Burst Traffic Optimization in Point-to-Point NoCsabstractNetwork-on-chip (NoC) is a promising solution to connect more than hundreds of processing elements (PEs). As the number of PEs increases, the high communication latency caused by the burst traffic hampers the speedup gained by computation acceleration. Although parallel multipath transmission is an effective method to reduce transmission latency, its advantages have not been fully exploited in previous works, especially for emerging point-to-point NoCs since: (1) Previous static message splitting strategy increases contentions when traffic loads are heavy, degrading NoC performance. (2) Only limited shortest paths are chosen, ignoring other possible paths without contentions. (3) The optimization of hardware that supports parallel multipath transmission is missing, resulting in additional overhead. Thus, we propose a software and hardware collaborated design to reduce latency in point-to-point NoCs through parallel multipath transmission. Specifically, we revise hardware design to support parallel multipath transmission efficiently. Moreover, we propose a reinforcement learning-based algorithm to decide when and how to split messages, and which path should be used according to traffic loads. Experiments show that our algorithm achieves a remarkable performance improvement (+12.1% to +21.0%) when compared with the state-of-the-art dual-path algorithm. Also, our hardware decreases power and area consumption by 23.2% and 10.3% over the dual-path hardware. Hui Chen 0016, Peng Chen 0027, Shien Zhu, Weichen Liu 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2020 | XOR-Net: An Efficient Computation Pipeline for Binary Neural Network Inference on Edge DevicesabstractAccelerating the inference of Convolution Neural Networks (CNNs) on edge devices is essential due to the small memory size and poor computation capability of these devices. Network quantization methods such as XNOR-Net, Bi-Real-Net, and XNOR-Net++ reduce the memory usage of CNNs by binarizing the CNNs. They also simplify the multiplication operations to bit-wise operations and obtain good speedup on edge devices. However, there are hidden redundancies in the computation pipeline of these methods, constraining the speedup of those binarized CNNs. In this paper, we propose XOR-Net as an optimized computation pipeline for binary networks both without and with scaling factors. As XNOR is realized by two instructions XOR and NOT on CPU/GPU platforms, XOR-Net avoids NOT operations by using XOR instead of XNOR, thus reduces bit-wise operations in both aforementioned kinds of binary convolution layers. For the binary convolution with scaling factors, our XOR-Net further rearranges the computation sequence of calculating and multiplying the scaling factors to reduce full-precision operations. Theoretical analysis shows that XOR-Net reduces one-third of the bit-wise operations compared with traditional binary convolution, and up to 40% of the full-precision operations compared with XNOR-Net. Experimental results show that our XOR-Net binary convolution without scaling factors achieves up to 135× speedup and consumes no more than 0.8% energy compared with parallel full-precision convolution. For the binary convolution with scaling factors, XOR-Net is up to 17% faster and 19% more energy-efficient than XNOR-Net. Shien Zhu, Luan H. K. Duong, Weichen Liu 0001 |
ICPADS | 1 |