Yuzong Chen 0001

dblp:67/237-1 · DBLP profile ↗
← Back
15ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0001-6387-327XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 10 first-author · 12 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 $\mathrm{P}^{3}$-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
Yuzong Chen 0001, Chao Fang 0005, Xilai Dai, Thierry Tambe, Marian Verhelst, Mohamed S. Abdelfattah
ISCA1
2026 Bit-Serial Acceleration of LLM Inference With Mixture-of-Datatype Quantization
abstract
Large language models (LLMs) have achieved significant breakthroughs on machine learning tasks. Yet the substantial memory footprint of LLMs significantly hinders their wide deployment. In this paper, we propose BitMoD, an algorithm-hardware co-design solution for efficient LLM deployment. On the algorithm side, BitMoD introduces “fine-grained data type adaptation”, which uses a different data type to quantize a group (e.g., 128) of weights and key-value-cache (KV-cache). Through the careful design of these data types, BitMoD is able to quantize LLM weights and KV-cache to sub-4-bit precision while maintaining high accuracy. On the hardware side, BitMoD employs the bit-serial computing paradigm to easily support multiple numerical precisions and data types, thus providing a flexible trade-off between model accuracy and hardware efficiency. Furthermore, we design low-cost hardware components to effectively handle online KV-cache quantization and per-group partial sum dequantization. Our evaluation on a diverse set of LLMs demonstrates that BitMoD significantly outperforms state-of-the-art LLM quantization methods on both discriminative and generative tasks. Combining the superior model performance with an efficient accelerator design, BitMoD surpasses the state-of-the-art LLM accelerator in terms of both hardware performance and energy efficiency.
Yuzong Chen 0001, Chi-Chih Chang, Xilai Dai, Ahmed F. AbouElhamayed, Marta Andronic, George A. Constantinides, Mohamed S. Abdelfattah
IEEE Trans. Computers1
2025 BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
abstract
Large language models (LLMs) have demonstrated remarkable performance across various machine learning tasks. Yet the substantial memory footprint of LLMs significantly hinders their deployment. In this paper, we improve the accessibility of LLMs through BitMoD1, an algorithm-hardware co-design solution that enables efficient LLM acceleration at low weight precision. On the algorithm side, BitMoD introduces fine-grained data type adaptation that uses a different numerical data type to quantize a group of (e.g., 128) weights. Through the careful design of these new data types, BitMoD is able to quantize LLM weights to very low precision (e.g., 4 bits and 3 bits) while maintaining high accuracy. On the hardware side, BitMoD employs a bitserial processing element to easily support multiple numerical precisions and data types; our hardware design includes two key innovations: First, it employs a unified representation to process different weight data types, thus reducing the hardware cost. Second, it adopts a bit-serial dequantization unit to rescale the per-group partial sum with minimal hardware overhead. Our evaluation on six representative LLMs demonstrates that BitMoD significantly outperforms state-of-the-art LLM quantization and acceleration methods. For discriminative tasks, BitMoD can quantize LLM weights to 4 -bit with1Code is available at: https://github.com/yc2367/BitMoD-HPCA-25
Yuzong Chen 0001, Ahmed F. AbouElhamayed, Xilai Dai, Yang Wang 0053, Marta Andronic, George A. Constantinides, Mohamed S. Abdelfattah
HPCA1
2024 FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
abstract
FPGAs offer a flexible platform for accelerating deep neural network (DNN) inference, particularly for non- uniform workloads featuring fine-grained unstructured sparsity and mixed arithmetic precision. To leverage these redundancies, an emerging approach involves partially or fully unrolling computations for each DNN layer. That way, parameter-level and bit-level ineffectual operations can be completely skipped thus saving the associated area and power. Regardless, unrolled implementations scale poorly thus limiting the size of a DNN that can be unrolled on an FPGA. This motivates the investigation of new reconfigurable architectures to improve the efficiency of unrolled DNNs, while taking advantage of sparsity and mixed precision. To enable this, we present Kratos: a focused FPGA benchmark of unrolled DNN primitives with varying levels of sparsity and different arithmetic precisions. Our analysis reveals that unrolled DNNs can operate at very high frequencies, reaching the maximum frequency limit of an Arria 10 device. Additionally, we found that substantial area reductions can be achieved through fine-grained sparsity and low bit-width. We build on those results to tailor the FPGA fabric for unrolled DNNs through an architectural case study demonstrating$\sim 2\times$area reduction when using smaller LUT sizes within current FPGAs, paving the way for the exploration of new purpose-built programmable architectures.
Xilai Dai, Yuzong Chen 0001, Mohamed S. Abdelfattah
FCCM2
2024 Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
abstract
FPGAs offer a flexible platform for accelerating deep neural network (DNN) inference, particularly for non-uniform workloads featuring fine-grained unstructured sparsity and mixed arithmetic precision. To leverage these redundancies, an emerging approach involves partially or fully unrolling computations for each DNN layer. That way, parameter-level and bit-level ineffectual operations can be completely skipped, thus saving the associated area and power. Regardless, unrolled implementations scale poorly and limit the size of a DNN that can be unrolled on an FPGA. This motivates the investigation of new reconfigurable architectures to improve the efficiency of unrolled DNNs, while taking advantage of sparsity and mixed precision. To enable this, we present Kratos: a focused FPGA benchmark of unrolled DNN primitives with varying levels of sparsity and different arithmetic precisions. Our analysis reveals that unrolled DNNs can operate at very high frequencies, reaching the maximum frequency limit of an Arria 10 device. Additionally, we found that substantial area reductions can be achieved through fine-grained sparsity and low bit-width. We build on those results to tailor the FPGA fabric for unrolled DNNs through an architectural case study demonstrating $\sim 2 \times$ area reduction when using smaller LUT sizes within current FPGAs. This paves the way for further exploration of new programmable architectures that are purpose-built for sparse and low-precision unrolled DNNs. Our source code and benchmark are available on github.com/abdelfattah-lab/Kratos-benchmark.
Xilai Dai, Yuzong Chen 0001, Mohamed S. Abdelfattah
FPL2
2024 Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs
abstract
The increasing size of large language models (LLMs) traditionally requires low-precision integer formats to meet strict latency and power demands. Yet recently, alternative formats such as Normal Float (NF4) have increased model accuracy at the cost of increased chip area. In this work, we first conduct a large-scale analysis of LLM weights and activations across 30 networks and conclude that most distributions follow a Student’s t-distribution. We then derive a new theoretically optimal format, Student Float (SF4), that improves over NF4 across modern LLMs, for example increasing the average accuracy on LLaMA2-7B by 0.76% across tasks. Using this format as a high-accuracy reference, we then propose augmenting E2M1 with two variants of supernormal support for higher model accuracy. Finally, we explore the quality and efficiency frontier across 11 datatypes by evaluating their model accuracy and hardware complexity. We discover a Pareto curve composed of INT4, E2M1, and E2M1 with supernormal support, which offers a continuous tradeoff between model accuracy and chip area. For example, E2M1 with supernormal support increases the accuracy of Phi-2 by up to 2.19% with 1.22% area overhead, enabling more LLM-based applications to be run at four bits. The supporting code is hosted at https://github.com/cornell-zhang/llm-datatypes.
Jordan Dotzel, Yuzong Chen 0001, Bahaa Kotb, Sushma Prasad, Sheng Li 0007, Mohamed S. Abdelfattah, Zhiru Zhang
ICML2
2024 BBS: Bi-Directional Bit-Level Sparsity for Deep Learning Acceleration
abstract
Bit-level sparsity methods skip ineffectual zero-bit operations and are typically applicable within bit-serial deep learning accelerators. This type of sparsity at the bit-level is especially interesting because it is both orthogonal and compatible with other deep neural network (DNN) efficiency methods such as quantization and pruning. Furthermore, it comes at little or no accuracy degradation and can be performed completely post-training. However, current bit-sparsity approaches lack practicality because of (1) load imbalance from the random distribution of zero bits, (2) unoptimized external memory access because all bits are fetched from off-chip memory, and (3) high hardware implementation overhead, including large multiplexers and shifters to support sparsity at the bit level. In this work, we improve the practicality and efficiency of bit-level sparsity through a novel algorithmic bit-pruning, averaging, and compression method, and a co-designed efficient bit-serial hardware accelerator. On the algorithmic side, we introduce bi-directional bit sparsity (BBS). The key insight of BBS is that we can leverage bit sparsity in a symmetrical way to prune either zero-bits or one-bits. This significantly improves the load balance of bit-serial computing and guarantees the level of sparsity to be more than 50%. On top of BBS, we further propose two bit-level binary pruning methods that require no retraining, and can be seamlessly applied to quantized DNNs. Combining binary pruning with a new tensor encoding scheme, BBS can both skip computation and reduce the memory footprint associated with bi-directional sparse bit columns. On the hardware side, we demonstrate the potential of BBS through BitVert, a bit-serial architecture with an efficient PE design to accelerate DNNs with low overhead, exploiting our proposed binary pruning. Evaluation on seven representative DNN models shows that our approach achieves: (1) on average 1.66× reduction in model size with negligible accuracy loss of < 0.5%; (2) up to 3.03× speedup and 2.44× energy saving compared to prior DNN accelerators.
Yuzong Chen 0001, Jian Meng, Jae-sun Seo, Mohamed S. Abdelfattah
MICRO1
2023 BRAMAC: Compute-in-BRAM Architectures for Multiply-Accumulate on FPGAs
abstract
Deep neural network (DNN) inference using reduced integer precision has been shown to achieve significant improvements in memory utilization and compute throughput with little or no accuracy loss compared to full-precision floating-point. Modern FPGA-based DNN inference relies heavily on the on-chip block RAM (BRAM) for model storage and the digital signal processing (DSP) unit for implementing the multiply-accumulate (MAC) operation, a fundamental DNN primitive. In this paper, we enhance the existing BRAM to also compute MAC by proposing BRAMAC (Compute-in-BRAM Architectures for Multiply-Accumulate). BRAMAC supports 2's complement 2- to 8-bit MAC in a small dummy BRAM array using a hybrid bit-serial & bit-parallel data flow. Unlike previous compute-in-BRAM architectures, BRAMAC allows read/write access to the main BRAM array while computing in the dummy BRAM array, enabling both persistent and tiling-based DNN inference. We explore two BRAMAC variants: BRAMAC-2SA (with 2 synchronous dummy arrays) and BRAMAC-1DA (with 1 double-pumped dummy array). BRAMAC-2SA/BRAMAC-1DA can boost the peak MAC throughput of a large Arria-10 FPGA by$2.6\times/2.1\times,2.3\times/2.0\times$, and$1.9\times/1.7$× for 2-bit, 4-bit, and 8-bit precisions, respectively at the cost of 6.8%/3.4% increase in the FPGA core area. By adding BRAMAC-2SA/BRAMAC-1DA to a state-of-the-art tiling-based DNN accelerator, an average speedup of$2.05\times/1.7\times$and$1.33\times/1.52\times$can be achieved for AlexNet and ResNet-34, respectively across different model precisions. Our code is available at: https://github.com/abdelfattah-lab/BRAMAC.
Yuzong Chen 0001, Mohamed S. Abdelfattah
FCCM1
2023 BP-SCIM: A Reconfigurable 8T SRAM Macro for Bit-Parallel Searching and Computing In-Memory
abstract
This work presents BP-SCIM: a reconfigurable 8T static random access memory (SRAM) macro for bit-parallel searching and computing in-memory (CIM). The decoupled read/write ports of the employed 8T SRAM bit-cell eliminate read disturbance during search and CIM operations. BP-SCIM can support both in-memory Boolean logic and arithmetic operations. Novel CIM-friendly algorithms and peripheral circuits are proposed to reduce the latency of complex arithmetic operations such as multiplication and division. In addition, BP-SCIM can be configured as either a binary content-addressable memory (CAM) or a ternary CAM for fast searching. A$256\times64$BP-SCIM test chip was implemented in 65-nm CMOS technology. The 8-bit addition and 8-bit multiplication operations can achieve the maximum energy efficiency of 3.11 TOPS/W and 0.17 TOPS/W, respectively at 0.7 V supply. For the binary CAM search operation, BP-SCIM can achieve the minimum energy consumption of 0.91 fJ/bit/search at 87 MHz and 0.8 V supply.
Yuzong Chen 0001, Junjie Mu, Lu Lu 0013, Tony Tae-Hyoung Kim
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 A Reconfigurable 8T SRAM Macro for Bit-Parallel Searching and Computing In-Memory
abstract
This work presents BP-SCIM: a reconfigurable 8T SRAM macro for bit-parallel searching and computing in-memory (CIM). BP-SCIM can perform in-memory Boolean logic, arithmetic, and content-addressable memory (CAM) operations. Novel peripheral circuits and algorithms are proposed to support complex arithmetic operations such as multiplication and division. A $256\times 64$ BP-SCIM test chip was implemented in 65-nm CMOS technology. The 8-bit addition and 8-bit multiplication operations can achieve the best energy efficiency of 3.11 TOPS/W and 0.17 TOPS/W, respectively at 0.7 V supply. For the binary CAM search operation, BP-SCIM can achieve the minimum energy consumption of 0.91 fJ/bit/search at 87 MHz and 0.8 V supply.
Yuzong Chen 0001, Junjie Mu, Lu Lu 0013, Tony Tae-Hyoung Kim
ISCAS1
2021 A Multi-Functional 4T2R ReRAM Macro Enabling 2-Dimensional Access and Computing In-Memory
abstract
This paper presents a multi-functional resistive random access memory (ReRAM) macro using a novel 4T2R bit-cell. The proposed 4T2R ReRAM enables 2-dimensional (2D) memory access, which offers significant latency and energy reductions for many applications such as matrix operations. Besides non-volatile storage, the proposed 4T2R ReRAM can support two types of computing in-memory operations: ternary content-addressable memory (TCAM) and logic in-memory (LIM). Evaluations on various matrix operations show that the proposed 4T2R ReRAM with 2-D access capability can reduce memory access latency and energy by up to 88% and 82%, respectively compared with conventional 1T1R ReRAM. For TCAM, the proposed 4T2R bit-cell takes a smaller area than SRAM-based TCAM cell, while achieves a comparable search speed. For LIM, we propose an optimized LIM full adder (LIM- FA) that improves the delay and the power by 3.2* and 1.6*, respectively compared with prior LIM-FAs.
Yuzong Chen 0001, Lu Lu 0013, Yuncheng Lu, Tony Tae-Hyoung Kim
ISCAS1
2021 A Configurable Randomness Enhanced RRAM PUF with Biased Current Sensing Scheme
abstract
This paper explains a resistive nonvolatile memory (RRAM)-based physical unclonable function (PUF). The proposed PUF involves more variations and enhances the randomness by using four 1T1R cells to generate one-bit random data. Moreover, we propose a configurable replica column scheme and a biased current sensing amplifier to further improve the randomness. The proposed RRAM PUF is designed in 40nm CMOS technology. The simulated RRAM PUF with configurable replica columns achieves the randomness of 0.5001 with σ of 0.005. It also has a Hamming distance of 0.486 over various responses for a single chip and 0.498 between different chips.
Lu Lu 0013, Yuzong Chen 0001, Tony Tae-Hyoung Kim
ISCAS2
2021 AND8T SRAM Macro with Improved Linearity for Multi-Bit In-Memory Computing
abstract
In this work, we propose a multi-bit precision (4b input, 4b weight and 4b output) in-memory computing (IMC) architecture, based on the voltage scaling and charge sharing scheme, for the artificial intelligence (AI) edge devices. To achieve the efficient computation, a new AND logic based 8T SRAM cell (AND8T) has been used which employs the charge-domain based computation. For such computation, AND8T incorporates an overlaying metal-oxide-metal capacitor (MOM cap) with no bit-cell area overhead. The proposed cell mitigates the linearity issue of multiply and accumulate (MAC) operation for the IMC unit which is highly desirable for the reliable operation of complex neural networks (CNN). Moreover, our high precision AND8T based IMC architecture allows 128 parallel MAC operations avoiding the need of serial multi-bits input implementation through multiple cycles. The proposed design has been successfully verified by the monte carlo simulation results while working at 50MHz clock frequency and 1V supply using standard 65nm node.
Vishal Sharma 0004, Ju Eon Kim, Yuzong Chen 0001, Tony Tae-Hyoung Kim
ISCAS4
2020 Reconfigurable 2T2R ReRAM with Split Word-Lines for TCAM Operation and In-Memory Computing
abstract
The increased latency and power consumption due to data movement between memory and ALU have become the major obstacle in modern big-data and machine learning applications. Beyond von-Neumann architectures, particularly in-memory computing, is under intensive research to overcome this memory access bottleneck. In this work, we propose a 2T2R ReRAM structure that supports ternary content addressable memory (TCAM), logic in-memory operations, and in-memory dot product for Deep Neural Networks (DNNs) besides the normal non-volatile memory (NVM) functionality. This is achieved by employing reconfigurable sense amplifiers and novel word-line drivers. The proposed architecture can serve as a high-density storage system as well as an accelerator for data-intensive applications. Simulation results verify that the proposed 2T2R structure functions correctly for TCAM search, logic in-memory operations and in-memory dot product.
Yuzong Chen 0001, Lu Lu 0013, Bongjin Kim, Tony Tae-Hyoung Kim
ISCAS1
2020 Reconfigurable 2T2R ReRAM Architecture for Versatile Data Storage and Computing In-Memory
abstract
Nonvolatile memory (NVM)-based computing in-memory (CIM) is a promising solution to data-intensive applications. This work proposes a 2T2R resistive random access memory (ReRAM) architecture that supports three types of CIM operations: 1) ternary content addressable memory (TCAM); 2) logic in-memory (LiM) primitives and arithmetic blocks such as full adder (FA) and full subtractor; and 3) in-memory dot-product for neural networks. The proposed architecture allows the NVM operations in both 2T2R and conventional 1T1R configurations. The proposed LiM full adder (LiM-FA) improves the delay, the static power, and the dynamic power by$3.2\times $,$1.2\times $, and$1.6\times $, respectively, compared with state-of-the-art LiM-FAs. Furthermore, based on different optimization techniques and robustness analysis, a lower precharge voltage is set for each mode. This reduces the TCAM search energy and 1T1R ReRAM access energy by$1.6\times $and$1.14\times $, respectively, compared with the case without optimizations.
Yuzong Chen 0001, Lu Lu 0013, Bongjin Kim, Tony Tae-Hyoung Kim
IEEE Trans. Very Large Scale Integr. Syst.1