Jiaqi Yang 0009

dblp:131/7234-9 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0008-1195-8114ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 M3DKV: Monolithic 3D Gain Cell Memory Enabled Efficient KV Cache & Processing
abstract
Transformer-based generative large language models (LLMs) have revolutionized natural language processing, yet their quadratic growth in computational complexity in context length creates severe inference bottlenecks. While LLM keyvalue cache (KV cache) enhances decoding efficiency, prolonged contexts infer frequent KV cache reloads that exacerbate memory bandwidth constraints. To address this hardware challenge, we propose M3DKV-a monolithic three-dimensional (3D) gain cell near-memory computing accelerator featuring back-end-of-line (BEOL) cache layers for in-situ KV matrix buffering and computation and a front-end-of-line (FEOL) base layer for full selfattention operations. Through optimized 3D data organization, inter-layer dataflow management, and intelligent computation scheduling, our design achieves $0.29 \mathrm{~TB} / \mathrm{s} /$ core on-die bandwidth while demonstrating $97.03 \times / 268.01 \times$ speedup over GPU/CPU in the decoding stage and $1.72 \times-262.16 \times$ better area efficiency per parameter versus state-of-the-art accelerators.
Jiaqi Yang 0009, Yanbo Su, Yihan Fu, Jianshi Tang, Bonan Yan
ASP-DAC1
2026 ESTroM: Element-Flow Architecture for Processing Sparse Tractable Probabilistic Models
abstract
Probabilistic Circuits (PCs) models are emerging popular tractable probabilistic models. Their internal connections are represented in the form of directed acyclic graphs (DAGs) with sum nodes and product nodes, ensuring their internal parameter efficiency and model expressiveness in terms of probabilistic inference. Despite these algorithmic advantages, executing PC still faces graph structure deployment issues. PyJuice on GPU with the block-sparse parallel computation methods causes a parallelism-sparsity gap, while DAG-style processing does not take advantage of the repetitive characteristics of PC internal nodes, resulting in low throughput. To address this challenge, this work proposes the ESTroM, an efficient architecture that provides novel graph-element (nodes/edges) parallelism with sparsity-aware compilation. Through analysis of the sum/product node computational requirements, ESTroM core uses compressed matrices for sum/product nodes DAG representations, edge-based dataflow for product node processing, and node-based dataflow for sum node processing. With intra-core rewind and intercore multicast optimizations, we develop a prototype ESTrom chip and a demonstrative system for a PC-based neural lossless compression application. Our ablation experiments show ESTrom offers a speed improvement of$2.11 \sim 3.79 \times$compared to the state-of-the-art DAG processing unit (DPU)-v2 with the same computing resources. Under various typical PC structures, ESTrom achieves a speedup of$18.7 \times$compared to DPU-v2 and$3.9 \times$compared to NVIDIA RTX 4090 GPU with PyJuice framework. In terms of neural lossless compression, ESTroM demonstrates a$1.39 \times$improvement in compression ratio compared to the industrial-standard Z-standard (Zstd) algorithms with the highest compression level, while offering$16.3 \sim 65.2 \times$improvement in compression speed compared to Zstd on Intel Xeon Gold 6230. In a nutshell, this work develops novel graph element parallelism and element-flow architecture theory with practical prototype chips and systems, revealing a new hardware-perspective path for the “scaling law” of emerging tractable probabilistic models.
Anjunyi Fan, Xuejie Liu, Anji Liu, Qiuping Wu, Jiaqi Yang 0009, Yuchao Qin, Guy Van den Broeck, Yitao Liang, Bonan Yan
HPCA5
2026 A 442.42 TOPS/W RRAM-based Digital Computing-in-Memory Accelerator for BF16×1-bit Vision Transformer
Jingyun Gu, Jiaqi Yang 0009, Jingyao Dong, Qilong Chen, Dingbang Liu, Wei Mao 0002, Hao Yu 0001
ISCAS3
2025 A Layer-wised Mixed-Precision CIM Accelerator with Bit-level Sparsity-aware ADCs for NAS-Optimized CNNs
abstract
Exploring multiple precisions as well as sparsities for a computingin-memory (CIM) based convolutional accelerators is challenging. To further improve energy efficiency with minimal accuracy loss, this paper develops a neural architecture search (NAS) method to identify precision for each layer of the CNN and further leverages bit-level sparsity. The results indicate that following this approach, ResNet-18 and VGG-16 not only maintain their accuracy but also implement layer-wised mixed-precision effectively. Furthermore, there is a substantial enhancement in the bit-level sparsity of weights within each layer, with an average bit-level sparsity exceeding 90% per bit, thus providing broader possibilities for hardware-level sparsity optimization. In terms of hardware design, a mixed-precision (2/4/8-bit) readout circuit as well as a bit-level sparsity-aware Analog-to-Digital Converter (ADC) are both proposed to reduce system power consumption. Based on bit-level sparsity mixed-precision CNNs benchmarks, post-layout simulation results in 28nm reveal that the proposed accelerator achieves up to 245.72 TOPS/W energy efficiency, which shows about 2.52 -- 6.57× improvement compared to the state-of-the-art SRAM-based CIM accelerators.
Haoxiang Zhou, Zikun Wei, Dingbang Liu, Liuyang Zhang, Chenchen Ding, Jiaqi Yang 0009, Wei Mao 0002, Hao Yu 0001
ASP-DAC6
2025 A 20.98TOPS/W Energy-Efficient Binary BERT Model on Group Vector Systolic CIM Accelerator
abstract
Transformer-based large language models (LLMs) impose significant bandwidth and compute challenges when deployed on edge devices. SRAM-based compute-in-memory (CIM) accelerators offer a promising solution to reduce data movement but are still limited by model size. This work develops a ternary weight splitting (TWS) binarization to obtain Brain-Floating-Point-16×INT1 (BF16×1-b) and INT8×INT1 (8-b×1-b) based transformers that exhibit competitive accuracy while significantly reducing model size compared to full precision counterparts. Then, a fully digital SRAM-based CIM accelerator is designed incorporating a bit-parallel SRAM macro within a highly efficient group vector systolic architecture, which can store one column of BERT-Tiny model with stationary systolic data reuse. The design in a 28nm technology only requires 2KB SRAM with an area of 2mm2. It achieves a throughput of 6.55TOPS and consumes a total power of 312.5mW and 221mW at 400MHz, resulting in a state-of-the-art area efficiency of 3.3TOPS/mm2and normalized energy efficiency of 20.98TOPS/W and 34.35TOPS/W for BF16×1-b and 8-b×1-b respectively on BERT-Tiny model, demonstrating a 10.25× improvement in area efficiency and a 2.23× improvement in energy efficiency compared to other state-of-the-art counterparts. Additionally, our proposed configuration compresses the model size by 32% with only a 0.5% accuracy loss on SST-2.
Dingbang Liu, Qilong Chen, Jingyun Gu, Jiaqi Yang 0009, Kai Li 0024, Wei Mao 0002, Ngai Wong 0001, Chang Wen Chen, Hao Yu 0001
ISLPED5
2025 A 28-nm 135.19 TOPS/W Bootstrapped-SRAM Compute-in-Memory Accelerator With Layer-Wise Precision and Sparsity
abstract
Artificial intelligence (AI) edge devices demand high energy efficiency as well as inference accuracy. SRAM-based compute-in-memory (CIM) accelerators have great potential for power reduction but still need to exploit higher throughput and better linearity performance. To meet edge-AI computing demands by CIM works, it is crucial to optimize algorithms and parameters for specific circuit systems to achieve hardware acceleration. This work firstly employs neural network search (NAS) method to find out the layer-wise optimized precisions and sparsities for convolutional neural networks (CNNs). Then, a 144-Kb charge-domain signed mixed-precision (2/4/8-bit) CIM accelerator employing bootstrapped SRAM cells with 9-transistors and 1-capacitor (9T1C) structure is proposed that incorporates a bit-level sparsity-aware analog-to-digital converter (ADC). This work not only achieves highly linear parallel accumulation operations to meet AI computing demands but also implements a hardware and software co-optimization system tailored to specific data characteristics. The design is verified on NAS-optimized networks VGG-16 and ResNet-18 using Cifar-10 dataset, which could achieve an equivalent accuracy at 4-bit of 68.68% while maintaining a high energy efficiency at 2-bit of 135.19TOPS/W by measurements.
Wei Mao 0002, Dingbang Liu, Haoxiang Zhou, Fuyi Li, Kai Li 0024, Qiuping Wu, Jiaqi Yang 0009, Liuyang Zhang, Hao Yu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.7