Jia Chen 0032

dblp:99/6879-32 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0003-4814-5420ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 first-author · 10 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
YearPublicationVenuePosition
2026 MemSearch: An Efficient Memristive In-memory Search Engine with Configurable Similarity Measures
Yingjie Yu, Houji Zhou, Jiancong Li, Tong Hu, Jia Chen 0032, Yi Li 0049, Xiangshui Miao
ASP-DAC5
2026 OAH-CIM: Outlier-Aware Hybrid RRAM-SRAM CIM Accelerator with Variation-Robust Sparsity
Tong Hu, Han Bao 0013, Houji Zhou, Yuyang Fu, Jiancong Li, Jia Chen 0032, Yi Li 0049, Xiangshui Miao
ASP-DAC7
2026 Bridging Efficiency and Scalability in Llm System Via 3D Hybrid Pim With 2D in-Transit Computation
abstract
Large Language Models (LLMs) have transformed society, but their computational and energy needs hinder efficient inference. The memory wall, the growing processor-memory speed disparity, remains a critical bottleneck for LLM. While Process-in-Memory (PIM) architectures address this challenge by co-locating computation with memory, achieving 5-20 × higher bandwidth than GPUs, existing scalable PIM solutions face critical trade-offs in flexibility, capacity, and efficiency when handling LLMs' dynamic memory-compute patterns and operator diversity. DRAM-PIM suffers from inter-bank communication overhead despite its vector parallelism. SRAM-PIM offers sub10ns latency for matrix operation but is constrained by limited capacity. This work introduces CompAir, a scalable PIM architecture that integrates DRAM-PIM and SRAM-PIM with hybrid bonding, enabling efficient linear computations while unlocking multi-granularity data pathways. We further develop CompAirNoC, an advanced network-on-chip (NoC) with an embedded arithmetic logic unit that performs non-linear operations during data movement. Such a design offloads the centralized communication bottleneck in the channel level to distributed banks, simultaneously reducing communication overhead and area cost for scalability. Finally, we develop a hierarchical Instruction Set Architecture that ensures both flexibility and programmability of the hybrid PIM. Experiments show CompAir delivers 1.83-7.98× faster prefill and 1.95 - 6.28 × faster decoding versus state-ofthe-art PIM designs, with 3.52 × lower energy than GPU-PIM hybrids. This work presents the first systematic exploration of hybrid DRAM-PIM and SRAM-PIM architectures with innetwork computation, paving the way towards a scalable PIM system for LLM inference.
Songchen Ma, Huanyu Qu, Jia Chen 0032, Junfeng Lin, Fengbin Tu
ISCA5
2026 A Scaling Annealing Method for Combinatorial Optimization in Asynchronous Memristive Hopfield Network
Han Bao 0013, Kehong Xu, Yibai Xue, Yuyang Fu, Jiancong Li, Jia Chen 0032, Yi Li 0049, Xiangshui Miao
ISCAS7
2025 ReSMiPS: A ReRAM-based Sparse Mixed-precision Solver with Fast Matrix Reordering Algorithm
abstract
The solution of sparse matrix equations is essential in scientific computing. However, traditional solvers on digital computing platforms are limited by memory bottlenecks in largescale sparse matrix storage and computation. Resistive Random Access Memory (ReRAM)-based computing-in-memory (CIM) offers a promising solution to this challenge but faces constraints in achieving high solution precision and energy efficiency simultaneously in sparse matrix computations. In this work, we propose ReSMiPS, a ReRAM-accelerated sparse mixed-precision solver. ReSMiPS incorporates a novel Fast Sparse Matrix Reordering (FSMR) algorithm and introduces an In-memory float64 (IF64) data format, enabling efficient floating-point sparse matrix computation directly within the analog ReRAM array. By combining our floating-point CIM macro design with a hybrid-domain solution framework, ReSMiPS achieves precision comparable to CPU and GPU-based BiCGSTAB solvers (with errors below $10^{-15}$) on real-world workloads, while demonstrating over two orders of magnitude improvement in both computational speed and energy efficiency.
Yuyang Fu, Jiancong Li, Jia Chen 0032, Houji Zhou, Wenlong Peng, Yi Li 0049, Xiangshui Miao
DAC3
2025 SynDCIM: A Performance-Aware Digital Computing-in-Memory Compiler with Multi-Spec-Oriented Subcircuit Synthesis
abstract
Digital Computing-in-Memory (DCIM) is an innovative technology that integrates multiply-accumulation (MAC) logic directly into memory arrays to enhance the performance of modern AI computing. However, the need for customized memory cells and logic components currently necessitates significant manual effort in DCIM design. Existing tools for facilitating DCIM macro designs struggle to optimize subcircuit synthesis to meet user-defined performance criteria, thereby limiting the potential system-level acceleration that DCIM can offer. To address these challenges and enable the agile design of DCIM macros with optimal architectures, we present SynDCIM - a performance-aware DCIM compiler that employs multi-spec-oriented subcircuit synthesis. SynDCIM features an automated performance-to-layout generation process that aligns with user-defined performance expectations. This is supported by a scalable subcircuit library and a multi-spec-oriented searching algorithm for effective subcircuit synthesis. The effectiveness of SynDCIM is demonstrated through extensive experiments and validated with a test chip fabricated in a 40nm CMOS process. Testing results reveal that designs generated by SynDCIM exhibit competitive performance when compared to state-of-the-art manually designed DCIM macros.
Kunming Shao, Fengshi Tian, Jiakun Zheng, Jia Chen 0032, Jingyu He, Hui Wu 0010, Jinbo Chen 0002, Xihao Guan, Fengbin Tu, Jie Yang 0033, Mohamad Sawan, Kwang-Ting Cheng, Chi-Ying Tsui
DATE5
2025 Exploiting the Memory-Compute-Coupling Feature for CIM Accelerator Design Optimization
abstract
SRAM computing-in-memory (CIM) accelerators have evolved as a promising solution to the memory wall problem in neural network (NN) models. By integrating memory and compute resources in each macro, CIM accelerators offer massive in-situ computing parallelism and large memory capacity, enabling spatial mapping with layer fusion and potentially keeping layers stationary in CIM. However, CIM’s memory-compute coupling (MCC) feature poses challenges in designing CIM accelerators. From an architecture aspect, designers must balance CIM’s memory and compute resources by optimizing the macro’s memory-compute ratio (MCR) configuration across diverse scenarios. From a mapping aspect, conventional mappings, which allocate each macro exclusively to each layer, face two major problems: a layer-fusion dilemma (the accelerator suffers from excessive memory access due to layer replications or performance degradation due to load imbalance) and a layer-eviction issue (storing layers stationary in CIM is usually infeasible due to limited CIM capacity). To address these challenges, this paper introduces MCC-DSE, an MCC-aware Design Space Exploration framework for architecture-mapping co-optimization of CIM accelerators. We also propose a three-axis CIM division mapping, which interleaves multiple layers in each macro to concurrently optimize memory access and performance during layer fusion as well as reserves a part of CIM memory in each macro for layer pinning. Compared to baseline architecture and mapping, MCC-DSE shows a 1.4x 8.3x EDP reduction across various workloads and chip areas. Moreover, MCC-DSE provides insights into CIM accelerator optimization, such as selecting optimal MCR and configuring CIM dynamically for different scenarios.
Yongkun Wu, Jia Chen 0032, Zhenhua Zhu 0002, Jingyu He, Pingcheng Dong, Yonghao Tan, Xin Zhao 0044, Liang Chang 0002, Yu Wang 0002, Fengbin Tu, Chi-Ying Tsui, Kwang-Ting Cheng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 LSMR: Synergy Randomness in Liquid State Machine and RRAM-based Analog-digital Accelerator
abstract
Bio-inspired event sensors are gaining popularity at the edge, such as in robots and wearable electronics. This trend necessitates learning vast amounts of sensory data on the edge, often in few-shot or even zero-shot scenarios, posing challenges in both software and hardware. This paper presents a novel software-hardware co-design to address these issues. Software-wise, we develop an SNN-ANN model, where the SNN encoder is a liquid state machine (LSM) that naturally processes events and significantly reduces learning complexity at the edge due to fixed random weights. The lightweight trainable ANN projection heads are optimized through contrastive learning, enabling zero-shot learning of multimodal events. Hardware-wise, we propose a hybrid analog (RRAM)-digital (CMOS) accelerator - LSMR. The analog in-memory computing core physically implements the LSM by leveraging RRAM stochasticity to generate fixed random weights. The digital core utilizes innovative reconfigurable systolic arrays to accelerate the contrastive learning of ANN projection heads. Extensive experimental outcomes from six neuromorphic datasets, encompassing visual, tactile, and auditory modalities, demonstrate that LSMR considerably improves energy efficiency by a range of 1.65× to 23.70×, in comparison to state-of-the-art edge devices. Simultaneously, it reduces training complexity by a range of 152.83× to 20,587.77× across various edge learning tasks.
Ning Lin, Songqi Wang, Xinyuan Zhang 0008, Shaocong Wang 0001, Yangu He, Woyu Zhang, Bo Wang 0153, Jiankun Li, Mingzi Li, Binbin Cui, Yi Li 0049, Jia Chen 0032, Chunwei Xia, Xiaoming Chen 0003, Dashan Shang
ICCAD12
2023 AutoDCIM: An Automated Digital CIM Compiler
abstract
Digital Computing-in-Memory (DCIM) is an emerging architecture that integrates digital logic into memory for efficient AI computing. However, current DCIM designs heavily rely on manual efforts. This increases DCIM design time and limits the optimization space, making it challenging to satisfy the user specifications of diverse AI applications. This paper presents AutoDCIM, the first automated DCIM compiler. Au-toDCIM takes the user specifications as inputs and generates a DCIM macro architecture with an optimized layout. AutoDCIM’s template-based generation balances handcrafted cell design and agile macro development. AutoDCIM’s layout exploration loop analyzes diverse DCIM array partitioning schemes to satisfy user specifications. The auto-generated DCIM macros present competitive efficiency results in comparison with state-of-the-art silicon-verified DCIM macros.
Jia Chen 0032, Fengbin Tu, Kunming Shao, Fengshi Tian, Xiao Huo, Chi-Ying Tsui, Kwang-Ting Cheng
DAC1
2023 Architecting Efficient Multi-modal AIoT Systems
abstract
Multi-modal computing (M2C) has recently exhibited impressive accuracy improvements in numerous autonomous artificial intelligence of things (AIoT) systems. However, this accuracy gain is often tethered to an incredible increase in energy consumption. Particularly, various highly-developed modality sensors devour most of the energy budget, which would make the deployment of M2C for real-world AIoT applications a difficult challenge.
Xiaofeng Hou, Jiacheng Liu 0001, Xuehan Tang, Chao Li 0009, Jia Chen 0032, Luhong Liang, Kwang-Ting Cheng, Minyi Guo
ISCA5