Xi Chen 0107

dblp:16/3283-107 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
7since 2021 · last 2026
0009-0009-0119-5966ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An LSGQ-FFS Framework for Adaptive Optimization of Hybrid INT-CIM Architecture
abstract
Hybrid computing-in-memory (CIM) has recently gained significant attention due to its ability to leverage the strengths of both digital CIM (DCIM) and analog CIM (ACIM). The multibit fusion (MF) scheme enhances energy efficiency by fusing low-bit results, which typically require multiple read-out cycles, into a single-cycle read out. However, the relationship between hybrid INT-CIM circuit design and network performance based on the MF scheme has not yet been systematically explored. In addition, we investigate how different MF configurations affect the performance of various neural networks. To address this gap, we first propose a less-significant group quantization (LSGQ) model, which defines and explores the design space of hybrid INT-CIM. Second, we develop a FastFuse-Search (FFS) algorithm, which optimizes configurations for different networks to strike a better balance between model accuracy and energy efficiency. Based on the experimental results, some key considerations on hybrid CIM design are derived. FFS yields a$1.72\times $energy-efficiency boost with negligible accuracy loss. Finally, we fabricate a 28-nm hybrid INT-CIM test chip, achieving 59.74 TOPS/W and 0.96 TOPS/mm2, with performance metrics of 23.21 perplexity for GPT-2, 68.69% accuracy for ResNet18, and 80.53% accuracy for ViT.
Shaochen Li, Xi Chen 0107, Yujia Xiong, Lingyi Kong, He Wang 0028, Tianhui Jiao, Yan Yan 0030, Xin Si
IEEE Trans. Very Large Scale Integr. Syst.2
2025 TRIFP-DCIM: A Toggle-Rate-Immune Floating-point Digital Compute-in-Memory Design with Adaptive-Asymmetric Compute-Tree
abstract
Floating-point compute-in-memory (FP-CIM) is regarded as an attractive approach to enhancing the energy efficiency of complex neural networks. Digital domain compute mechanism has been widely utilized in CIM designs owing to its high robustness to PVT variations. However, the energy consumption of digital CIM is significantly influenced by the toggle rate of the compute-tree. This work proposes a toggle-rate-immune floating-point digital CIM (TRIFP-DCIM) design with 34.03% compute energy reduction on average. Combined with the TRIFP-DCIM design, a toggle-rate gathering method is employed in the neural network training/inference process with almost no accuracy loss. Experiment results show that the TRIFP-DCIM can achieve 14.51--36.83 TFLOPS/W @BF16 in 28nm technology process.
Tianhui Jiao, Shaochen Li, Zhican Zhang, Xi Chen 0107, Xin Si
ASP-DAC7
2025 AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM
abstract
SRAM Processing-in-Memory (PIM) has emerged as the most promising implementation for high-performance PIM, delivering superior computing density, energy efficiency, and computational precision.However, the pursuit of higher performance necessitates more complex circuit designs and increased operating frequencies, which exacerbate IR-drop issues.Severe IR-drop can significantly degrade chip performance and even threaten reliability.Conventional circuit-level IR-drop mitigation methods, such as back-end optimizations, are resource-intensive and often compromise power, performance, and area (PPA).To address these challenges, we propose AIM, comprehensive software and hardware co-design for architecture-level IR-drop mitigation in high-performance PIM.Initially, leveraging the bit-serial and in-situ dataflow processing properties of PIM, we introduce R tog and HR, which establish a direct correlation between PIM workloads and IR-drop.Building on this foundation, we propose LHR and WDS, enabling extensive exploration of architecture-level IR-drop mitigation while maintaining computational accuracy through software optimization.Subsequently, we develop IR-Booster, a dynamic adjustment mechanism that integrates software-level HR information with hardwarebased IR-drop monitoring to adapt the V-f pairs of the PIM macro, achieving enhanced energy efficiency and performance.Finally, we propose the HR-aware task mapping method, bridging software and hardware designs to achieve optimal improvement.Post-layout simulation results on a 7nm 256-TOPS PIM chip demonstrate that AIM achieves up to 69.2% IR-drop mitigation, resulting in 2.29× energy efficiency improvement and 1.152× speedup.
Yuanpeng Zhang 0002, Xing Hu 0010, Xi Chen 0107, Zhihang Yuan, Cong Li 0008, Jingchen Zhu, Xin Si, Wei Gao 0058, Qiang Wu 0012, Runsheng Wang, Guangyu Sun 0003
ISCA3
2025 Modeling of Less-Significant Group Quantization for Hybrid CIM Architecture
abstract
Hybrid computing-in-memory (CIM) has gained growing interest in recent times due to its ability to combine the strengths of both digital CIM (DCIM) and analog CIM (ACIM). The Multi-Bit Fusion (MF) scheme enhances energy efficiency by fusing low-bit results, which would typically require multiple readout cycles, into a single-cycle readout. However, the relationship between hybrid INT-CIM circuit design and network performance based on the MF scheme has yet to be systematically explored. To fill this gap, a less-significant group quantization (LSGQ) model is proposed, defining and exploring the hybrid INT-CIM design space. Experimental results demonstrate up to 1.83x improvement in energy efficiency with minimal impact on performance. A 28nm hybrid INT-CIM test chip is fabricated, achieving 52.16 TOPS/W, 0.96 TOPS/mm2, with respective performance metrics of 48.58 perplexity for GPT-2, 75.51% accuracy for ResNet18, and 81.5% accuracy for ViT.
Shaochen Li, Xi Chen 0107, Lingyi Kong, He Wang 0028, Yi Yang 0001, Xin Si
ISCAS2
2025 A 22-nm 64-kB lightning-like hybrid computing-in-memory macro with a compressed adder tree and analog-storage quantizers for transformer and CNNs
An Guo 0001, Xi Chen 0107, Fangyuan Dong, Jinwu Chen, Zhihang Yuan, Xing Hu 0010, Guangyu Sun 0003, Arindam Basu, Jun Yang 0006, Xin Si
Sci. China Inf. Sci.2
2023 From macro to microarchitecture: reviews and trends of SRAM-based compute-in-memory circuits
Zhaoyang Zhang 0008, Jinwu Chen, Xi Chen 0107, An Guo 0001, Bo Wang 0023, Tianzhu Xiong, Yuyao Kong, Xingyu Pu, Shengnan He, Xin Si, Jun Yang 0006
Sci. China Inf. Sci.3
2022 VCCIM: a voltage coupling based computing-in-memory architecture in 28 nm for edge AI applications
An Guo 0001, Xi Chen 0107, Xin Si
CCF Trans. High Perform. Comput.3
2017 FP-DNN: An Automated Framework for Mapping Deep Neural Networks onto FPGAs with RTL-HLS Hybrid Templates
abstract
DNNs (Deep Neural Networks) have demonstrated great success in numerous applications such as image classification, speech recognition, video analysis, etc. However, DNNs are much more computation-intensive and memory-intensive than previous shallow models. Thus, it is challenging to deploy DNNs in both large-scale data centers and real-time embedded systems. Considering performance, flexibility, and energy efficiency, FPGA-based accelerator for DNNs is a promising solution. Unfortunately, conventional accelerator design flows make it difficult for FPGA developers to keep up with the fast pace of innovations in DNNs. To overcome this problem, we propose FP-DNN (Field Programmable DNN), an end-to-end framework that takes TensorFlow-described DNNs as input, and automatically generates the hardware implementations on FPGA boards with RTL-HLS hybrid templates. FP-DNN performs model inference of DNNs with our high-performance computation engine and carefully-designed communication optimization strategies. We implement CNNs, LSTM-RNNs, and Residual Nets with FPDNN, and experimental results show the great performance and flexibility provided by our proposed FP-DNN framework.
Yijin Guan, Hao Liang 0003, Ningyi Xu, Shaoshuai Shi, Xi Chen 0107, Guangyu Sun 0003, Wei Zhang 0012, Jason Cong
FCCM6
2017 FxpNet: Training a deep convolutional neural network in fixed-point representation
abstract
We introduce FxpNet, a framework to train deep convolutional neural networks with low bit-width arithmetics in both forward pass and backward pass. During training FxpNet further reduces the bit-width of stored parameters (also known as primal parameters) by adaptively updating their fixed-point formats. These primal parameters are usually represented in the full resolution of floating-point values in previous binarized and quantized neural networks. In FxpNet, during forward pass fixed-point primal weights and activations are first binarized before computation, while in backward pass all gradients are represented as low resolution fixed-point values and then accumulated to corresponding fixed-point primal parameters. To have highly efficient implementations in FPGAs, ASICs and other dedicated devices, FxpNet introduces Integer Batch Normalization (IBN) and Fixed-point ADAM (FxpADAM) methods to further reduce the required floating-point operations, which will save considerable power and chip area. The evaluation on CIFAR-10 dataset indicates the effectiveness that FxpNet with 12-bit primal parameters and 12-bit gradients achieves comparable prediction accuracy with state-of-the-art binarized and quantized neural networks.
Xi Chen 0107, Hucheng Zhou, Ningyi Xu
IJCNN1