Chenjia Xie

dblp:304/3142 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0003-4982-6770ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2025 A Fast-Convergence Near-Memory-Computing Accelerator for Solving Partial Differential Equations
abstract
Solving partial differential equations (PDEs) is omnipresent in scientific research and engineering and requires expensive numerical iteration for memory and computation. The primary concerns for solving PDEs are convergence speed, data movement, and power consumption. This work proposed the first fast-convergence PDE solver with an automatic adjustment multiple-stride iteration method, significantly increasing the PDE convergence speed. A dynamic-precision near-memory-computing architecture with booth encoding is proposed to reduce iterated intermediate data movement. A customized 32T compressor and a 14T full adder are designed to reduce the power and hardware cost of the solver. The processor is fabricated using 65-nm CMOS technology and occupies a 6.25 mm2 die area. It can achieve a convergence speedup by$4\times $compared with the existing work.
Chenjia Xie, Xingyuan Hu, Yuan Du
IEEE Trans. Very Large Scale Integr. Syst.1
2024 An Efficient GCN Accelerator Based on Workload Reorganization and Feature Reduction
abstract
The irregular adjacency matrix and the mismatched computation patterns of Aggregation and Combination phases make Graph Neural Networks (GNNs) challenging to compute efficiently. This paper proposes a software and hardware co-design system to reduce computational latency and memory access based on workload reorganization and feature reduction. In software, the adjacency matrix is preprocessed, and the workload in both feature and node dimensions is concentrated to optimize memory access and hardware utilization. The interlayer nodes are analyzed using Principal Component Analysis (PCA) to explore the minimum feature vector length based on information redundancy, and a unique weight initialization is utilized for retraining to trim the feature vector to the minimum length. In hardware, an efficient GCN accelerator is designed to fully support the reorganized workload by reconfigurable output node computation. The hardware accelerator is implemented using 28-nm CMOS technology. It achieves 3.3 TOPS peak throughput and 2.6 TOPS/W energy efficiency. Compared with HyGCN, this result shows that the proposed method can improve the overall performance by$5\times $with a negligible accuracy loss of less than 0.5%.
Chenjia Xie, Zihan Ning, Liang Chang 0002, Yuan Du
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 An Energy-Efficient Spiking Neural Network Accelerator Based on Spatio-Temporal Redundancy Reduction
abstract
The neurons of spiking neural networks (SNNs) carry sparsity to the activation from temporal and spatial. To achieve high energy efficiency, this work proposes layer-wise configurable timesteps (LCTs) and channel-regrouped sparse convolution (CRSC) to explore and exploit the redundancy of temporal and spatial dimensions. In the temporal dimension, principal component analysis (PCA) is utilized to analyze redundant information in each layer, and LCT is adopted to balance and eliminate variable redundancies between layers. In the spatial dimension, channels are regrouped by spike activation frequencies to ensure workload balance, and sparse convolution is implemented to further accelerate SNN’s computation flow. A layer-fuse method that embeds the parameter of batch normalization in leaky integrate-and-fire (BLIF) is also proposed to reduce data movement and processing time. An energy-efficient SNN accelerator integrating the above methods is designed. Compared with the baseline, the hardware equipped with the proposed LCT can achieve$3.22\times $acceleration by redundancy elimination. The CRSC and BLIF can also reduce the hardware computation and memory access by$3.24\times $during inference. Implemented in TSMC 28 nm technology, this accelerator can achieve an energy efficiency of 36.89 TOPS/W at 650 MHz.
Chenjia Xie, Yuan Du
IEEE Trans. Very Large Scale Integr. Syst.1
2023 An Efficient CNN Inference Accelerator Based on Intra- and Inter-Channel Feature Map Compression
abstract
Deep convolutional neural networks (CNNs) generate intensive inter-layer data during inference, which results in substantial on- chip memory size and off-chip bandwidth. To solve the memory constraint, this paper proposes an accelerator adopting a compression technique that can reduce the inter-layer data by removing both intra- and inter-channel redundant information. Principal component analysis (PCA) is utilized in the compression process to concentrate inter-channel information. The spatial differences, truncation, and reconfigurable bit-width coding are implemented inside every feature map to eliminate the intra-channel data redundancy. Moreover, a particular data arrangement is introduced to enhance data continuity to optimize PCA analysis and improve compression performance. A CNN accelerator with the proposed compression technique is designed to support the on- the-fly compression process by pipelining the reconstruction, CNN computation, and compression operation. The prototype accelerator is implemented using 28-nm CMOS technology. It achieves 819.2GOPS peak throughput and 3.75TOPS/W energy efficiency with 218.5mW. Experiments show that the proposed compression technique achieves compression ratios of 21.5%$\sim $43.0% (8-bit mode) and 9.8%$\sim $19.3% (16-bit mode) on state-of-the-art CNNs with a negligible accuracy loss.
Chenjia Xie, Yuan Du
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 Deep Neural Network Interlayer Feature Map Compression Based on Least-Squares Fitting
abstract
Deep convolutional neural networks (CNNs) have brought a significant amount of interlayer data during computation, resulting in a large data-exchange delay and power consumption. This paper proposes a Least-Squares Fitting Compression (LSFC) method to compress the interlayer data to resolve the above problem. In LSFC, the feature maps are firstly divided into block groups; then, two base blocks are selected for each block group. Finally, the LSFC core is applied to get the fitting parameters, and the fitting parameters are selectively stored in the on-chip memory according to the mean-squared error (MSE) results. The proposed compression method is hardware-implemented and integrated into an AI accelerator to support the on-the-fly compression process with a slight hardware overhead and latency. Experiments show that the LSFC can reduce the required on-chip storage space by 21.9% $\sim$ 33.6% during CNN computation without loss of network prediction.
Chenjia Xie, Yuan Du, Zhongfeng Wang 0001
ISCAS1
2022 Memory-Efficient CNN Accelerator Based on Interlayer Feature Map Compression
abstract
Existing deep convolutional neural networks (CNNs) generate massive interlayer feature data during network inference. To maintain real-time processing in embedded systems, large on-chip memory is required to buffer the interlayer feature maps. In this paper, we propose an efficient hardware accelerator with an interlayer feature compression technique to significantly reduce the required on-chip memory size and off-chip memory access bandwidth. The accelerator compresses interlayer feature maps through transforming the stored data into frequency domain using hardware-implemented$8\times 8$discrete cosine transform (DCT). The high-frequency components are removed after the DCT through quantization. Sparse matrix compression is utilized to further compress the interlayer feature maps. The on-chip memory allocation scheme is designed to support dynamic configuration of the feature map buffer size and scratch pad size according to different network-layer requirements. The hardware accelerator combines compression, decompression, and CNN acceleration into one computing stream, achieving minimal compressing and processing delay. A prototype accelerator is implemented on an FPGA platform and also synthesized in TSMC 28-nm COMS technology. It achieves 403GOPS peak throughput and$1.4\times \sim 3.3\times $interlayer feature map reduction by adding light hardware area overhead, making it a promising hardware accelerator for intelligent IoT devices.
Yuan Du, Huadong Wei, Chenjia Xie, Zhongfeng Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.8