Xingyu Xu 0008

dblp:361/7512 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0006-9890-3567ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 3D-TANoC: Thermal-Aware 3D LLM Accelerator with Hierarchical NoC and Operator-Aware Dataflow Mapping Strategy
Xingyu Xu 0008, Dengke Liu, Zihan Zou, Xilong Kang, Hui Kou, Hao Cai 0001, Bo Liu 0019
ISCAS2
2025 SUArch: Accelerating Layer-wise N: M Sparse Pattern with a Unified Architecture for Deep-learning Edge Device
abstract
Deep neural networks are of the essence for user applications on edge devices. However, the computation and memory-intensive nature of deep neural networks conflicts with the resource-constrained devices. Moreover, the heterogeneity across different models imposes new challenges on deployment on edge devices. To boost the capabilities of edge devices, we propose SUArch, which innovates on three fronts: 1) a layer-wise N:M sparsity aware training approach to strike a balance between accuracy and training cost; 2) a sparsity alignment unit based on the butterfly network to maximize hardware utilization and eliminate extra overhead; 3) a mode-heterogenous processing element array to effectively accomplish the unified support for Convolution Neural Network and Transformer. The experimental results demonstrate that when running convolution-based and attention-based models under an industrial 28-nm process, the proposed SUArch realizes an energy efficiency of 52.1 TOPS/W. Compared to state-of-the-art architecture, SUArch achieves an energy efficiency improvement of 2.07× while accuracy loss is within 0.7%.
Xilong Kang, Qingwen Wei, Ningyuan Li 0004, Xingyu Xu 0008, Hao Cai 0001, Bo Liu 0019
ASP-DAC4
2025 SparCIM: A Heterogeneous CIM-Based Accelerator for Large Language Models with Contextual and Unstructured Bit Sparsity
abstract
Transformer-based Large Language Models (LLMs) exhibit high computation and memory demands, especially on resource-constrained edge devices. Sparsification techniques like contextual sparsity have shown potential in reducing these demands by selectively pruning model components, but their integration with efficient hardware remains challenging. Firstly, sparsity prediction introduces extra energy and latency cost. Secondly, prefilling and decoding stages in LLM inference cause imbalanced workload. Thirdly, Computing-in-Memory (CIM) suffers from low utilization due to unstructured bit sparsity. To address the above challenges, this paper proposes SparCIM, a heterogeneous CIM accelerator optimized for decoder-only LLMs with contextual and unstructured bit sparsity, leveraging both Analog-Domain and Digital-Domain CIM architectures. The main contributions are: (1) a Heterogeneous CIM Architecture (HCA) which utilizes Analog CIM for sparse prediction of attention heads and Feed Forward Network parameters, while Digital CIM processes essential computations based on contextual sparsity, enhancing energy efficiency; (2) a Dynamic Workload Allocator (DWA) which configures DCIM memory and computing modes dynamically to reduce external memory access and token generation latency; (3) a Bit-level Butterfly Sparsity Alignment Unit (BSAU) which leverages bit-level sparsity, improving CIM utilization and throughput. The proposed accelerator is implemented with an industry 28-nm technology, achieving 1.27-5.98× higher energy efficiency compared to state-of-the-art transformer accelerators with minimal accuracy loss.
Xingyu Xu 0008, Zihan Zou, Xin Si, Bo Liu 0019
ICCAD1
2024 FDCA: Fine-grained Digital-CIM based CNN Accelerator with Hybrid Quantization and Weight-Stationary Dataflow
abstract
Digital-Compute-in-memory (DCIM) has demonstrated significant energy and area efficiency in convolutional neural network (CNN) accelerators, particularly for high precision applications. However, to mitigate parasitic effects on word and bit lines, most DCIMs employ fine-grained multiply-accumulate operations, which introduces new challenges and opportunities but has not been widely explored. This paper proposes FDCA: a fine-grained digital-CIM based CNN accelerator with hybrid quantization and weight-stationary dataflow, in which the key contributions are :1) a hybrid quantization approach for CNNs leveraging hessian trace and approximation is utilized. This method incorporates the ratio of computation time and storage time into quantization, achieving high efficiency while maintaining accuracy; 2) a Cartesian Genetic Programming based approximate shift and accumulate with error compensation is proposed, where an approximate adder tree is generated to compensate for errors introduced by DCIM; 3) an optimized weight-stationary dataflow is used to improve the utilization of CIM and eliminate dataflow stalls. The experimental results demonstrate that under 28-nm process, when running VGG16 and ResNet50 on CIFAR100, the proposed FDCA achieves 17.1TOPS/W and 18.79TOPS/W with only a slight decrease in accuracy by 0.71% and 0.98%, respectively. Compared to previous works, this work achieves 1.76× and 1.28× better in energy efficiency with less accuracy loss.
Bo Liu 0019, Qingwen Wei, Yang Zhang 0132, Xingyu Xu 0008, Zihan Zou, Xinxiang Huang, Xin Si, Hao Cai 0001
DAC4
2023 Work-in-Process: Error-Compensation-Based Energy-Efficient MAC Unit for CNNs
abstract
Approximate circuits sacrifice accuracy in exchange for energy efficiency and have been widely used in hardware deployment of neural networks (NNs). Since convolution accounts for most of the power consumption in NNs, it is necessary to design an approximate multiplication and accumulation (MAC) unit which improve the energy efficiency of hardware with ignorable accuarcy loss. In this work, an error-compensation-based energy-efficient MAC unit is proposed in which approximate multipliers are designed by Boolean matrix factorization and approximate adders are generated by Cartesian genetic programming. The proposed MAC unit is conducted on CIFAR10 using ResNet-18, where PDP is reduced by 58.8% with an accuracy loss of 0.81%.
Xingyu Xu 0008, Qingwen Wei, Yang Zhang 0132, Hao Cai 0001, Bo Liu 0019
CASES1