Xingzhou Cheng

dblp:194/0883 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
2since 2021 · last 2023
0000-0001-7623-2033ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Hardware accelerators and domain-specific architectures · 53% Memory systems · 47%
Artificial intelligence
2 papers
Efficient and distributed learning · 84% Deep learning architectures and training · 16%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.732023
NAND-SPIN-based processing-in-MRAM architecture for convolutional neural network acceleration · Sci. China Inf. Sci. 2023
S2 Engine: A Novel Systolic Architecture for Sparse Convolutional Neural Networks · IEEE Trans. Computers 2022
SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks Training · DAC 2020
Memory systems › processing-in-memory
processing-in-MRAM
1.122023
NAND-SPIN-based processing-in-MRAM architecture for convolutional neural network acceleration · Sci. China Inf. Sci. 2023
TCIM: Triangle Counting Acceleration With Processing-In-MRAM Architecture · DAC 2020
Memory systems
processing-in-memory
0.712023
NAND-SPIN-based processing-in-MRAM architecture for convolutional neural network acceleration · Sci. China Inf. Sci. 2023
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator
0.612022
S2 Engine: A Novel Systolic Architecture for Sparse Convolutional Neural Networks · IEEE Trans. Computers 2022
Hardware accelerators and domain-specific architectures
systolic array
0.612022
S2 Engine: A Novel Systolic Architecture for Sparse Convolutional Neural Networks · IEEE Trans. Computers 2022
Machine learning › Efficient and distributed learning
model compression
0.412020
SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks Training · DAC 2020
Machine learning › Efficient and distributed learning › model compression
pruning
0.412020
SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks Training · DAC 2020
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
CNN training accelerator
0.412020
SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks Training · DAC 2020
Memory systems
in-memory computing
0.412020
TCIM: Triangle Counting Acceleration With Processing-In-MRAM Architecture · DAC 2020
Graph algorithms and graph theory › subgraph counting
triangle counting
0.412020
TCIM: Triangle Counting Acceleration With Processing-In-MRAM Architecture · DAC 2020
Memory systems › non-volatile memory
magnetic random access memory
0.322023
NAND-SPIN-based processing-in-MRAM architecture for convolutional neural network acceleration · Sci. China Inf. Sci. 2023
TCIM: Triangle Counting Acceleration With Processing-In-MRAM Architecture · DAC 2020
Memory systems
non-volatile memory
0.322023
NAND-SPIN-based processing-in-MRAM architecture for convolutional neural network acceleration · Sci. China Inf. Sci. 2023
TCIM: Triangle Counting Acceleration With Processing-In-MRAM Architecture · DAC 2020
Machine learning › Deep learning architectures and training › convolutional neural network › convolutional neural network architecture
sparse convolutional networks
0.212022
S2 Engine: A Novel Systolic Architecture for Sparse Convolutional Neural Networks · IEEE Trans. Computers 2022

Methods — techniques the papers use, named apart from their topics

dynamic data selection · 1.1compressed dataflow · 1.1stochastic pruning · 0.9graph slicing · 0.9data mapping · 0.9bitwise logic operations · 0.91-dimensional convolution dataflow · 0.9
YearPublicationVenuePosition
2023 NAND-SPIN-based processing-in-MRAM architecture for convolutional neural network acceleration
Yinglin Zhao, Jianlei Yang 0001, Bing Li 0017, Xingzhou Cheng, Xucheng Ye, Xiaotao Jia, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001
Sci. China Inf. Sci.4
2022 S2 Engine: A Novel Systolic Architecture for Sparse Convolutional Neural Networks
abstract
Convolutional neural networks (CNNs) have achieved great success in performing cognitive tasks. However, execution of CNNs requires a large amount of computing resources and generates heavy memory traffic, which impose a severe challenge on computing system design. Through optimizing parallel executions and data reuse in convolution, systolic architecture demonstrates great advantages in accelerating CNN computations. However, regular internal data transmission path in traditional systolic architecture prevents the systolic architecture from completely leveraging the benefits introduced by neural network sparsity.Deployment of fine-grained sparsity on the existing systolic architectures is greatly hindered by the incurred computational overheads.In this work, we propose S2Engine a novel systolic architecture that can fully exploit the sparsity in CNNs with maximized data reuse. S2Engine transmits compressed data internally and allows each processing element to dynamically select an aligned data from the compressed dataflow in convolution. Compared to the naive systolic array, S2Engine achieves about 3.2 and about 3.0 improvements on speed and power efficiency, respectively.
Jianlei Yang 0001, Wenzhi Fu, Xingzhou Cheng, Xucheng Ye, Pengcheng Dai, Weisheng Zhao 0001
IEEE Trans. Computers3
2020 SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks Training
abstract
Training Convolutional Neural Networks (CNNs) usually requires a large number of computational resources. In this paper, SparseTrain is proposed to accelerate CNN training by fully exploiting the sparsity. It mainly involves three levels of innovations: activation gradients pruning algorithm, sparse training dataflow, and accelerator architecture. By applying a stochastic pruning algorithm on each layer, the sparsity of back-propagation gradients can be increased dramatically without degrading training accuracy and convergence rate. Moreover, to utilize both natural sparsity (resulted from ReLU or Pooling layers) and artificial sparsity (brought by pruning algorithm), a sparse-aware architecture is proposed for training acceleration. This architecture supports forward and back-propagation of CNN by adopting 1-Dimensional convolution dataflow. We have built a cycle-accurate architecture simulator to evaluate the performance and efficiency based on the synthesized design with 14nm FinFET technologies. Evaluation results on AlexNet/ResNet show that SparseTrain could achieve about 2.7× speedup and 2.2× energy efficiency improvement on average compared with the original training process.
Pengcheng Dai, Jianlei Yang 0001, Xucheng Ye, Xingzhou Cheng, Junyu Luo 0002, Linghao Song, Yiran Chen 0001, Weisheng Zhao 0001
DAC4
2020 TCIM: Triangle Counting Acceleration With Processing-In-MRAM Architecture
abstract
Triangle counting (TC) is a fundamental problem in graph analysis and has found numerous applications, which motivates many TC acceleration solutions in the traditional computing platforms like GPU and FPGA. However, these approaches suffer from the bandwidth bottleneck because TC calculation involves a large amount of data transfers. In this paper, we propose to overcome this challenge by designing a TC accelerator utilizing the emerging processing-in-MRAM (PIM) architecture. The true innovation behind our approach is a novel method to perform TC with bitwise logic operations (such as AND), instead of the traditional approaches such as matrix computations. This enables the efficient in-memory implementations of TC computation, which we demonstrate in this paper with computational Spin-Transfer Torque Magnetic RAM (STT-MRAM) arrays. Furthermore, we develop customized graph slicing and mapping techniques to speed up the computation and reduce the energy consumption. We use a device-to-architecture co-simulation framework to validate our proposed TC accelerator. The results show that our data mapping strategy could reduce 99.99% of the computation and 72% of the memory WRITE operations. Compared with the existing GPU or FPGA accelerators, our in-memory accelerator achieves speedups of 9× and 23.4×, respectively, and a 20.6× energy efficiency improvement over the FPGA accelerator.
Jianlei Yang 0001, Yinglin Zhao, Yingjie Qi, Meichen Liu, Xingzhou Cheng, Xiaotao Jia, Gang Qu 0001, Weisheng Zhao 0001
DAC6