Fang Zheng 0007

dblp:29/5037-7 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
4since 2021 · last 2022
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2022 Haica: A High Performance Computing & Artificial Intelligence Fused Computing Architecture
Zhengbo Chen, Fang Zheng 0007, Zuoning Chen
ICA3PP2
2022 DSSA: Dual-Side Sparse Systolic Array Architecture for Accelerating Convolutional Neural Network Training
abstract
Ever-growing CNN size incurs a significant amount of redundancy in model parameters, which in turn, puts considerable burden on hardware. Unstructured pruning is widely used to reduce model sparsity. While, the irregularity introduced by unstructured pruning makes it difficult to accelerate sparse CNNs on systolic array. To address this issue, a variety of accelerators have been proposed. SIGMA, the state-of-the-art sparse GEMM accelerator, achieves significant speedup over systolic array. However, SIGMA suffers from two disadvantages: 1) it only supports one-side sparsity, leaving potential for further performance gains; 2) SIGMA improves utilization of large-sized systolic arrays at the cost of extra overhead.
Zhengbo Chen, Fang Zheng 0007, Zuoning Chen
ICPP3
2022 Evaluating performance of AI operators using roofline model
Zhengbo Chen, Fang Zheng 0007, Rujun Sun, Zuoning Chen
Appl. Intell.2
2022 AMT: asynchronous in-place matrix transpose mechanism for sunway many-core processor
Zhengbo Chen, Fang Zheng 0007, Zuoning Chen
J. Supercomput.4
2015 Cooperative Computing Techniques for a Deeply Fused and Heterogeneous Many-Core Processor Architecture
Fang Zheng 0007, Xiao-Hong Xu, Xianghui Xie 0001
J. Comput. Sci. Technol.1