Xiaoqiang Dan

dblp:351/5977 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Dynamic Scheduling for AI Accelerators via TISA
Guanghui Song, Xiaoqiang Dan, Chengke Wang, Wenyuan Lv, Zhongzhou Jiang, Jianjian Guan, Teng Lu, Weixing Pan, Zirong Shen, Jie Zhao 0002
ISCA2
2023 Effectively Scheduling Computational Graphs of Deep Neural Networks toward Their Domain-Specific Accelerators
Jie Zhao 0002, Siyuan Feng 0007, Xiaoqiang Dan, Chengke Wang, Sheng Yuan, Wenyuan Lv, Qikai Xie
OSDI3
2022 Cost-Aware TVM (CAT) Tensorization for Modern Deep Learning Accelerators
abstract
TVM is proposed to efficiently deploy deep learning (DL) networks to diverse hardware devices. Tensorization in TVM enables designers to manually set up mappings between tensor computation patterns in DL networks and specific hardware tensor instructions, so as to fully exploit the potential of DL accelerators. However, with the development of modem DL accelerators, a much richer set of hardware instructions are being proposed. This leads to a one-to-many mapping between a tensor computation pattern and hardware instructions, making it a non-trivial task to choose a proper mapping. To this end, we propose Cost Aware TVM (CAT) tensorization. The key insight of CAT is that both tensor computation patterns and instructions share similar tensor features. Based on this insight, CAT first extracts the tensor features of tensor computation patterns and locates eligible hardware instructions accordingly, and then selects the least time consuming hardware instruction via a cost model. Our experimental results across 8 deep learning models show that CAT tensorization can automatically map tensor computation patterns to proper hardware instructions, and our cost model based selection policy achieves an average speed up of 6.4x compared with randomly selecting an eligible hardware instruction for a tensor computation pattern.
Yahang Hu, Xiaoqiang Dan
ICCD3