Chenggang Zhao

dblp:254/2607 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0005-2297-0790ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Syno: Structured Synthesis for Neural Operators
abstract
The desires for better prediction accuracy and higher execution performance in neural networks never end. Neural architecture search (NAS) and tensor compilers are two popular techniques to optimize these two goals, but they are both limited to composing or optimizing existing manually designed operators rather than coming up with completely new designs. In this work, we explore the less studied direction of neural operator synthesis, which aims to automatically and efficiently discover novel neural operators with better accuracy and/or speed. We develop an end-to-end framework Syno, to realize practical neural operator synthesis. Syno makes use of a novel set of fine-grained primitives defined on tensor dimensions, which ensure various desired properties to ease model training, and also enable expression canonicalization techniques to avoid redundant candidates during search. Syno further adopts a novel guided synthesis flow to obtain valid operators matched with the specified input/output dimension sizes, and leverages efficient stochastic tree search algorithms to quickly explore the design space. We demonstrate that Syno discovers better operators with average speedups of 1.37× to 2.06× on various hardware and compiler choices, while keeping less than 1% accuracy loss even on NAS-optimized models.
Yongqi Zhuo, Zhengyuan Su, Chenggang Zhao, Mingyu Gao 0001
ASPLOS (3)3
2025 Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
abstract
The rapid scaling of large language models (LLMs) has unveiled critical limitations in current hardware architectures, including constraints in memory capacity, computational efficiency, and interconnection bandwidth.DeepSeek-V3, trained on 2,048 NVIDIA H800 GPUs, demonstrates how hardware-aware model co-design can effectively address these challenges, enabling cost-efficient training and inference at scale.This paper presents an in-depth analysis of the DeepSeek-V3/R1 model architecture and its AI infrastructure, highlighting key innovations such as Multi-head Latent Attention (MLA) for enhanced memory efficiency, Mixture of Experts (MoE) architectures for optimized computation-communication trade-offs, FP8 mixed-precision training to unlock the full potential of hardware capabilities, and a Multi-Plane Network Topology to minimize * Yuqing Wang and Liyue Zhang are the corresponding authors of this paper.
Chenggang Zhao, Chengqi Deng, Chong Ruan, Damai Dai, Huazuo Gao, Jiashi Li, Panpan Huang, Shangyan Zhou, Shirong Ma, Wenfeng Liang, Ying He 0018, Yuxuan Liu 0019, Y. X. Wei
ISCA1
2024 DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
abstract
Damai Dai, Chengqi Deng, Chenggang Zhao, R.x. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y.k. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, Wenfeng Liang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, Wenfeng Liang
ACL (1)3
2024 Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning
abstract
The rapid progress in Deep Learning (DL) and Large Language Models (LLMs) has exponentially increased demands of computational power and bandwidth. This, combined with the high costs of faster computing chips and interconnects, has significantly inflated High Performance Computing (HPC) construction costs. To address these challenges, we introduce the Fire-Flyer AI-HPC architecture, a synergistic hardware-software co-design framework and its best practices. For DL training, we deployed the Fire-Flyer 2 with 10,000 PCIe A100 GPUs, achieved performance approximating the DGX-A100 while reducing costs by half and energy consumption by $40 \%$. We specifically engineered HFReduce to accelerate allreduce communication and implemented numerous measures to keep our Computation-Storage Integrated Network congestion-free. Through our software stack, including HaiScale, 3FS, and HAI-Platform, we achieved substantial scalability by overlapping computation and communication. Our system-oriented experience from DL training provides valuable insights to drive future advancements in AI-HPC.
Xiao Bi, Guanting Chen 0002, Shanhuang Chen, Chengqi Deng, Honghui Ding, Kai Dong 0003, Qiushi Du, Kang Guan, Jianzhong Guo, Yongqiang Guo, Zhe Fu 0009, Ying He 0018, Panpan Huang, Jiashi Li, Wenfeng Liang, Xiaodong Liu 0021, Xin Liu 0126, Yiyuan Liu, Yuxuan Liu 0019, Shanghao Lu, Xiaotao Nie, Tian Pei, Junjie Qiu, Zehui Ren, Zhangli Sha, Xuecheng Su, Xiaowen Sun, Yixuan Tan, Minghui Tang, Ziwei Xie, Yiliang Xiong, Shengfeng Ye, Shuiping Yu, Yukun Zha, Mingchuan Zhang, Yichao Zhang 0004, Chenggang Zhao, Yao Zhao 0005, Shangyan Zhou, Shunfeng Zhou, Yuheng Zou
SC48
2021 Critique of "Planetary Normal Mode Computation: Parallel Algorithms, Performance, and Reproducibility" by SCC Team From Tsinghua University
abstract
In this article we present our results from the SC19 Student Cluster Competition Reproducibility Challenge. The challenge entails reproducing the article entitled “Computing Planetary Interior Normal Modes with A Highly Parallel Polynomial Filtering Eigensolver” presented at SC'18, which proposes a parallel polynomial filtered Lanczos algorithm to directly calculate the planetary normal modes of heterogeneous planets. The proposed algorithm showed excellent performance with relatively low memory consumption and high parallel efficiency. In this work, we reproduce the scaling tests in that article on a cluster using Intel Cascade Lake architecture and use the proposed algorithm to illustrate specific normal modes of Mars. We compare the results obtained on our cluster with those in the original article. We also design a new metric to better analyze the results. In addition, we use the profiling tool Intel VTune Amplifier to explain our discoveries. Our results demonstrate that the given models show great scalability, which is similar to the original article. The required normal modes of Mars are also successfully calculated and visualized.
Chen Zhang 0001, Chenggang Zhao, Jiaao He, Shengqi Chen 0001, Liyan Zheng 0001, Kezhao Huang, Jidong Zhai
IEEE Trans. Parallel Distributed Syst.2
2019 Student Cluster Competition 2018, Team Tsinghua University: Reproducing performance of multi-physics simulations of the Tsunamigenic 2004 Sumatra megathrust earthquake on the Intel Skylake Architecture
Jiaao He, Chenggang Zhao, Jiping Yu, Xinjian Yu, Liyan Zheng 0001, Chenyao Lou, Shizhi Tang, Jidong Zhai
Parallel Comput.2