Shizhi Tang

dblp:252/1992 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-6543-0859ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ParDiff: Efficiently Parallelizing Reverse-Mode Automatic Differentiation with Direct Indexing
abstract
Automatic Differentiation (AD) is a technique that computes the derivatives of numerical programs by systematically applying the chain rule, playing a critical role in domains such as machine learning, simulation, and control systems. However, parallelizing differentiated programs remains a significant challenge due to the conflict between tapes (a data structure for intermediate variable storage) and summations: the differentiation process inherently introduces inter-thread summation patterns, which require prohibitively expensive atomic operations; and traditional tape designs tightly couple data retrieval with the program’s control flow, preventing code restructuring needed to eliminate these costly dependencies.
Shuhong Huang, Shizhi Tang, Yuan Wen, Huanqi Cao, Ruibai Tang, Yidong Chen 0003, Jiping Yu, Jidong Zhai
PPoPP2
2025 IntelliGen: Instruction-Level Auto-tuning for Tensor Program with Monotonic Memory Optimization
abstract
Tensor compilers play a critical role in optimizing deep neural networks (DNNs), with memory performance emerging as a key bottleneck in code generation for DNN models. Existing tensor compilers are constrained by inefficient auto-tuning algorithms. They either must deploy coarse-grained descriptions, thus miss potential optimization, or struggle with vast search spaces, rendering auto-tuning inapplicable. Tensor compilers require a more holistic optimization of memory performance to overcome these constraints. To address this issue, we focus our optimization objective on memory performance, which allows us to design monotonic optimization methods, significantly enhancing the efficiency of auto-tuning and thus enabling auto-tuning on a fine-granularity description. Based on these observations, we propose IntelliGen, a tensor compiler with instruction-level auto-tuning and monotonic memory optimization. We design an instruction-level graph description, and a monotonic optimization method for optimization on . Benefiting from auto-tuning techniques with fine-grained description, IntelliGen demonstrates significant speedup of up to 3.13×, 3.55×, and 16.9× (averaging 1.46×, 1.85×, and 2.30×, respectively) on NVIDIA GPUs, AMD GPUs, and Cambricon MLUs over the most efficient existing frameworks.
Zixuan Ma, Haojie Wang 0004, Jingze Xing, Shuhong Huang, Liyan Zheng 0001, Chen Zhang 0001, Huanqi Cao, Kezhao Huang, Mingshu Zhai, Shizhi Tang, Penghan Wang, Jidong Zhai
CGO10
2023 EINNET: Optimizing Tensor Programs with Derivation-Based Transformations
Liyan Zheng 0001, Haojie Wang 0004, Jidong Zhai, Muyan Hu, Zixuan Ma, Tuowei Wang, Shuhong Huang, Xupeng Miao, Shizhi Tang, Kezhao Huang
OSDI9
2023 Unified Programming Models for Heterogeneous High-Performance Computers
Zixuan Ma, Yuyang Jin 0001, Shizhi Tang, Haojie Wang 0004, Wei-Cheng Xue, Jidong Zhai
J. Comput. Sci. Technol.3
2023 Mat2Stencil: A Modular Matrix-Based DSL for Explicit and Implicit Matrix-Free PDE Solvers on Structured Grid
abstract
Partial differential equation (PDE) solvers are extensively utilized across numerous scientific and engineering fields. However, achieving high performance and scalability often necessitates intricate and low-level programming, particularly when leveraging deterministic sparsity patterns in structured grids. In this paper, we propose an innovative domain-specific language (DSL), Mat2Stencil, with its compiler, for PDE solvers on structured grids. Mat2Stencil introduces a structured sparse matrix abstraction, facilitating modular, flexible, and easy-to-use expression of solvers across a broad spectrum, encompassing components such as Jacobi or Gauss-Seidel preconditioners, incomplete LU or Cholesky decompositions, and multigrid methods built upon them. Our DSL compiler subsequently generates matrix-free code consisting of generalized stencils through multi-stage programming. The code allows spatial loop-carried dependence in the form of quasi-affine loops, in addition to the Jacobi-style stencil’s embarrassingly parallel on spatial dimensions. We further propose a novel automatic parallelization technique for the spatially dependent loops, which offers a compile-time deterministic task partitioning for threading, calculates necessary inter-thread synchronization automatically, and generates an efficient multi-threaded implementation with fine-grained synchronization. Implementing 4 benchmarking programs, 3 of them being the pseudo-applications in NAS Parallel Benchmarks with 6.3% lines of code and 1 being matrix-free High Performance Conjugate Gradients with 16.4% lines of code, we achieve up to 1.67× and on average 1.03× performance compared to manual implementations.
Huanqi Cao, Shizhi Tang, Qianchao Zhu, Bowen Yu 0003
Proc. ACM Program. Lang.2
2023 Optimizing DNNs With Partially Equivalent Transformations and Automated Corrections
abstract
Deep neural network (DNN) applications are typically represented by tensor programs. To boost the performance of DNN computations, existing works adopt fully equivalent transformations for tensor program optimization by guaranteeing the equivalence on each element of tensors. However, as there are thousands of elements in a tensor, such optimization misses the opportunities that allow the in-equivalence of minority elements. In this work, we proposePet, the first work that introduces partially equivalent transformations to optimize tensor programs. To maintain the functional equivalence of tensor programs,Petautomatically finds and corrects the in-equivalent positions by leveraging the multi-linearity of DNN computations.Petfurther uses a mutation manager to improve search efficiency. Evaluation results show thatPetcan achieve up to 1.98$\times$and 2.20$\times$speedups on NVIDIA Tesla A100 and V100 respectively compared with existing DNN frameworks by introducing new optimization opportunities of partially equivalent transformations.
Haojie Wang 0004, Jidong Zhai, Mingyu Gao 0001, Feng Zhang 0007, Tuowei Wang, Zixuan Ma, Shizhi Tang, Liyan Zheng 0001, Kaiyuan Rong, Yuanyong Chen
IEEE Trans. Computers7
2022 FreeTensor: a free-form DSL with holistic optimizations for irregular tensor programs
abstract
Tensor programs are of critical use in many domains. Existing frameworks, such as PyTorch, TensorFlow, and JAX, adopt operator-based programming to ease programming, increase performance, and perform automatic differentiation. However, as the rapid development of tensor programs, operator-based programming shows significant limitations for irregular patterns since a large amount of redundant computation or memory access is introduced.
Shizhi Tang, Jidong Zhai, Haojie Wang 0004, Liyan Zheng 0001, Zhenhao Yuan, Chen Zhang 0001
PLDI1
2022 BaGuaLu: targeting brain scale pretrained models with over 37 million cores
abstract
Large-scale pretrained AI models have shown state-of-the-art accuracy in a series of important applications. As the size of pretrained AI models grows dramatically each year in an effort to achieve higher accuracy, training such models requires massive computing and memory capabilities, which accelerates the convergence of AI and HPC. However, there are still gaps in deploying AI applications on HPC systems, which need application and system co-design based on specific hardware features.
Zixuan Ma, Jiaao He, Jiezhong Qiu, Huanqi Cao, Yuanwei Wang, Zhenbo Sun, Liyan Zheng 0001, Haojie Wang 0004, Shizhi Tang, Tianyu Zheng, Junyang Lin, Guanyu Feng, Zeqiang Huang, Aohan Zeng, Jianwei Zhang 0012, Runxin Zhong, Tianhui Shi, Jie Tang 0001, Hongxia Yang, Xin Liu 0086, Jidong Zhai
PPoPP9
2021 PET: Optimizing Tensor Programs with Partially Equivalent Transformations and Automated Corrections
Haojie Wang 0004, Jidong Zhai, Mingyu Gao 0001, Zixuan Ma, Shizhi Tang, Liyan Zheng 0001, Yuanzhi Li, Kaiyuan Rong, Yuanyong Chen
OSDI5
2019 Student Cluster Competition 2018, Team Tsinghua University: Reproducing performance of multi-physics simulations of the Tsunamigenic 2004 Sumatra megathrust earthquake on the Intel Skylake Architecture
Jiaao He, Chenggang Zhao, Jiping Yu, Xinjian Yu, Liyan Zheng 0001, Chenyao Lou, Shizhi Tang, Jidong Zhai
Parallel Comput.7