Liyan Zheng 0001

dblp:254/2588-1 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0001-7327-748XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 1 first-author · 11 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 ChituDiffusion: A Data-Characteristic-Aware Serving System for Diffusion Models
abstract
Diffusion models have become the dominant approach for generative tasks in images, videos, and other domains. However, diverse data properties in generation requests, which are critical for efficient serving, remain underexploited. To address this issue, we propose a diffusion model serving system ChituDiffusion. ChituDiffusion leverages the locality of data properties to recompose a diffusion pipeline into dGraphs with shared optimization opportunities, enabling thorough compile-time and runtime co-optimizations. During compilation, ChituDiffusion compiles each dGraph into multiple execution engines optimized for specific data properties. At runtime, heterogeneous requests are elaborately reorganized into fine-grained batching tasks with similar properties and then efficiently executed by matched engines. Evaluation on five diffusion applications shows that ChituDiffusion improves the throughput by up to 2.13× (1.58× on average) on A100 and 2.19× (1.51× on average) on H100 compared with existing frameworks. The code for ChituDiffusion and the production traces have been made open-source at https://github.com/thu-pacman/chitu/tree/Diffusion.
Chengzhang Wu, Liyan Zheng 0001, Haojie Wang 0004, Kezhao Huang, Zixuan Ma, Dong Dong 0001, Jidong Zhai
PPoPP2
2025 IntelliGen: Instruction-Level Auto-tuning for Tensor Program with Monotonic Memory Optimization
abstract
Tensor compilers play a critical role in optimizing deep neural networks (DNNs), with memory performance emerging as a key bottleneck in code generation for DNN models. Existing tensor compilers are constrained by inefficient auto-tuning algorithms. They either must deploy coarse-grained descriptions, thus miss potential optimization, or struggle with vast search spaces, rendering auto-tuning inapplicable. Tensor compilers require a more holistic optimization of memory performance to overcome these constraints. To address this issue, we focus our optimization objective on memory performance, which allows us to design monotonic optimization methods, significantly enhancing the efficiency of auto-tuning and thus enabling auto-tuning on a fine-granularity description. Based on these observations, we propose IntelliGen, a tensor compiler with instruction-level auto-tuning and monotonic memory optimization. We design an instruction-level graph description, and a monotonic optimization method for optimization on . Benefiting from auto-tuning techniques with fine-grained description, IntelliGen demonstrates significant speedup of up to 3.13×, 3.55×, and 16.9× (averaging 1.46×, 1.85×, and 2.30×, respectively) on NVIDIA GPUs, AMD GPUs, and Cambricon MLUs over the most efficient existing frameworks.
Zixuan Ma, Haojie Wang 0004, Jingze Xing, Shuhong Huang, Liyan Zheng 0001, Chen Zhang 0001, Huanqi Cao, Kezhao Huang, Mingshu Zhai, Shizhi Tang, Penghan Wang, Jidong Zhai
CGO5
2025 mTuner: Accelerating Parameter-Efficient Fine-Tuning on Multi-GPU Servers with Elastic Tensor
Kezhao Huang, Siqi Zhu, Mingshu Zhai, Liyan Zheng 0001, Kinman Lei, Jiaao He, Yuyang Jin 0001, Jidong Zhai
USENIX ATC4
2024 Optimal Kernel Orchestration for Tensor Programs with Korch
abstract
Kernel orchestration is the task of mapping the computation defined in different operators of a deep neural network (DNN) to the execution of GPU kernels on modern hardware platforms. Prior approaches optimize kernel orchestration by greedily applying operator fusion, which fuses the computation of multiple operators into a single kernel, and miss a variety of optimization opportunities in kernel orchestration.
Muyan Hu, Ashwin Venkatram, Shreyashri Biswas, Balamurugan Marimuthu, Bohan Hou, Gabriele Oliaro, Haojie Wang 0004, Liyan Zheng 0001, Xupeng Miao, Jidong Zhai
ASPLOS (3)8
2024 WiseGraph: Optimizing GNN with Joint Workload Partition of Graph and Operations
abstract
Graph Neural Network (GNN) has emerged as an important workload for learning on graphs. With the size of graph data and the complexity of GNN model architectures increasing, developing an efficient GNN system grows more important. As GNN has heavy neural computation workloads on a large graph, it is crucial to partition the entire workload into smaller parts for parallel execution and optimization. However, existing approaches separately partition graph data and GNN operations, resulting in inefficiency and large data movement overhead.
Kezhao Huang, Jidong Zhai, Liyan Zheng 0001, Haojie Wang 0004, Yuyang Jin 0001, Qihao Zhang, Runqing Zhang, Zhen Zheng, Youngmin Yi, Xipeng Shen
EuroSys3
2023 EINNET: Optimizing Tensor Programs with Derivation-Based Transformations
Liyan Zheng 0001, Haojie Wang 0004, Jidong Zhai, Muyan Hu, Zixuan Ma, Tuowei Wang, Shuhong Huang, Xupeng Miao, Shizhi Tang, Kezhao Huang
OSDI1
2023 Optimizing DNNs With Partially Equivalent Transformations and Automated Corrections
abstract
Deep neural network (DNN) applications are typically represented by tensor programs. To boost the performance of DNN computations, existing works adopt fully equivalent transformations for tensor program optimization by guaranteeing the equivalence on each element of tensors. However, as there are thousands of elements in a tensor, such optimization misses the opportunities that allow the in-equivalence of minority elements. In this work, we proposePet, the first work that introduces partially equivalent transformations to optimize tensor programs. To maintain the functional equivalence of tensor programs,Petautomatically finds and corrects the in-equivalent positions by leveraging the multi-linearity of DNN computations.Petfurther uses a mutation manager to improve search efficiency. Evaluation results show thatPetcan achieve up to 1.98$\times$and 2.20$\times$speedups on NVIDIA Tesla A100 and V100 respectively compared with existing DNN frameworks by introducing new optimization opportunities of partially equivalent transformations.
Haojie Wang 0004, Jidong Zhai, Mingyu Gao 0001, Feng Zhang 0007, Tuowei Wang, Zixuan Ma, Shizhi Tang, Liyan Zheng 0001, Kaiyuan Rong, Yuanyong Chen
IEEE Trans. Computers8
2022 FreeTensor: a free-form DSL with holistic optimizations for irregular tensor programs
abstract
Tensor programs are of critical use in many domains. Existing frameworks, such as PyTorch, TensorFlow, and JAX, adopt operator-based programming to ease programming, increase performance, and perform automatic differentiation. However, as the rapid development of tensor programs, operator-based programming shows significant limitations for irregular patterns since a large amount of redundant computation or memory access is introduced.
Shizhi Tang, Jidong Zhai, Haojie Wang 0004, Liyan Zheng 0001, Zhenhao Yuan, Chen Zhang 0001
PLDI5
2022 BaGuaLu: targeting brain scale pretrained models with over 37 million cores
abstract
Large-scale pretrained AI models have shown state-of-the-art accuracy in a series of important applications. As the size of pretrained AI models grows dramatically each year in an effort to achieve higher accuracy, training such models requires massive computing and memory capabilities, which accelerates the convergence of AI and HPC. However, there are still gaps in deploying AI applications on HPC systems, which need application and system co-design based on specific hardware features.
Zixuan Ma, Jiaao He, Jiezhong Qiu, Huanqi Cao, Yuanwei Wang, Zhenbo Sun, Liyan Zheng 0001, Haojie Wang 0004, Shizhi Tang, Tianyu Zheng, Junyang Lin, Guanyu Feng, Zeqiang Huang, Aohan Zeng, Jianwei Zhang 0012, Runxin Zhong, Tianhui Shi, Jie Tang 0001, Hongxia Yang, Xin Liu 0086, Jidong Zhai
PPoPP7
2022 Vapro: performance variance detection and diagnosis for production-run parallel applications
abstract
Performance variance is a serious problem for parallel applications, which can cause performance degradation and make applications' behavior hard to understand. Therefore, detecting and diagnosing performance variance are of crucial importance for users and application developers. However, previous detection approaches either bring too large overhead and hurt applications' performance, or rely on nontrivial source code analysis that is impractical for production-run parallel applications.
Liyan Zheng 0001, Jidong Zhai, Xiongchao Tang, Haojie Wang 0004, Yuyang Jin 0001, Shuaiwen Song
PPoPP1
2022 Leveraging Code Snippets to Detect Variations in the Performance of HPC Systems
abstract
Variations in the performance of parallel and distributed systems are becoming increasingly challenging. The runtimes of different executions can vary greatly even with a fixed number of computing nodes. Many HPC applications on supercomputers exhibit such variance. This not only leads to unpredictable execution times, but also renders the system’s behavior unintuitive. The efficient online detection of variations in performance is an open problem in HPC research. To solve it, we propose an approach, calledvSensor, to detect variations in the performance of systems. The key finding of this study is that the source code of programs can better represent performance at runtime than an external detector. Specifically, many HPC applications contain code snippets that are fixed workload patterns of execution, e.g., the workload of an invariant quantity and a linearly growing workload. This observation allows us to automatically identify these snippets of workload-related code and use them to detect variations in performance. We evaluatevSensoron the Tianhe-2A system with a large number of parallel applications, and the results indicate that it can efficiently identify variations in system performance. The average overhead of 4,096 processes is less than 6% for fixed-workload v-sensors. We identify a problematic node with slow memory by usingvSensorthat degrades the performance of the program by 21%. A serious issue with network performance is also detected that slows down the Tianhe-2A system by 3.37 times for an HPC kernel.
Jidong Zhai, Liyan Zheng 0001, Jinghan Sun, Feng Zhang 0007, Xiongchao Tang, Xuehai Qian, Bingsheng He, Wei Xue 0003
IEEE Trans. Parallel Distributed Syst.2
2022 Detecting Performance Variance for Parallel Applications Without Source Code
abstract
For parallel applications, performance variance is a critical issue that can degrade performance and make applications’ behavior difficult to explain. Therefore, users and application developers should be able to detect and diagnose performance variance. Previous detection methods either introduce too much overhead and slow down applications, or rely on nontrivial source code analysis, which is impractical for production-run parallel systems. In this article, we proposeVapro, a framework for detecting and diagnosing performance variance in production-run parallel systems. Our method is based on an observation that most parallel programs contain code snippets that are executed repeatedly with a fixed workload and can be utilized to detect performance variance. We present State Transition Graph (STG) to track program execution and then do light-weight workload analysis on STG to locate performance variance.Vaprois able to successfully identify these snippets at runtime even without program source code. To diagnose the discovered variation,Vaprouses a progressive diagnosis method based on a hybrid model combining variance breakdown and statistical analysis. According to evaluating results,Vapro's performance overhead is only 1.38% on average.Vaprocan identify performance variance in real applications caused by hardware issues, such as memory and IO. The standard deviation of the execution time is decreased by up to 73.5% when the identified variance is fixed.Vaproachieves 30.0% larger detection coverage than the state-of-the-art variance detection approach based on source code analysis.
Jidong Zhai, Liyan Zheng 0001, Feng Zhang 0007, Xiongchao Tang, Haojie Wang 0004, Yuyang Jin 0001, Shuaiwen Song
IEEE Trans. Parallel Distributed Syst.2
2021 PET: Optimizing Tensor Programs with Partially Equivalent Transformations and Automated Corrections
Haojie Wang 0004, Jidong Zhai, Mingyu Gao 0001, Zixuan Ma, Shizhi Tang, Liyan Zheng 0001, Yuanzhi Li, Kaiyuan Rong, Yuanyong Chen
OSDI6
2021 Critique of "Planetary Normal Mode Computation: Parallel Algorithms, Performance, and Reproducibility" by SCC Team From Tsinghua University
abstract
In this article we present our results from the SC19 Student Cluster Competition Reproducibility Challenge. The challenge entails reproducing the article entitled “Computing Planetary Interior Normal Modes with A Highly Parallel Polynomial Filtering Eigensolver” presented at SC'18, which proposes a parallel polynomial filtered Lanczos algorithm to directly calculate the planetary normal modes of heterogeneous planets. The proposed algorithm showed excellent performance with relatively low memory consumption and high parallel efficiency. In this work, we reproduce the scaling tests in that article on a cluster using Intel Cascade Lake architecture and use the proposed algorithm to illustrate specific normal modes of Mars. We compare the results obtained on our cluster with those in the original article. We also design a new metric to better analyze the results. In addition, we use the profiling tool Intel VTune Amplifier to explain our discoveries. Our results demonstrate that the given models show great scalability, which is similar to the original article. The required normal modes of Mars are also successfully calculated and visualized.
Chen Zhang 0001, Chenggang Zhao, Jiaao He, Shengqi Chen 0001, Liyan Zheng 0001, Kezhao Huang, Jidong Zhai
IEEE Trans. Parallel Distributed Syst.5
2019 Student Cluster Competition 2018, Team Tsinghua University: Reproducing performance of multi-physics simulations of the Tsunamigenic 2004 Sumatra megathrust earthquake on the Intel Skylake Architecture
Jiaao He, Chenggang Zhao, Jiping Yu, Xinjian Yu, Liyan Zheng 0001, Chenyao Lou, Shizhi Tang, Jidong Zhai
Parallel Comput.5