Mingshu Zhai

dblp:275/3272 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
10since 2021 · last 2026
0009-0009-7573-2250ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 10 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RoMeo: Mitigating Dual-dimensional Outliers with Rotated Mixed Precision Quantization
abstract
Mixed precision quantization has been adopted to accelerate large language models (LLMs) serving by leveraging high-throughput low-precision compute units in GPUs while preserving outliers in higher precision to maintain model accuracy. However, existing methods focus on mitigating single-dimensional channel-wise outliers, leading to model accuracy degradation when scaled to 4-bit precision.
Qihao Zhang, Mingliang Tang, Mingshu Zhai, Kinman Lei, Jidong Zhai
PPoPP3
2025 IntelliGen: Instruction-Level Auto-tuning for Tensor Program with Monotonic Memory Optimization
abstract
Tensor compilers play a critical role in optimizing deep neural networks (DNNs), with memory performance emerging as a key bottleneck in code generation for DNN models. Existing tensor compilers are constrained by inefficient auto-tuning algorithms. They either must deploy coarse-grained descriptions, thus miss potential optimization, or struggle with vast search spaces, rendering auto-tuning inapplicable. Tensor compilers require a more holistic optimization of memory performance to overcome these constraints. To address this issue, we focus our optimization objective on memory performance, which allows us to design monotonic optimization methods, significantly enhancing the efficiency of auto-tuning and thus enabling auto-tuning on a fine-granularity description. Based on these observations, we propose IntelliGen, a tensor compiler with instruction-level auto-tuning and monotonic memory optimization. We design an instruction-level graph description, and a monotonic optimization method for optimization on . Benefiting from auto-tuning techniques with fine-grained description, IntelliGen demonstrates significant speedup of up to 3.13×, 3.55×, and 16.9× (averaging 1.46×, 1.85×, and 2.30×, respectively) on NVIDIA GPUs, AMD GPUs, and Cambricon MLUs over the most efficient existing frameworks.
Zixuan Ma, Haojie Wang 0004, Jingze Xing, Shuhong Huang, Liyan Zheng 0001, Chen Zhang 0001, Huanqi Cao, Kezhao Huang, Mingshu Zhai, Shizhi Tang, Penghan Wang, Jidong Zhai
CGO9
2025 TraceFlow: Efficient Trace Analysis for Large-Scale Parallel Applications via Interaction Pattern-Aware Trace Distribution
abstract
Trace analysis of large-scale parallel applications is crucial for understanding and optimizing performance. It primarily focuses on the interaction behaviors between different parallel processes, such as synchronization waits and asynchronous overlaps. The trace size explodes as the parallel scale of applications, thus current methods analyze traces in parallel to ensure analysis speed. However, due to the interaction pattern-agnostic trace distribution, they often introduce inter-process communications to fetch non-local event data during interaction analysis, leading to excessively long trace analysis time.
Yuyang Jin 0001, Xirui Shui, Mingshu Zhai, Zan Zong, Feng Zhang 0007, Felix Wolf 0001, Jidong Zhai
SC3
2025 mTuner: Accelerating Parameter-Efficient Fine-Tuning on Multi-GPU Servers with Elastic Tensor
Kezhao Huang, Siqi Zhu, Mingshu Zhai, Liyan Zheng 0001, Kinman Lei, Jiaao He, Yuyang Jin 0001, Jidong Zhai
USENIX ATC3
2025 QFactory: Accelerating Quantized Large Language Model Serving with Qtile Graphs
Qihao Zhang, Mingshu Zhai, Jidong Zhai
USENIX ATC2
2024 PUZZLE: Efficiently Aligning Large Language Models through Light-Weight Context Switch
Kinman Lei, Yuyang Jin 0001, Mingshu Zhai, Kezhao Huang, Haoxing Ye, Jidong Zhai
USENIX ATC3
2023 GraphSet: High Performance Graph Mining through Equivalent Set Transformations
abstract
Graph mining is of critical use in a number of fields such as social networks, knowledge graphs, and fraud detection. As an NP-complete problem, accelerating computation performance is the main target for current optimizations. Due to excellent performance, state-of-the-art graph mining systems mainly rely on pattern-aware algorithms. Despite previous efforts, complex control flows introduced by pattern-aware algorithms bring significant overhead and also impede further acceleration on heterogeneous hardware.
Tianhui Shi, Jidong Zhai, Haojie Wang 0004, Qiqian Chen, Mingshu Zhai, Zixu Hao
SC5
2023 SmartMoE: Efficiently Training Sparsely-Activated Models through Combining Offline and Online Parallelization
Mingshu Zhai, Jiaao He, Zixuan Ma, Zan Zong, Runqing Zhang, Jidong Zhai
USENIX ATC1
2023 Critique of "A Parallel Framework for Constraint-Based Bayesian Network Learning via Markov Blanket Discovery" by SCC Team From Tsinghua University
abstract
Srivastava et al. propose a parallel framework to optimize Bayesian network learning in the SC20 article entitled “A Parallel Framework for Constraint-Based Bayesian Network Learning via Markov Blanket Discovery”. They parallelize all the phases in network constructing algorithms to achieve high performance and scalability. In this article, we reproduce the strong scaling and weak scaling experiments in that SC article. We conduct experiments on a 4-node cluster with Intel CPUs provided by the SCC committee. We further analyze the results of communication overhead. Our results show that the proposed method in that SC article scales well on the provided cluster, in accordance with the SC article.Author: Please confirm or add details for any funding or financial support for the research of this article. ?>
Juncheng Cao, Kaiyuan Rong, Mingshu Zhai, Yanyu Ren, Yuxi Zhu, Jidong Zhai
IEEE Trans. Parallel Distributed Syst.3
2022 Critique of "MemXCT: Memory-Centric X-Ray CT Reconstruction With Massive Parallelization" by SCC Team From Tsinghua University
abstract
Hidayetoğluet al.propose a novel memory-centric algorithm to reconstruct X-ray CT images in the SC19 article entitled “MemXCT: Memory-Centric X-ray CT Reconstruction with Massive Parallelization”. They formulate the reconstruction with several SpMVs, and propose two memory-centric optimizations to improve cache locality for better memory bandwidth utilization, i.e., a two-level pseudo-Hilbert ordering and a multi-stage input buffering. In this article, we present our results on reproducing that article to show its effectiveness and generality, as part of the SC20 Student Cluster Competition Reproducibility Challenge. We reproduce the execution time and memory bandwidth tests in that article on various architectures, including Intel CPUs, AMD CPUs, and NVIDIA GPUs. We further analyze the bottleneck on different architectures by comparing the achieved memory bandwidth with the peak bandwidth on those architectures. We then reproduce the strong scaling test on CPU and GPU clusters with different scales, and use the proposed algorithm to reconstruct three new X-ray computed tomograms.
Runxin Zhong, Chen Zhang 0001, Mingshu Zhai, Lin Gan 0001, Jidong Zhai
IEEE Trans. Parallel Distributed Syst.4
2020 GraphPi: high performance graph pattern matching through effective redundancy elimination
abstract
Graph pattern matching, which aims to discover structural patterns in graphs, is considered one of the most fundamental graph mining problems in many real applications. Despite previous efforts, existing systems face two main challenges. First, inherent symmetry existing in patterns can introduce a large amount of redundant computation. Second, different matching orders for a pattern have significant performance differences and are quite hard to predict. When these factors are mixed, this problem becomes extremely complicated. High efficient pattern matching remains an open problem currently. To address these challenges, we propose GraphPi, a high performance distributed pattern matching system. GraphPi utilizes a new algorithm based on 2-cycles in group theory to generate multiple sets of asymmetric restrictions, where each set can eliminate redundant computation completely. We further design an accurate performance model to determine the optimal matching order and asymmetric restriction set for efficient pattern matching. We evaluate GraphPi on Tianhe-2A supercomputer. Results show that GraphPi outperforms the state-of-the-art system, by up to $ 105\times$ for 6 real-world graph datasets on a single node. We also scale GraphPi to 1,024 computing nodes (24,576 cores).
Tianhui Shi, Mingshu Zhai, Jidong Zhai
SC2