Juntao Zhao 0002

dblp:06/9805-2 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-3376-0607ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MegaScale-Data: Scaling DataLoader for Multisource Large Foundation Model Training
abstract
Modern frameworks for training large foundation models (LFMs) employ dataloaders in a data-parallel manner, with each loader processing a disjoint subset of training data. When preparing data for LFM training that originates from multiple, distinct sources, two fundamental challenges arise. First, due to the quadratic computational complexity of the attention operator, the non-uniform sample distribution over data-parallel ranks leads to significant workload imbalance among dataloaders, degrading the training efficiency. Second, supporting diverse data sources requires per-dataset file access states that are redundantly replicated across parallel loaders, consuming excessive memory. This also hinders dynamic data mixing (e.g., curriculum learning) and causes redundant access/memory overhead in hybrid parallelism.
Juntao Zhao 0002, Borui Wan, Lei Zuo 0004, Junda Feng, Jianyu Jiang, Yangrui Chen, Shuaishuai Cao, Jialing He, Kaihua Jiang, Shibiao Nong, Yanghua Peng, Haibin Lin, Chuan Wu 0001
EuroSys1
2026 BROS: Efficient LLM Serving on Hybrid Real-time and Best-effort Requests
Borui Wan, Juntao Zhao 0002, Chenyu Jiang 0002, Chuanxiong Guo, Chuan Wu 0001
INFOCOM2
2025 SplitQuant: Resource-Efficient LLM Offline Serving on Heterogeneous GPUs via Phase-Aware Model Partition and Adaptive Quantization
abstract
Modern large language models (LLMs) serving systems address distributed deployment challenges through two key techniques: distributed model partitioning for parallel computation across accelerators and quantization for reducing parameter size. While existing systems assume homogeneous GPU environments, we reveal significant untapped potential in heterogeneous systems with mixed-capacity accelerators where two critical limitations persist: (1) uniform partitioning and quantization strategies fail to adapt to hardware heterogeneity, exacerbating resource imbalance, and (2) decoupled optimization of partitioning and quantization overlooks critical performance synergies between these techniques. We present SplitQuant, a phase-aware distributed serving system that co-optimizes mixedprecision quantization, phase-aware model partitioning, and micro-batch sizing for heterogeneous environments. Our approach combines analytical modeling of quality-runtime tradeoffs with a lightweight planning algorithm to maximize throughput while preserving user-specified model quality targets. Evaluations across 10 production clusters show SplitQuant achieves up to$2.34 \times(1.61 \times$mean) higher throughput than state-of-theart approaches without violating accuracy targets. Our results underscore the value of co-designing quantization and model partitioning strategies for heterogeneous environments.
Juntao Zhao 0002, Borui Wan, Yanghua Peng, Haibin Lin, Chuan Wu 0001
CLUSTER1
2024 CDMPP: A Device-Model Agnostic Framework for Latency Prediction of Tensor Programs
abstract
Deep Neural Networks (DNNs) have shown excellent performance in a wide range of machine learning applications. Knowing the latency of running a DNN model or tensor program on a specific device is useful in various tasks, such as DNN graph- or tensor-level optimization and device selection. Considering the large space of DNN models and devices that impedes direct profiling of all combinations, recent efforts focus on building a predictor to model the performance of DNN models on different devices. However, none of the existing attempts have achieved a cost model that can accurately predict the performance of various tensor programs while supporting both training and inference accelerators. We propose CDMPP, an efficient tensor program latency prediction framework for both cross-model and cross-device prediction. We design an informative but efficient representation of tensor programs, called compact ASTs, and a pre-order-based positional encoding method, to capture the internal structure of tensor programs. We develop a domain-adaption-inspired method to learn domain-invariant representations and devise a KMeans-based sampling algorithm, for the predictor to learn from different domains (i.e., different DNN operators and devices). Our extensive experiments on a diverse range of DNN models and devices demonstrate that CDMPP significantly outperforms state-of-the-art baselines with 14.03% and 10.85% prediction error for cross-model and cross-device prediction, respectively, and one order of magnitude higher training efficiency. The implementation and the expanded dataset are available at https://github.com/joapolarbear/cdmpp.
Hanpeng Hu, Junwei Su, Juntao Zhao 0002, Yanghua Peng, Yibo Zhu 0001, Haibin Lin, Chuan Wu 0001
EuroSys3
2024 QSync: Quantization-Minimized Synchronous Distributed Training Across Hybrid Devices
abstract
A number of production deep learning clusters have attempted to explore inference hardware for DNN training, at the off-peak serving hours with many inference GPUs idling. Conducting DNN training with a combination of heterogeneous training and inference GPUs, known as hybrid device training, presents considerable challenges due to disparities in compute capability and significant differences in memory capacity. We propose QSync, a training system that enables efficient synchronous data-parallel DNN training over hybrid devices by strategically exploiting quantized operators. According to each device’s available resource capacity, QSync selects a quantization-minimized setting for operators in the distributed DNN training graph, minimizing model accuracy degradation but keeping the training efficiency brought by quantization. We carefully design a predictor with a bi-directional mixed-precision indicator to reflect the sensitivity of DNN layers on fixed-point and floating-point low-precision operators, a replayer with a neighborhood-aware cost mapper to accurately estimate the latency of distributed hybrid mixed-precision training, and then an allocator that efficiently synchronizes workers with minimized model accuracy degradation. QSync bridges the computational graph on PyTorch to an optimized backend for quantization kernel performance and flexible support for various GPU architectures. Extensive experiments show that QSync’s predictor can accurately simulate distributed mixed-precision training with < 5% error, with a consistent 0.27 − 1.03% accuracy improvement over the from-scratch training tasks compared to uniform precision.
Juntao Zhao 0002, Borui Wan, Yanghua Peng, Haibin Lin, Yibo Zhu 0001, Chuan Wu 0001
IPDPS1
2024 POSTER: LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
abstract
The immense sizes of Large-scale language models (LLMs) have led to high resource demand and cost for running the models. Though the models are largely served using uniform high-caliber GPUs nowadays, utilizing a heterogeneous cluster with a mix of available high- and low-capacity GPUs can potentially substantially reduce the serving cost. This paper proposes LLM-PQ, a system that advocates adaptive model quantization and phase-aware partition to improve LLM serving efficiency on heterogeneous GPU clusters. Extensive experiments on production inference workloads demonstrate throughput improvement in inference, showing great advantages over state-of-the-art works.
Juntao Zhao 0002, Borui Wan, Chuan Wu 0001, Yanghua Peng, Haibin Lin
PPoPP1
2023 CryptoArcade: A Cloud Gaming System With Blockchain-Based Token Economy
abstract
Cloud gaming is a novel service provisioning technology that offloads parts of game software from terminals to powerful cloud infrastructures. However, the commercial charging model for cloud gaming is still in its infancy. In this paper, we reveal the deficiencies of existing cloud gaming pricing models and propose CryptoArcade, a token-based cloud gaming system that adopts cryptocurrency as a payment method. Using cryptocurrency, CryptoArcade provides a transparent and resource-aware pricing method, enabling a time irrelevant silent payment on the floating price to protect players' interests, which avoids the Quality of Experience (QoE) degradation caused by traditional dynamic models. While CryptoArcade can solve the problem of pricing strategies, players still face decision headaches caused by having commission overhead and pre-deposit amounts on blockchains. To better understand players' trading behaviors in this decision-making, we consider a marketplace where players trade tokens through smart contracts before gaming sessions. Considering the uncertainty of future token consumption, we use Prospect Theory (PT) in modeling and obtain the optimal solution in closed form. When comparing with the benchmark expect utility theory (EUT), we show that with the same external factors, EUT players are more likely to buy tokens than PT ones.
Sizheng Fan, Juntao Zhao 0002, Zehua Wang 0001, Wei Cai 0002
IEEE Trans. Cloud Comput.2