Youshan Miao

dblp:06/11190 · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0002-2395-9965ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Computer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2025 LoRA Decompose: Serving Fine-Tuned Models into LoRA-Like
abstract
Large language models (LLMs) achieve remarkable performance across diverse tasks but face increasing GPU-memory demands due to the growing variety and complexity of downstream tasks. Efficient inference has thus become essential, especially for resource-limited settings. In this paper, we propose LoRA Decompose, a novel compression approach based on a key insight: instruction-fine-tuned models share a common pretrained-like base component and differ primarily through low-rank, LoRA-like delta components. Leveraging this observation, we reformulate the inference problem as a constrained optimization task that jointly identifies a shared low-rank structure across multiple models, significantly reducing their memory footprints. We solve this optimization efficiently using a custom-designed block coordinate descent algorithm, converging quickly within a few iterations. Empirical experiments with Llama-2 7B and 13B models demonstrate that our method achieves a remarkable >32x GPU memory reduction while preserving task accuracy, allowing substantial efficiency gains for practical deployment.
Yibo Han, Tangzhi Xu, Zenan Li, Youshan Miao, Yuan Yao 0001, Ningyi Xu
ECAI5
2025 AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN Training
Guanbin Xu, Zhihao Le, Yinhe Chen, Zewen Jin, Youshan Miao, Cheng Li 0001
NSDI6
2025 TrainVerify: Equivalence-Based Verification for Distributed LLM Training
abstract
Training large language models (LLMs) at scale requires parallel execution across thousands of devices, incurring enormous computational costs. Yet, these costly distributed trainings are prone to correctness bugs, causing silent errors and potentially wasting millions of GPU hours. These bugs are challenging to expose through testing.
Yunchi Lu, Youshan Miao, Cheng Tan 0005, Peng Huang 0005, Xian Zhang 0001, Fan Yang 0024
SOSP2
2024 Aceso: Efficient Parallel DNN Training through Iterative Bottleneck Alleviation
abstract
Many parallel mechanisms, including data parallelism, tensor parallelism, and pipeline parallelism, have been proposed and combined together to support training increasingly large deep neural networks (DNN) on massive GPU devices. Given a DNN model and GPU cluster, finding the optimal configuration by combining these parallelism mechanisms is an NP-hard problem. Widely adopted mathematical programming approaches have been proposed to search in a configuration subspace, but they are still too costly when scaling to large models over numerous devices.
Youshan Miao, Xiaoxiang Shi, Saeed Maleki, Fan Yang 0024, Yungang Bao, Sa Wang
EuroSys2
2024 Tessel: Boosting Distributed Execution of Large DNN Models via Flexible Schedule Search
abstract
Increasingly complex and diverse deep neural network (DNN) models necessitate distributing the execution across multiple devices for training and inference tasks, and also require carefully planned schedules for performance. However, existing practices often rely on predefined schedules that may not fully exploit the benefits of emerging diverse model-aware operator placement strategies. Handcrafting high-efficiency schedules can be challenging due to the large and varying schedule space. This paper presents Tessel, an automated system that searches for efficient schedules for distributed DNN training and inference for diverse operator placement strategies. To reduce search costs, Tessel leverages the insight that the most efficient schedules often exhibit repetitive pattern (repetend) across different data inputs. This leads to a two-phase approach: repetend construction and schedule completion. By exploring schedules for various operator placement strategies, Tessel significantly improves both training and inference performance. Experiments with representative DNN models demonstrate that Tessel achieves up to 5.5 x training performance speedup and up to 38 % inference latency reduction.
Youshan Miao, Guanbin Xu, Cheng Li 0001, Olli Saarikivi, Saeed Maleki, Fan Yang 0024
HPCA2
2024 nnScaler: Constraint-Guided Parallelization Plan Generation for Deep Learning Training
Youshan Miao, Quanlu Zhang, Fan Yang 0024, Cheng Li 0001, Saeed Maleki, Yilei Yang, Weijiang Xu, Mao Yang 0004, Lidong Zhou
OSDI2
2024 Efficient Schedule Construction for Distributed Execution of Large DNN Models
abstract
Increasingly complex and diverse deep neural network (DNN) models necessitate distributing the execution across multiple devices for training and inference tasks, and also require carefully planned schedules for performance. However, existing practices often rely on predefined schedules that may not fully exploit the benefits of emerging diverse model-aware operator placement strategies. Handcrafting high-efficiency schedules can be challenging due to the large and varying schedule space. This paper presents Tessel, an automated system that searches for efficient schedules for distributed DNN training and inference for diverse operator placement strategies. To reduce search costs, Tessel leverages the insight that the most efficient schedules often exhibit repetitive pattern (repetend) across different data inputs. This leads to a two-phase approach: repetend construction and schedule completion. By exploring schedules for various operator placement strategies, Tessel significantly improves both training and inference performance. Experiments with representative DNN models demonstrate that Tessel achieves up to 5.5× training performance speedup and up to 38% inference latency reduction.
Youshan Miao, Guanbin Xu, Cheng Li 0001, Olli Saarikivi, Saeed Maleki, Fan Yang 0024
IEEE Trans. Parallel Distributed Syst.2
2023 Adam Accumulation to Reduce Memory Footprints of Both Activations and Gradients for Large-Scale DNN Training
abstract
Running out of GPU memory has become a main bottleneck for large-scale DNN training. How to reduce the memory footprint during training has received intensive research attention. We find that previous gradient accumulation reduces activation memory but fails to be compatible with gradient memory reduction due to a contradiction between preserving gradients and releasing gradients. To address this issue, we propose a novel optimizer accumulation method for Adam, named Adam Accumulation (AdamA), which enables reducing both activation and gradient memory. Specifically, AdamA directly integrates gradients into optimizer states and accumulates optimizer states over micro-batches, so that gradients can be released immediately after use. We mathematically and experimentally demonstrate AdamA yields the same convergence properties as Adam. Evaluated on transformer-based models, AdamA achieves up to 23% memory reduction compared to gradient accumulation with less than 2% degradation in training throughput. Notably, AdamA can work together with memory reduction methods for optimizer states to fit 1.26×~3.14× larger models over PyTorch and DeepSpeed baseline on GPUs with different memory capacities.
Yibo Han, Shijie Cao, Guohao Dai 0001, Youshan Miao, Ting Cao 0003, Fan Yang 0024, Ningyi Xu
ECAI5
2022 Breaking the computation and communication abstraction barrier in distributed machine learning workloads
abstract
Recent trends towards large machine learning models require both training and inference tasks to be distributed. Considering the huge cost of training these models, it is imperative to unlock optimizations in computation and communication to obtain best performance. However, the current logical separation between computation and communication kernels in machine learning frameworks misses optimization opportunities across this barrier. Breaking this abstraction can provide many optimizations to improve the performance of distributed workloads. However, manually applying these optimizations requires modifying the underlying computation and communication libraries for each scenario, which is both time consuming and error-prone.
Abhinav Jangda, Amir Hossein Nodehi Sabet, Saeed Maleki, Youshan Miao, Madan Musuvathi, Todd Mytkowicz, Olli Saarikivi
ASPLOS6
2021 Efficient Data Loader for Fast Sampling-Based GNN Training on Large Graphs
abstract
Emerging graph neural networks (GNNs) have extended the successes of deep learning techniques against datasets like images and texts to more complex graph-structured data. By leveraging GPU accelerators, existing frameworks combine mini-batch and sampling for effective and efficient model training on large graphs. However, this setup faces a scalability issue since loading rich vertex features from CPU to GPU through a limited bandwidth link usually dominates the training cycle. In this article, we propose PaGraph, a novel, efficient data loader that supports general and efficient sampling-based GNN training on single-server with multi-GPU. PaGraph significantly reduces the data loading time by exploiting available GPU resources to keep frequently-accessed graph data with a cache. It also embodies a lightweight yet effective caching policy that takes into account graph structural information and data access patterns of sampling-based GNN training simultaneously. Furthermore, to scale out on multiple GPUs, PaGraph develops a fast GNN-computation-aware partition algorithm to avoid cross-partition access during data-parallel training and achieves better cache efficiency. Finally, it overlaps data loading and GNN computation for further hiding loading costs. Evaluations on two representative GNN models, GCN and GraphSAGE, using two sampling methods, Neighbor and Layer-wise, show that PaGraph could eliminate the data loading time from the GNN training pipeline, and achieve up to 4.8× performance speedup over the state-of-the-art baselines. Together with preprocessing optimization, PaGraph further delivers up to 16.0× end-to-end speedup.
Youhui Bai, Cheng Li 0001, Yufei Wu 0011, Youshan Miao, Yunxin Liu 0001, Yinlong Xu 0001
IEEE Trans. Parallel Distributed Syst.5
2020 PaGraph: Scaling GNN training on large graphs via computation-aware caching
abstract
Emerging graph neural networks (GNNs) have extended the successes of deep learning techniques against datasets like images and texts to more complex graph-structured data. By leveraging GPU accelerators, existing frameworks combine both mini-batch and sampling for effective and efficient model training on large graphs. However, this setup faces a scalability issue since loading rich vertices features from CPU to GPU through a limited bandwidth link usually dominates the training cycle. In this paper, we propose PaGraph, a system that supports general and efficient sampling-based GNN training on single-server with multi-GPU. PaGraph significantly reduces the data loading time by exploiting available GPU resources to keep frequently accessed graph data with a cache. It also embodies a lightweight yet effective caching policy that takes into account graph structural information and data access patterns of sampling-based GNN training simultaneously. Furthermore, to scale out on multiple GPUs, PaGraph develops a fast GNN-computation-aware partition algorithm to avoid cross-partition access during data parallel training and achieves better cache efficiency. Evaluations on two representative GNN models, GCN and GraphSAGE, show that PaGraph achieves up to 96.8% data loading time reductions and up to 4.8X performance speedup over the state-of-the-art baselines. Together with preprocessing optimization, PaGraph further delivers up to 16.0X end-to-end speedup.
Cheng Li 0001, Youshan Miao, Yunxin Liu 0001, Yinlong Xu 0001
SoCC3
2020 Motif-Preserving Temporal Network Embedding
abstract
Network embedding, mapping nodes in a network to a low-dimensional space, achieves powerful performance. An increasing number of works focus on static network embedding, however, seldom attention has been paid to temporal network embedding, especially without considering the effect of mesoscopic dynamics when the network evolves. In light of this, we concentrate on a particular motif --- triad --- and its temporal dynamics, to study the temporal network embedding. Specifically, we propose MTNE, a novel embedding model for temporal networks. MTNE not only integrates the Hawkes process to stimulate the triad evolution process that preserves motif-aware high-order proximities, but also combines attention mechanism to distinguish the importance of different types of triads better. Experiments on various real-world temporal networks demonstrate that, compared with several state-of-the-art methods, our model achieves the best performance in both static and dynamic tasks, including node classification, link prediction, and link recommendation.
Hong Huang 0001, Zixuan Fang, Xiao Wang 0017, Youshan Miao, Hai Jin 0001
IJCAI4
2020 Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks
Lingxiao Ma, Zhi Yang 0001, Jilong Xue, Youshan Miao, Wenxiang Hu, Fan Yang 0024, Lidong Zhou
OSDI5
2020 Distributed Graph Computation Meets Machine Learning
abstract
TuX2is a new distributed graph engine that bridges graph computation and distributed machine learning.TuX2inherits the benefits of elegant graph computation model, efficient graph layout, and balanced parallelism to scale to billion-edge graphs, while extended and optimized for distributed machine learning to support heterogeneity in data model, Stale Synchronous Parallel in scheduling, and a new Mini-batch, Exchange, GlobalSync, and Apply (MEGA) model for programming.TuX2further introduces a hybrid vertex-cut graph optimization and supports various consistency models in fault tolerance for machine learning. We have developed a set of representative distributed machine learning algorithms inTuX2, covering both supervised and unsupervised learning. Compared to the implementations on distributed machine learning platforms, writing those algorithms inTuX2takes only about 25 percent of the code: our graph computation model hides the detailed management of data layout, partitioning, and parallelism from developers. The extensive evaluation ofTuX2, using large datasets with up to 64 billion of edges, shows thatTuX2outperforms PowerGraph/PowerLyra, the state-of-the-art distributed graph engines, by an order of magnitude, while beating two state-of-the-art distributed machine learning systems by at least 60 percent.
Wencong Xiao, Jilong Xue, Youshan Miao, Ming Wu 0007, Wei Li 0022, Lidong Zhou
IEEE Trans. Parallel Distributed Syst.3
2019 Fast Distributed Deep Learning over RDMA
abstract
Deep learning emerges as an important new resource-intensive workload and has been successfully applied in computer vision, speech, natural language processing, and so on. Distributed deep learning is becoming a necessity to cope with growing data and model sizes. Its computation is typically characterized by a simple tensor data abstraction to model multi-dimensional matrices, a dataflow graph to model computation, and iterative executions with relatively frequent synchronizations, thereby making it substantially different from Map/Reduce style distributed big data computation.
Jilong Xue, Youshan Miao, Ming Wu 0007, Lidong Zhou
EuroSys2
2019 NeuGraph: Parallel Deep Neural Network Computation on Large Graphs
Lingxiao Ma, Zhi Yang 0001, Youshan Miao, Jilong Xue, Ming Wu 0007, Lidong Zhou, Yafei Dai
USENIX ATC3
2017 Tux2: Distributed Graph Computation for Machine Learning
Wencong Xiao, Jilong Xue, Youshan Miao, Ming Wu 0007, Wei Li 0022, Lidong Zhou
NSDI3
2015 GraM: scaling graph computation to the trillions
abstract
GraM is an efficient and scalable graph engine for a large class of widely used graph algorithms. It is designed to scale up to multicores on a single server, as well as scale out to multiple servers in a cluster, offering significant, often over an order-of-magnitude, improvement over existing distributed graph engines on evaluated graph algorithms. GraM is also capable of processing graphs that are significantly larger than previously reported. In particular, using 64 servers (1,024 physical cores), it performs a PageRank iteration in 140 seconds on a synthetic graph with over one trillion edges, setting a new milestone for graph engines.
Ming Wu 0007, Fan Yang 0024, Jilong Xue, Wencong Xiao, Youshan Miao, Haoxiang Lin, Yafei Dai, Lidong Zhou
SoCC5
2015 ImmortalGraph: A System for Storage and Analysis of Temporal Graphs
abstract
Temporal graphs that capture graph changes over time are attracting increasing interest from research communities, for functions such as understanding temporal characteristics of social interactions on a time-evolving social graph. ImmortalGraph is a storage and execution engine designed and optimized specifically for temporal graphs. Locality is at the center of ImmortalGraph’s design: temporal graphs are carefully laid out in both persistent storage and memory, taking into account data locality in both time and graph-structure dimensions. ImmortalGraph introduces the notion of locality-aware batch scheduling in computation, so that common “bulk” operations on temporal graphs are scheduled to maximize the benefit of in-memory data locality. The design of ImmortalGraph explores an interesting interplay among locality, parallelism, and incremental computation in supporting common mining tasks on temporal graphs. The result is a high-performance temporal-graph system that is up to 5 times more efficient than existing database solutions for graph queries. The locality optimizations in ImmortalGraph offer up to an order of magnitude speedup for temporal iterative graph mining compared to a straightforward application of existing graph engines on a series of snapshots.
Youshan Miao, Ming Wu 0007, Fan Yang 0024, Lidong Zhou, Vijayan Prabhakaran, Enhong Chen
ACM Trans. Storage1
2014 Chronos: a graph engine for temporal graph analysis
abstract
Temporal graphs capture changes in graphs over time and are becoming a subject that attracts increasing interest from the research communities, for example, to understand temporal characteristics of social interactions on a time-evolving social graph. Chronos is a storage and execution engine designed and optimized specifically for running in-memory iterative graph computation on temporal graphs. Locality is at the center of the Chronos design, where the in-memory layout of temporal graphs and the scheduling of the iterative computation on temporal graphs are carefully designed, so that common "bulk" operations on temporal graphs are scheduled to maximize the benefit of in-memory data locality. The design of Chronos further explores the interesting interplay among locality, parallelism, and incremental computation in supporting common mining tasks on temporal graphs. The result is a high-performance temporal-graph system that offers up to an order of magnitude speedup for temporal iterative graph mining compared to a straightforward application of existing graph engines on a series of snapshots.
Youshan Miao, Ming Wu 0007, Fan Yang 0024, Lidong Zhou, Vijayan Prabhakaran, Enhong Chen
EuroSys2
2012 Kineograph: taking the pulse of a fast-changing and connected world
abstract
Kineograph is a distributed system that takes a stream of incoming data to construct a continuously changing graph, which captures the relationships that exist in the data feed. As a computing platform, Kineograph further supports graph-mining algorithms to extract timely insights from the fast-changing graph structure. To accommodate graph-mining algorithms that assume a static underlying graph, Kineograph creates a series of consistent snapshots, using a novel and efficient epoch commit protocol. To keep up with continuous updates on the graph, Kineograph includes an incremental graph-computation engine. We have developed three applications on top of Kineograph to analyze Twitter data: user ranking, approximate shortest paths, and controversial topic detection. For these applications, Kineograph takes a live Twitter data feed and maintains a graph of edges between all users and hashtags. Our evaluation shows that with 40 machines processing 100K tweets per second, Kineograph is able to continuously compute global properties, such as user ranks, with less than 2.5-minute timeliness guarantees. This rate of traffic is more than 10 times the reported peak rate of Twitter as of October 2011.
Raymond Cheng 0001, Aapo Kyrola, Youshan Miao, Xuetian Weng, Ming Wu 0007, Fan Yang 0024, Lidong Zhou, Feng Zhao 0001, Enhong Chen
EuroSys4