Yinmin Zhong

dblp:339/0691 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-2504-7652ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FastServe: Iteration-Level Preemptive Scheduling for Large Language Model Inference
Bingyang Wu, Yinmin Zhong, Fangyue Liu, Yuanhang Sun, Gang Huang 0001, Xuanzhe Liu, Xin Jin 0008
NSDI2
2026 DistRS: Disaggregated Reward Service for RLVR with Batch-Level Constraint
Ruidong Zhu, Mingcong Han, Yinmin Zhong, Wencong Xiao, Xuanzhe Liu, Xin Jin 0008
NSDI3
2026 DualPath: Accelerating Agentic LLM Inference by Harvesting Disaggregated KV-Cache Storage I/O
abstract
The performance of multi-turn, agentic LLM inference is increasingly dominated by KV-Cache storage I/O rather than computation. In prevalent disaggregated architectures, loading the massive KV-Cache from external storage creates a fundamental imbalance: storage NICs on prefill engines become bandwidth-saturated, while those on decoding engines remain idle. This asymmetry severely constrains overall system throughput.
Yongtong Wu, Shaoyuan Chen, Rilin Huang, Yixuan Tan, Yinmin Zhong, Xin Jin 0008, Panpan Huang
SIGCOMM5
2025 Optimizing RLHF Training for Large Language Models with Stage Fusion
Yinmin Zhong, Bingyang Wu, Changyi Wan, Hanpeng Hu, Ranchen Ming, Yibo Zhu 0001, Xin Jin 0008
NSDI1
2025 DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
abstract
Multimodal large language models (LLMs) empower LLMs to ingest inputs and generate outputs in multiple forms, such as text, image, and audio. However, the integration of multiple modalities introduces heterogeneity in both the model and training data, creating unique systems challenges.
Yinmin Zhong, Hanpeng Hu, Jianjian Sun, Zheng Ge, Yibo Zhu 0001, Daxin Jiang, Xin Jin 0008
SIGCOMM2
2024 MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUs
Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang 0001, Yangrui Chen, Zhi Zhang 0005, Yanghua Peng, Xiang Li 0067, Shibiao Nong, Yulu Jia, Sun He, Hongmin Chen, Zhihao Bai, Qi Hou, Shipeng Yan, Yiyao Sheng, Zhuo Jiang, Haohan Xu, Zhang Zhang 0003, Pengfei Nie, Leqi Zou, Sida Zhao, Zherui Liu, Xiaoying Jia 0001, Jianxi Ye, Xin Jin 0008, Xin Liu 0086
NSDI3
2024 DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
Yinmin Zhong, Junda Chen, Jianbo Hu, Yibo Zhu 0001, Xuanzhe Liu, Xin Jin 0008, Hao Zhang 0025
OSDI1
2024 LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
abstract
The context window of large language models (LLMs) is rapidly increasing, leading to a huge variance in resource usage between different requests as well as between different phases of the same request. Restricted by static parallelism strategies, existing LLM serving systems cannot efficiently utilize the underlying resources to serve variable-length requests in different phases. To address this problem, we propose a new parallelism paradigm, elastic sequence parallelism (ESP), to elastically adapt to the variance across different requests and phases. Based on ESP, we design and build LoongServe, an LLM serving system that (1) improves computation efficiency by elastically adjusting the degree of parallelism in real-time, (2) improves communication efficiency by reducing key-value cache migration overhead and overlapping partial decoding communication with computation, and (3) improves GPU memory efficiency by reducing key-value cache fragmentation across instances. Our evaluation under diverse real-world datasets shows that LoongServe improves the throughput by up to 3.85× compared to chunked prefill and 5.81× compared to prefill-decoding disaggregation.
Bingyang Wu, Yinmin Zhong, Peng Sun 0006, Xuanzhe Liu, Xin Jin 0008
SOSP3
2024 DistMind: Efficient Resource Disaggregation for Deep Learning Workloads
abstract
Deep learning (DL) systems suffer from low resource utilization due to 1) monolithic server model that tightly couples compute and memory; and 2) limited sharing between different inference applications, and across inference and training, because of strict service level objectives (SLOs). To address this problem, we present, a disaggregated DL system that enables efficient multiplexing of DL applications with near-optimal resource utilization. decouples compute from host memory, and exposes the abstractions of a GPU pool and a memory pool, each of which can be independently provisioned. The key challenge is to dynamically allocate GPU resources to different applications based on their real-time demands while meeting strict SLOs. We tackle this challenge by exploiting the power of high-speed 100 Gbps networks, and design three-stage pipelining, cache-aware load balancing, and DNN-aware sharding mechanisms based on the characteristics of DL workloads, to achieve millisecond-scale application loading overhead and improve system efficiency. We have implemented a prototype of and integrated it with PyTorch. Experimental results on AWS EC2 show that achieves near 100% resource utilization, and compared with NVIDIA MPS and Ray, improves the throughput by up to 279% and reduces the inference latency by up to 94%.
Xin Jin 0008, Zhihao Bai, Zhen Zhang 0063, Yibo Zhu 0001, Yinmin Zhong, Xuanzhe Liu
IEEE/ACM Trans. Netw.5
2024 Aquifer: Transparent Microsecond-Scale Scheduling for vRAN Workloads
abstract
Virtual Radio Access Network (vRAN) is an emerging approach offered by cloud providers to accelerate 5G services deployment. Despite significant microsecond-scale traffic variations, vRAN instances are provisioned based on peak load to meet strict latency requirements, leading to significant resource waste. Conceivably, vRAN can share CPUs with other applications to increase CPU utilization. Yet, existing sharing solutions require modifications to vRAN source code, hindering their deployment on public clouds. We present Aquifer, a microsecond-scale scheduler providing transparent CPU sharing for vRAN workloads. Our key observation is a common producer-consumer task execution pattern in mainstream vRAN implementations. We exploit this pattern to reclaim CPU cores from worker threads only at the boundary of processing different tasks. This guarantees run-to-completion task processing, which is critical for vRAN to achieve low latency and stability. Aquifer intercepts system calls invoked by vRAN at the OS layer to achieve transparent load monitoring and core reallocation. Aquifer employs a set of system-level optimizations on thread state detection, signal transmission and core selection, which reduces the scheduling cycle to 2$\mu s$. Experimental results show that Aquifer reclaims up to 88.31% of wasted CPU resources for two mainstream vRAN implementations, FlexRAN and OAI, without any source code modifications.
Yunshan Jia, Yinmin Zhong, Meng Wang 0018, Xuanzhe Liu, Xin Jin 0008
IEEE Trans. Serv. Comput.2
2023 ElasticFlow: An Elastic Serverless Training Platform for Distributed Deep Learning
abstract
This paper proposes ElasticFlow, an elastic serverless training platform for distributed deep learning. ElasticFlow provides a serverless interface with two distinct features: (i) users specify only the deep neural network (DNN) model and hyperparameters for a job, but not the number of GPUs; (ii) users specify the deadline for a job, but not the amount of time to occupy GPUs. In contrast to existing server-centric platforms, ElasticFlow provides performance guarantees in terms of meeting deadlines while alleviating tedious, low-level, and manual resource management for deep learning developers. The characteristics of distributed training introduce two challenges. First, the training throughput scales non-linearly with the number of GPUs. Second, the scaling efficiency is affected by worker placement. To address these challenges, we propose Minimum Satisfactory Share to capture the resource usage of training jobs to meet deadlines, and ElasticFlow performs admission control based on it. We develop a greedy algorithm that dynamically allocates resources to admitted jobs based on diminishing returns. We apply buddy allocation to worker placement to eliminate the effect of topology. Evaluation results on a cluster of 128 GPUs show that ElasticFlow increases the number of jobs that can meet their deadlines by 1.46–7.65× compared to existing solutions.
Diandian Gu, Yinmin Zhong, Yifan Xiong 0001, Zhenhua Han, Peng Cheng 0005, Fan Yang 0024, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu
ASPLOS (2)3
2023 AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving
Zhuohan Li 0001, Lianmin Zheng, Yinmin Zhong, Vincent Liu 0001, Ying Sheng 0007, Xin Jin 0008, Yanping Huang, Hao Zhang 0025, Joseph Gonzalez 0001, Ion Stoica
OSDI3