VLDB 2026 Research / reviewers in the wild / expert
Yinmin Zhong
dblp:339/0691
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-2504-7652ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FastServe: Iteration-Level Preemptive Scheduling for Large Language Model Inference
Bingyang Wu, Yinmin Zhong, Fangyue Liu, Yuanhang Sun, Gang Huang 0001, Xuanzhe Liu, Xin Jin 0008 |
NSDI | 2 |
| 2026 | DistRS: Disaggregated Reward Service for RLVR with Batch-Level Constraint
Ruidong Zhu, Mingcong Han, Yinmin Zhong, Wencong Xiao, Xuanzhe Liu, Xin Jin 0008 |
NSDI | 3 |
| 2026 | DualPath: Accelerating Agentic LLM Inference by Harvesting Disaggregated KV-Cache Storage I/OabstractThe performance of multi-turn, agentic LLM inference is increasingly dominated by KV-Cache storage I/O rather than computation. In prevalent disaggregated architectures, loading the massive KV-Cache from external storage creates a fundamental imbalance: storage NICs on prefill engines become bandwidth-saturated, while those on decoding engines remain idle. This asymmetry severely constrains overall system throughput. Yongtong Wu, Shaoyuan Chen, Rilin Huang, Yixuan Tan, Yinmin Zhong, Xin Jin 0008, Panpan Huang |
SIGCOMM | 5 |
| 2025 | Optimizing RLHF Training for Large Language Models with Stage Fusion
Yinmin Zhong, Bingyang Wu, Changyi Wan, Hanpeng Hu, Ranchen Ming, Yibo Zhu 0001, Xin Jin 0008 |
NSDI | 1 |
| 2025 | DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language ModelsabstractMultimodal large language models (LLMs) empower LLMs to ingest inputs and generate outputs in multiple forms, such as text, image, and audio. However, the integration of multiple modalities introduces heterogeneity in both the model and training data, creating unique systems challenges. Yinmin Zhong, Hanpeng Hu, Jianjian Sun, Zheng Ge, Yibo Zhu 0001, Daxin Jiang, Xin Jin 0008 |
SIGCOMM | 2 |
| 2024 | MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUs
Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang 0001, Yangrui Chen, Zhi Zhang 0005, Yanghua Peng, Xiang Li 0067, Shibiao Nong, Yulu Jia, Sun He, Hongmin Chen, Zhihao Bai, Qi Hou, Shipeng Yan, Yiyao Sheng, Zhuo Jiang, Haohan Xu, Zhang Zhang 0003, Pengfei Nie, Leqi Zou, Sida Zhao, Zherui Liu, Xiaoying Jia 0001, Jianxi Ye, Xin Jin 0008, Xin Liu 0086 |
NSDI | 3 |
| 2024 | DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
Yinmin Zhong, Junda Chen, Jianbo Hu, Yibo Zhu 0001, Xuanzhe Liu, Xin Jin 0008, Hao Zhang 0025 |
OSDI | 1 |
| 2024 | LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence ParallelismabstractThe context window of large language models (LLMs) is rapidly increasing, leading to a huge variance in resource usage between different requests as well as between different phases of the same request. Restricted by static parallelism strategies, existing LLM serving systems cannot efficiently utilize the underlying resources to serve variable-length requests in different phases. To address this problem, we propose a new parallelism paradigm, elastic sequence parallelism (ESP), to elastically adapt to the variance across different requests and phases. Based on ESP, we design and build LoongServe, an LLM serving system that (1) improves computation efficiency by elastically adjusting the degree of parallelism in real-time, (2) improves communication efficiency by reducing key-value cache migration overhead and overlapping partial decoding communication with computation, and (3) improves GPU memory efficiency by reducing key-value cache fragmentation across instances. Our evaluation under diverse real-world datasets shows that LoongServe improves the throughput by up to 3.85× compared to chunked prefill and 5.81× compared to prefill-decoding disaggregation. Bingyang Wu, Yinmin Zhong, Peng Sun 0006, Xuanzhe Liu, Xin Jin 0008 |
SOSP | 3 |
| 2024 | DistMind: Efficient Resource Disaggregation for Deep Learning WorkloadsabstractDeep learning (DL) systems suffer from low resource utilization due to 1) monolithic server model that tightly couples compute and memory; and 2) limited sharing between different inference applications, and across inference and training, because of strict service level objectives (SLOs). To address this problem, we present, a disaggregated DL system that enables efficient multiplexing of DL applications with near-optimal resource utilization. decouples compute from host memory, and exposes the abstractions of a GPU pool and a memory pool, each of which can be independently provisioned. The key challenge is to dynamically allocate GPU resources to different applications based on their real-time demands while meeting strict SLOs. We tackle this challenge by exploiting the power of high-speed 100 Gbps networks, and design three-stage pipelining, cache-aware load balancing, and DNN-aware sharding mechanisms based on the characteristics of DL workloads, to achieve millisecond-scale application loading overhead and improve system efficiency. We have implemented a prototype of and integrated it with PyTorch. Experimental results on AWS EC2 show that achieves near 100% resource utilization, and compared with NVIDIA MPS and Ray, improves the throughput by up to 279% and reduces the inference latency by up to 94%. Xin Jin 0008, Zhihao Bai, Zhen Zhang 0063, Yibo Zhu 0001, Yinmin Zhong, Xuanzhe Liu |
IEEE/ACM Trans. Netw. | 5 |
| 2024 | Aquifer: Transparent Microsecond-Scale Scheduling for vRAN WorkloadsabstractVirtual Radio Access Network (vRAN) is an emerging approach offered by cloud providers to accelerate 5G services deployment. Despite significant microsecond-scale traffic variations, vRAN instances are provisioned based on peak load to meet strict latency requirements, leading to significant resource waste. Conceivably, vRAN can share CPUs with other applications to increase CPU utilization. Yet, existing sharing solutions require modifications to vRAN source code, hindering their deployment on public clouds. We present Aquifer, a microsecond-scale scheduler providing transparent CPU sharing for vRAN workloads. Our key observation is a common producer-consumer task execution pattern in mainstream vRAN implementations. We exploit this pattern to reclaim CPU cores from worker threads only at the boundary of processing different tasks. This guarantees run-to-completion task processing, which is critical for vRAN to achieve low latency and stability. Aquifer intercepts system calls invoked by vRAN at the OS layer to achieve transparent load monitoring and core reallocation. Aquifer employs a set of system-level optimizations on thread state detection, signal transmission and core selection, which reduces the scheduling cycle to 2$\mu s$. Experimental results show that Aquifer reclaims up to 88.31% of wasted CPU resources for two mainstream vRAN implementations, FlexRAN and OAI, without any source code modifications. Yunshan Jia, Yinmin Zhong, Meng Wang 0018, Xuanzhe Liu, Xin Jin 0008 |
IEEE Trans. Serv. Comput. | 2 |
| 2023 | ElasticFlow: An Elastic Serverless Training Platform for Distributed Deep LearningabstractThis paper proposes ElasticFlow, an elastic serverless training platform for distributed deep learning. ElasticFlow provides a serverless interface with two distinct features: (i) users specify only the deep neural network (DNN) model and hyperparameters for a job, but not the number of GPUs; (ii) users specify the deadline for a job, but not the amount of time to occupy GPUs. In contrast to existing server-centric platforms, ElasticFlow provides performance guarantees in terms of meeting deadlines while alleviating tedious, low-level, and manual resource management for deep learning developers. The characteristics of distributed training introduce two challenges. First, the training throughput scales non-linearly with the number of GPUs. Second, the scaling efficiency is affected by worker placement. To address these challenges, we propose Minimum Satisfactory Share to capture the resource usage of training jobs to meet deadlines, and ElasticFlow performs admission control based on it. We develop a greedy algorithm that dynamically allocates resources to admitted jobs based on diminishing returns. We apply buddy allocation to worker placement to eliminate the effect of topology. Evaluation results on a cluster of 128 GPUs show that ElasticFlow increases the number of jobs that can meet their deadlines by 1.46–7.65× compared to existing solutions. Diandian Gu, Yinmin Zhong, Yifan Xiong 0001, Zhenhua Han, Peng Cheng 0005, Fan Yang 0024, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu |
ASPLOS (2) | 3 |
| 2023 | AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving
Zhuohan Li 0001, Lianmin Zheng, Yinmin Zhong, Vincent Liu 0001, Ying Sheng 0007, Xin Jin 0008, Yanping Huang, Hao Zhang 0025, Joseph Gonzalez 0001, Ion Stoica |
OSDI | 3 |