Minchen Yu

dblp:227/0752 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-6797-9028ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 8 since 2021Computer networks · 4 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric Approach
abstract
Serverless computing offers a compelling paradigm for deploying machine learning inference workflows composed of heterogeneous CPU and GPU functions. However, existing data-passing solutions in serverless systems primarily rely on host memory for data exchange (host-centric), leading to substantial data movement and salient I/O overhead. Moreover, modern GPU communication libraries (e.g., NCCL, NVSHMEM, UCX) are ill-suited to serverless environments, suffering from redundant data copies, underutilized transfer bandwidth, and inefficient temporary GPU storage.
Hao Wu 0032, Yaochen Liu, Minchen Yu, Qizhen Weng 0001, Junxiao Deng, Hao Fan 0006, Song Wu 0001, Wei Wang 0030, Hai Jin 0001
EuroSys3
2026 ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
Xin Tan 0004, Minchen Yu, Jingzong Li, Hong Xu 0001
IWQoS3
2026 Enabling Low-Latency, GPU-Efficient Serverless Inference with Model Swapping
abstract
Serverless computing offers a compelling cloud model for online inference services. However, existing serverless platforms lack efficient support for GPUs, hindering their ability to deliver high-performance inference. In this article, we present Torpor , a serverless platform for GPU-efficient, low-latency inference. To enable efficient sharing of a node’s GPUs among numerous inference functions, Torpor maintains models in main memory and dynamically swaps them onto GPUs upon request arrivals (i.e., late binding with model swapping). Torpor uses various techniques, including asynchronous API redirection, GPU runtime sharing, pipelined model execution, and efficient GPU memory management, to minimize latency overhead caused by model swapping. Additionally, we design an interference-aware request scheduling algorithm that utilizes high-speed GPU interconnects to meet latency service-level objectives (SLOs) for individual inference functions. We have implemented Torpor and evaluated its performance in a production environment. Utilizing late binding and model swapping, Torpor can concurrently serve hundreds of inference functions on a worker node with 4 GPUs, while achieving latency performance comparable to native execution, where each model is cached exclusively on a GPU. Pilot deployment in a leading commercial serverless cloud shows that Torpor reduces the GPU provisioning cost by 70% and 65% for users and the platform, respectively.
Minchen Yu, Bohui Wu, Haoxuan Yu, Wei Wang 0030, Ruichuan Chen, Dapeng Nie
ACM Trans. Archit. Code Optim.1
2025 AdaSpec: Adaptive Speculative Decoding for Fast, SLO-Aware Large Language Model Serving
abstract
Cloud-based Large Language Model (LLM) services often face challenges in achieving low inference latency and meeting Service Level Objectives (SLOs) under dynamic request patterns. Speculative decoding, which exploits lightweight models for drafting and LLMs for verification, has emerged as a compelling technique to accelerate LLM inference. However, existing speculative decoding solutions often fail to adapt to fluctuating workloads and dynamic system environments, resulting in impaired performance and SLO violations. In this paper, we introduce AdaSpec, an efficient LLM inference system that dynamically adjusts speculative strategies according to real-time request loads and system configurations. AdaSpec proposes a theoretical model to analyze and predict the efficiency of speculative strategies across diverse scenarios. Additionally, it implements intelligent drafting and verification algorithms to maximize performance while ensuring high SLO attainment. Experimental results on real-world LLM service traces demonstrate that AdaSpec consistently meets SLOs and achieves substantial performance improvements, delivering up to 66% speedup compared to state-of-the-art speculative inference systems. The source code is publicly available at https://github.com/cerebellumking/AdaSpec
Hao Wu 0032, Zhubo Shi, Han Zou, Minchen Yu, Qingjiang Shi
SoCC5
2025 Toppings: CPU-Assisted, Rank-Aware Adapter Serving for LLM Inference
Suyi Li 0002, Hanfeng Lu, Tianyuan Wu, Minchen Yu, Qizhen Weng 0001, Xusheng Chen, Yizhou Shan, Binhang Yuan, Wei Wang 0030
USENIX ATC4
2025 Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
Minchen Yu, Haoxuan Yu, Zhuohao Li, Wei Wang 0030, Ruichuan Chen, Dapeng Nie
USENIX ATC1
2025 Pheromone: Restructuring Serverless Computing With Data-Centric Function Orchestration
abstract
Serverless applications are typically composed of function workflows in which multiple short-lived functions are triggered to exchange data in response to events or state changes. Current serverless platforms coordinate and trigger functions by following high-level invocation dependencies but are oblivious to the underlying data exchanges between functions. This design is neither efficient nor easy to use in orchestrating complex workflows – developers often have to manage complex function interactions by themselves, with customized implementation and unsatisfactory performance. Therefore, we argue that function orchestration should follow a data-centric approach. In our design, the platform provides a data bucket abstraction to hold the intermediate data generated by functions. Developers can use a rich set of data trigger primitives to control when and how the output of each function should be passed to the next functions in a workflow. By making data consumption explicit and allowing it to trigger functions and drive the workflow, complex function interactions can be easily and efficiently supported. We presentPheromone– a scalable, low-latency serverless platform following this data-centric design. Compared to well-established commercial and open-source platforms,Pheromonecuts the latencies of function interactions and data exchanges by orders of magnitude, scales to large workflows, and enables easy implementation of complex applications.
Minchen Yu, Tingjia Cao, Wei Wang 0030, Ruichuan Chen
IEEE Trans. Netw.1
2023 Following the Data, Not the Function: Rethinking Function Orchestration in Serverless Computing
Minchen Yu, Tingjia Cao, Wei Wang 0030, Ruichuan Chen
NSDI1
2022 Enabling Cost-Effective, SLO-Aware Machine Learning Inference Serving on Public Cloud
abstract
The remarkable advances of Machine Learning (ML) have spurred an increasing demand for ML- as-a-Service on public cloud: developers train and publish ML models as online services to provide low-latency inference for dynamic queries. The primary challenge of ML model serving is to meet the response-time Service-Level Objectives (SLOs) of inference workloads while minimizing serving cost. In this article, we proposes MArk (Model Ark), a general-purpose inference serving system, to tackle the dual challenge of SLO compliance and cost effectiveness. MArk employs three design choices tailored to inference workload. First, MArk dynamically batches requests and opportunistically serves them using expensive hardware accelerators (e.g., GPU) for improved performance-cost ratio. Second, instead of relying on feedback control scaling or over-provisioning to serve dynamic workload, which can be too slow or too expensive, MArk employs predictive autoscaling to hide the provisioning latency at low cost. Third, given the stateless nature of inference serving, MArk exploits the flexible, yet costly serverless instances to cover occasional load spikes that are hard to predict. We evaluated the performance of MArk using several state-of-the-art ML models trained in TensorFlow, MXNet, and Keras. Compared with the premier industrial ML serving platform SageMaker, MArk reduces the serving cost up to$7.8\times$while achieving even better latency performance.
Chengliang Zhang, Minchen Yu, Wei Wang 0030, Feng Yan 0001
IEEE Trans. Cloud Comput.2
2021 Gillis: Serving Large Neural Networks in Serverless Functions with Automatic Model Partitioning
abstract
The increased use of deep neural networks has stimulated the growing demand for cloud-based model serving platforms. Serverless computing offers a simplified solution: users deploy models as serverless functions and let the platform handle provisioning and scaling. However, serverless functions have constrained resources in CPU and memory, making them inefficient or infeasible to serve large neural networks-which have become increasingly popular. In this paper, we present Gillis, a serverless-based model serving system that automatically partitions a large model across multiple serverless functions for faster inference and reduced memory footprint per function. Gillis employs two novel model partitioning algorithms that respectively achieve latency-optimal serving and cost-optimal serving with SLO compliance. We have implemented Gillis on three serverless platforms-AWS Lambda, Google Cloud Functions, and KNIX-with MXNet as the serving backend. Experimental evaluations against popular models show that Gillis supports serving very large neural networks, reduces the inference latency substantially, and meets various SLOs with a low serving cost.
Minchen Yu, Zhifeng Jiang 0001, Hok Chun Ng, Wei Wang 0030, Ruichuan Chen, Bo Li 0001
ICDCS1
2021 CrystalPerf: Learning to Characterize the Performance of Dataflow Computation through Code Analysis
Huangshi Tian, Minchen Yu, Wei Wang 0030
USENIX ATC2
2020 RepBun: Load-Balanced, Shuffle-Free Cluster Caching for Structured Data
abstract
Cluster caching systems increasingly store structured data objects in the columnar format. However, these systems routinely face the imbalanced load that significantly impairs the I/O performance. Existing load-balancing solutions, while effective for reading unstructured data objects, fall short in handling columnar data. Unlike unstructured data that can only be read through a full-object scan, columnar data supports direct query of specific columns with two distinct access patterns: (1) columns have the heavily skewed popularity, and (2) hot columns are likely accessed together in a query job. Based on these two access patterns, we propose an effective load-balancing solution for structured data. Our solution, which we call RepBun, groups hot columns into a bundle. It then copies multiple replicas of the column bundle and stores them uniformly across servers. We show that RepBun achieves improved load balancing with reduced memory overhead, while avoiding data shuffling between cache servers. We implemented RepBun atop Alluxio, a popular in-memory distributed storage, and evaluate its performance through EC2 deployment against the TPC-H benchmark work-load. Experimental results show that RepBun outperforms the existing load-balancing solutions with significantly shorter read latency and faster query completion.
Minchen Yu, Yinghao Yu, Yunchuan Zheng, Baichen Yang, Wei Wang 0030
INFOCOM1
2019 MArk: Exploiting Cloud Services for Cost-Effective, SLO-Aware Machine Learning Inference Serving
Chengliang Zhang, Minchen Yu, Wei Wang 0030, Feng Yan 0001
USENIX ATC2
2018 Continuum: A Platform for Cost-Aware, Low-Latency Continual Learning
abstract
Many machine learning applications operate in dynamic environments that change over time, in which models must be continually updated to capture the recent trend in data. However, most of today's learning frameworks perform training offline, without a system support for continual model updating.
Huangshi Tian, Minchen Yu, Wei Wang 0030
SoCC2