Haoxuan Yu

dblp:310/2360 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Compass: Dissecting Communication and Computation Operators for Efficient LLM Training
abstract
Overlapping communication and computation operators is a common practice to hide communication overheads, accelerating large language models (LLMs) training on GPU clusters. Existing systems achieve this through either intra-operator fusion (IntraFusion), which packs operators into a single large kernel, or inter-operator decomposition (InterDecom), which splits a tensor into multiple parts for pipelined execution. However, current IntraFusion methods underutilize network topology, causing suboptimal bandwidth usage on multi-GPU systems, while InterDecom struggles to determine the optimal number of decomposed parts for peak performance. To address these issues, we introduce Compass, which employs systematic optimization and comprehensive modeling. First, we design a novel IntraFusion algorithm leveraging double-ring communications to maximize bandwidth utilization in hybrid NVLink-PCIe systems, achieving 1.5x-2.5x speedups. Second, we develop a decomposition model that mathematically derives the optimal tensor decomposition degree for InterDecom, improving performance by up to 1.3x. Finally, we develop a unified performance framework that accurately determines the best strategy for different scenarios. We validate Compass through extensive evaluation across 288 configurations and end-to-end experiments on real-world applications. The results demonstrate that Compass consistently selects the optimal strategy, achieving up to a 1.42x end-to-end speedup compared to the Megatron-LM baseline.
Guangyu Xiang, Lin Zhang 0059, Haoxuan Yu, Xinglin Pan, Shaohuai Shi, Xiaowen Chu 0001
INFOCOM3
2026 Enabling Low-Latency, GPU-Efficient Serverless Inference with Model Swapping
abstract
Serverless computing offers a compelling cloud model for online inference services. However, existing serverless platforms lack efficient support for GPUs, hindering their ability to deliver high-performance inference. In this article, we present Torpor , a serverless platform for GPU-efficient, low-latency inference. To enable efficient sharing of a node’s GPUs among numerous inference functions, Torpor maintains models in main memory and dynamically swaps them onto GPUs upon request arrivals (i.e., late binding with model swapping). Torpor uses various techniques, including asynchronous API redirection, GPU runtime sharing, pipelined model execution, and efficient GPU memory management, to minimize latency overhead caused by model swapping. Additionally, we design an interference-aware request scheduling algorithm that utilizes high-speed GPU interconnects to meet latency service-level objectives (SLOs) for individual inference functions. We have implemented Torpor and evaluated its performance in a production environment. Utilizing late binding and model swapping, Torpor can concurrently serve hundreds of inference functions on a worker node with 4 GPUs, while achieving latency performance comparable to native execution, where each model is cached exclusively on a GPU. Pilot deployment in a leading commercial serverless cloud shows that Torpor reduces the GPU provisioning cost by 70% and 65% for users and the platform, respectively.
Minchen Yu, Bohui Wu, Haoxuan Yu, Wei Wang 0030, Ruichuan Chen, Dapeng Nie
ACM Trans. Archit. Code Optim.6
2025 ZipBatch: Multi-Tenant GPU Batching with Dual-Resource Regulation
abstract
GPU multiplexing is a widely-adopted strategy in GPU clusters for improving overall throughput and lowering the total cost of ownership. To mitigate inter-task interference in compute power and memory bandwidth on multiplexed GPUs, existing techniques divide a GPU into instances with limited predefined rigid configurations. Low utilization arises from the mismatch between heterogeneous burstiness and immutable resource configurations: 1) bursty inference traffic forces the scheduler to launch underfilled batches that cannot saturate the instance; 2) bursty kernel resource utilization leads to bubbles in compute power and memory bandwidth.
Haoxuan Yu, Sheng Yao 0006, Wei Wang 0030
SoCC1
2025 SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUs
abstract
Cloud service providers heavily colocate high-priority, latency sensitive (LS), and low-priority, best-effort (BE) DNN inference services on the same GPU to improve resource utilization in data centers. Among the critical shared GPU resources, there has been very limited analysis on the dynamic allocation of compute units and VRAM bandwidth, mainly for two reasons: (1) The native GPU resource management solutions are either hardware-specific, or unable to dynamically allocate resources to different tenants, or both; (2) NVIDIA doesn't expose interfaces for VRAM bandwidth allocation, and the software stack and VRAM channel architectures are black-box, both of which limit the software-level resource management. These drive prior work to design either conservative sharing policies detrimental to throughput, or static resource partitioning only applicable to a few GPU models.
Yongkang Zhang 0003, Haoxuan Yu, Chenxia Han, Baotong Lu, Zhifeng Jiang 0001, Yang Li 0090, Xiaowen Chu 0001, Huaicheng Li
PPoPP2
2025 Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
Minchen Yu, Haoxuan Yu, Zhuohao Li, Wei Wang 0030, Ruichuan Chen, Dapeng Nie
USENIX ATC4