Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hao Wu 0032

dblp:72/4250-32 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0000-0003-2570-4648ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 33% GPUs and heterogeneous computing · 33% Distributed systems · 33%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
data transmission
1.012026
Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric Approach · EuroSys 2026
GPUs and heterogeneous computing
GPU communication
1.012026
Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric Approach · EuroSys 2026
Cloud and datacenter computing
serverless computing
1.012026
Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric Approach · EuroSys 2026
Machine learning › Efficient and distributed learning
inference serving
0.312026
Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric Approach · EuroSys 2026

Methods — techniques the papers use, named apart from their topics

NVSHMEM · 2.0NCCL · 2.0GPU-centric data passing · 2.0
YearPublicationVenuePosition
2026 Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric Approach
abstract
Serverless computing offers a compelling paradigm for deploying machine learning inference workflows composed of heterogeneous CPU and GPU functions. However, existing data-passing solutions in serverless systems primarily rely on host memory for data exchange (host-centric), leading to substantial data movement and salient I/O overhead. Moreover, modern GPU communication libraries (e.g., NCCL, NVSHMEM, UCX) are ill-suited to serverless environments, suffering from redundant data copies, underutilized transfer bandwidth, and inefficient temporary GPU storage.
Hao Wu 0032, Yaochen Liu, Minchen Yu, Qizhen Weng 0001, Junxiao Deng, Hao Fan 0006, Song Wu 0001, Wei Wang 0030, Hai Jin 0001
EuroSys1
2025 AdaSpec: Adaptive Speculative Decoding for Fast, SLO-Aware Large Language Model Serving
abstract
Cloud-based Large Language Model (LLM) services often face challenges in achieving low inference latency and meeting Service Level Objectives (SLOs) under dynamic request patterns. Speculative decoding, which exploits lightweight models for drafting and LLMs for verification, has emerged as a compelling technique to accelerate LLM inference. However, existing speculative decoding solutions often fail to adapt to fluctuating workloads and dynamic system environments, resulting in impaired performance and SLO violations. In this paper, we introduce AdaSpec, an efficient LLM inference system that dynamically adjusts speculative strategies according to real-time request loads and system configurations. AdaSpec proposes a theoretical model to analyze and predict the efficiency of speculative strategies across diverse scenarios. Additionally, it implements intelligent drafting and verification algorithms to maximize performance while ensuring high SLO attainment. Experimental results on real-world LLM service traces demonstrate that AdaSpec consistently meets SLOs and achieves substantial performance improvements, delivering up to 66% speedup compared to state-of-the-art speculative inference systems. The source code is publicly available at https://github.com/cerebellumking/AdaSpec
Hao Wu 0032, Zhubo Shi, Han Zou, Minchen Yu, Qingjiang Shi
SoCC2