Zhuofu Chen

dblp:367/4172 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0004-1735-4443ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 92% Hardware accelerators and domain-specific architectures · 8%
Artificial intelligence
1 paper
Efficient and distributed learning · 50% Deep learning architectures and training · 38% Language models and text generation · 12%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%
Software engineering, system software, and programming languages
1 paper
Programming languages and type systems · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning and data management
inference serving
1.012026
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding · EuroSys 2026
Cloud and datacenter computing › inference serving
LLM serving
1.012026
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding · EuroSys 2026
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention · ICLR 2025
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention
0.912025
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention · ICLR 2025
Cloud and datacenter computing
serverless computing
0.912025
Featherlight Stateful WebAssembly for Serverless Inference Workflows · IEEE Trans. Parallel Distributed Syst. 2025
Cloud and datacenter computing › serverless computing
serverless inference
0.912025
Featherlight Stateful WebAssembly for Serverless Inference Workflows · IEEE Trans. Parallel Distributed Syst. 2025
Cloud and datacenter computing
workflow scheduling
0.912025
Featherlight Stateful WebAssembly for Serverless Inference Workflows · IEEE Trans. Parallel Distributed Syst. 2025
Machine learning › Efficient and distributed learning
KV cache management
0.312025
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention · ICLR 2025
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context language model
0.312025
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention · ICLR 2025
Programming languages and type systems
webassembly
0.312025
Featherlight Stateful WebAssembly for Serverless Inference Workflows · IEEE Trans. Parallel Distributed Syst. 2025

Methods — techniques the papers use, named apart from their topics

speculative decoding · 2.0constrained optimization · 2.0process-level virtualization · 1.7lock-free zero-copy communication · 1.7affinity-aware scheduling · 1.7token selection · 0.9sparse attention · 0.9
YearPublicationVenuePosition
2026 AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
abstract
Modern large language model (LLM) applications exhibit diverse service-level objectives (SLOs), from low-latency requirements in interactive coding assistants to more relaxed constraints in data wrangling tasks. Existing LLM serving systems, which rely on uniform batching and scheduling strategies, often fail to meet these heterogeneous SLOs concurrently. We present AdaServe, the first LLM serving system designed to support efficient multi-SLO serving through SLO-customized speculative decoding. AdaServe formulates multi-SLO serving as a constrained optimization problem and introduces a hardware-aware algorithm that constructs a speculation tree tailored to each request's latency target. It features a speculate-select-verify pipeline that enables fine-grained control over decoding speed while maximizing system throughput. AdaServe further adapts to workload variation by dynamically adjusting speculation parameters. Evaluations across diverse workloads show that AdaServe reduces SLO violations by up to 4.3X and improves goodput by up to 1.9X compared to the best-performing baselines, highlighting its effectiveness in multi-SLO serving.
Zikun Li, Zhuofu Chen, Remi Delacourt, Gabriele Oliaro, Qinghan Chen, Shuhuai Lin, April Yang, Zhihao Zhang 0001, Zhuoming Chen, Yi-Hsiang Lai, Xinhao Cheng, Xupeng Miao
EuroSys2
2025 TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
abstract
Large language models (LLMs) have driven significant advancements across diverse NLP tasks, with long-context models gaining prominence for handling extended inputs. However, the expanding key-value (KV) cache size required by Transformer architectures intensifies the memory constraints, particularly during the decoding phase, creating a significant bottleneck. Existing sparse attention mechanisms designed to address this bottleneck have two limitations: (1) they often fail to reliably identify the most relevant tokens for attention, and (2) they overlook the spatial coherence of token selection across consecutive Transformer layers, which can lead to performance degradation and substantial overhead in token selection. This paper introduces TidalDecode, a simple yet effective algorithm and system for fast and accurate LLM decoding through position persistent sparse attention. TidalDecode leverages the spatial coherence of tokens selected by existing sparse attention methods and introduces a few token selection layers that perform full attention to identify the tokens with the highest attention scores, while all other layers perform sparse attention with the pre-selected tokens. This design enables TidalDecode to substantially reduce the overhead of token selection for sparse attention without sacrificing the quality of the generated results. Evaluation on a diverse set of LLMs and tasks shows that TidalDecode closely matches the generative performance of full attention methods while reducing the LLM decoding latency by up to $2.1\times$.
Lijie Yang 0003, Zhihao Zhang 0001, Zhuofu Chen, Zikun Li
ICLR3
2025 Featherlight Stateful WebAssembly for Serverless Inference Workflows
abstract
In serverless inference, complex prediction tasks are executed as workflows, relying on efficient state transfer across multiple functions. Serverless platforms typically deploy each function in a separate stateless container, depending on external processes for state management, which often results in suboptimal system utilization and increased latency. We introduce WasmFlow, a novel framework designed for serverless inference that ensures low latency and high throughput. This is achieved through process-level virtualization using WebAssembly. WasmFlow operates functions on a per-thread basis within compact WebAssembly modules, significantly reducing startup times and memory usage. The framework has two key features. (1) Efficient Memory Sharing: WasmFlow facilitates direct and rapid state transfer between functions using threads within the WebAssembly runtime. This is enabled through lightweight, lock-free, zero-copy intra-process communication, complemented by effective inter-process RPC. (2) System Optimizations: We further optimize WasmFlow with an advanced synchronization technique between functions, an affinity-aware workflow scheduler, and adaptive request batching. Implemented and integrated within the Kubernetes ecosystem, WasmFlow's performance was evaluated using synthetic workloads and realworld Azure traces, including typical serverless workflows and ML models. Our results demonstrate that WasmFlow dramatically outperforms existing serverless frameworks. It reduces P90 end-to-end latency by 74x and 78x, increases function density by n1.7x and 223x compared to Faasm and SPRIGHT, and improves system throughput by 12.3x and 8.8x over Knative and WasmEdge, respectively.
Xingguo Pang, Yanze Zhang, Zhuofu Chen, Zhijun Ding, Dazhao Cheng, Xiaobo Zhou 0002
IEEE Trans. Parallel Distributed Syst.4