Remi Delacourt

dblp:397/9604 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0008-1319-4832ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 87% Hardware accelerators and domain-specific architectures · 13%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › inference serving
LLM serving
2.022026
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees · NSDI 2026
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding · EuroSys 2026
Machine learning and data management
inference serving
1.012026
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding · EuroSys 2026
Machine learning › Efficient and distributed learning
inference serving
0.312026
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees · NSDI 2026

Methods — techniques the papers use, named apart from their topics

token-level scheduling · 2.0speculative decoding · 2.0constrained optimization · 2.0
YearPublicationVenuePosition
2026 AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
abstract
Modern large language model (LLM) applications exhibit diverse service-level objectives (SLOs), from low-latency requirements in interactive coding assistants to more relaxed constraints in data wrangling tasks. Existing LLM serving systems, which rely on uniform batching and scheduling strategies, often fail to meet these heterogeneous SLOs concurrently. We present AdaServe, the first LLM serving system designed to support efficient multi-SLO serving through SLO-customized speculative decoding. AdaServe formulates multi-SLO serving as a constrained optimization problem and introduces a hardware-aware algorithm that constructs a speculation tree tailored to each request's latency target. It features a speculate-select-verify pipeline that enables fine-grained control over decoding speed while maximizing system throughput. AdaServe further adapts to workload variation by dynamically adjusting speculation parameters. Evaluations across diverse workloads show that AdaServe reduces SLO violations by up to 4.3X and improves goodput by up to 1.9X compared to the best-performing baselines, highlighting its effectiveness in multi-SLO serving.
Zikun Li, Zhuofu Chen, Remi Delacourt, Gabriele Oliaro, Qinghan Chen, Shuhuai Lin, April Yang, Zhihao Zhang 0001, Zhuoming Chen, Yi-Hsiang Lai, Xinhao Cheng, Xupeng Miao
EuroSys3
2026 FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
Gabriele Oliaro, Xupeng Miao, Xinhao Cheng, Vineeth Kada, Mengdi Wu, Ruohan Gao, Yingyi Huang, Remi Delacourt, April Yang, Yingcheng Wang, Colin Unger
NSDI8