EDBT 2026 Demo / reviewers in the wild / expert
Xavier Servot
dblp:399/8477
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0003-3830-7219ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 56% Memory systems · 44% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM inference accelerator |
0.9 | 1 | 2025 | PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference · ASPLOS (2) 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 1 | 2025 | PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference · ASPLOS (2) 2025 |
Memory systems
processing-in-memory |
0.9 | 1 | 2025 | PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference · ASPLOS (2) 2025 |
Memory systems › cache
key-value cache |
0.3 | 1 | 2025 | PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference · ASPLOS (2) 2025 |
Memory systems
memory bandwidth |
0.3 | 1 | 2025 | PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference · ASPLOS (2) 2025 |
Methods — techniques the papers use, named apart from their topics
CXL · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model InferenceabstractLarge Language Model (LLM) inference uses an autoregressive manner to generate one token at a time, which exhibits notably lower operational intensity compared to earlier Machine Learning (ML) models such as encoder-only transformers and Convolutional Neural Networks. At the same time, LLMs possess large parameter sizes and use key-value caches to store context information. Modern LLMs support context windows with up to 1 million tokens to generate versatile text, audio, and video content. A large key-value cache unique to each prompt requires a large memory capacity, limiting the inference batch size. Both low operational intensity and limited batch size necessitate a high memory bandwidth. However, contemporary hardware systems for ML model deployment, such as GPUs and TPUs, are primarily optimized for compute throughput. This mismatch challenges the efficient deployment of advanced LLMs and makes users to pay for expensive compute resources that are poorly utilized for the memory-bound LLM inference tasks. Yufeng Gu, Alireza Khadem, Sumanth Umesh, Xavier Servot, Onur Mutlu, Ravi R. Iyer 0001, Reetuparna Das |
ASPLOS (2) | 5 |