EDBT 2026 Demo / reviewers in the wild / expert
Dong Liu 0049
dblp:98/1737-49
· DBLP profile ↗
3ranked-venue papers
3as first author
3since 2021 · last 2026
0009-0009-6815-8297ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 59% Hardware accelerators and domain-specific architectures · 33% Reconfigurable computing and FPGAs · 8% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › memory disaggregation
CXL memory disaggregation |
1.0 | 1 | 2026 | CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving · FPGA 2026 |
Memory systems
memory disaggregation |
1.0 | 1 | 2026 | CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving · FPGA 2026 |
Machine learning › Efficient and distributed learning
inference serving |
0.9 | 1 | 2025 | TinyServe: Query-Aware Cache Selection for Efficient LLM Serving · ACM Multimedia 2025 |
Machine learning › Efficient and distributed learning › inference serving
large language model serving |
0.9 | 1 | 2025 | TinyServe: Query-Aware Cache Selection for Efficient LLM Serving · ACM Multimedia 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | TinyServe: Query-Aware Cache Selection for Efficient LLM Serving · ACM Multimedia 2025 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.3 | 1 | 2026 | CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving · FPGA 2026 |
Memory systems
cache management |
0.3 | 1 | 2025 | TinyServe: Query-Aware Cache Selection for Efficient LLM Serving · ACM Multimedia 2025 |
Methods — techniques the papers use, named apart from their topics
query-aware page selection · 1.7fused CUDA kernel · 1.7bounding-box metadata · 1.7speculative prefetching · 1.0KV-cache compression · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MKA: Memory-Keyed Attention for Efficient Long-Context ReasoningabstractAs long-context language modeling becomes increasingly important, the cost of maintaining and attending to large Key/Value (KV) caches grows rapidly, becoming a major bottleneck in both training and inference. While prior works such as Multi-Query Attention (MQA) and Multi-Latent Attention (MLA) reduce memory by sharing or compressing KV features, they often trade off representation quality or incur runtime overhead. We propose Memory-Keyed Attention (MKA), a hierarchical attention mechanism that integrates multi-level KV caches—local, session, and long-term—and learns to route attention across them dynamically. We further introduce Route-Fused MKA (FastMKA), a broadcast-routed variant that fuses memory sources before attention computation for enhanced efficiency. Experiments on different sequence lengths show that FastMKA achieves a favorable accuracy-efficiency trade-off: comparable perplexity to MLA while achieving up to 5 × faster training throughput and 1.8 × lower evaluation latency. These results highlight MKA as a practical and extensible framework for efficient long-context attention. Dong Liu 0049, Yanxuan Yu, Ben Lengerich, Ying Nian Wu |
CF | 1 |
| 2026 | CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM ServingabstractLarge Language Models (LLMs) have revolutionized natural language processing tasks, but their deployment in datacenter environments faces significant challenges due to the massive memory requirements of key-value (KV) caches. During the autoregressive decoding process, KV caches consume substantial GPU memory, limiting batch sizes and overall system throughput. To address these challenges, we propose CXL-SpecKV, a novel disaggregated KV-cache architecture that leverages Compute Express Link (CXL) interconnects and FPGA accelerators to enable efficient speculative execution and memory disaggregation. Our approach introduces three key innovations: (i) a CXL-based memory disaggregation framework that offloads KV-caches to remote FPGA memory with low latency, (ii) a speculative KV-cache prefetching mechanism that predicts and preloads future tokens' cache entries, and (iii) an FPGA-accelerated KV-cache compression and decompression engine that reduces memory bandwidth requirements by up to 4×. When evaluated on state-of-the-art LLM models, CXL-SpecKV achieves up to 3.2× higher throughput compared to GPU-only baselines, while reducing memory costs by 2.8× and maintaining accuracy. Our system demonstrates that intelligent memory disaggregation combined with speculative execution can effectively address the memory wall challenge in large-scale LLM serving. Dong Liu 0049, Yanxuan Yu |
FPGA | 1 |
| 2025 | TinyServe: Query-Aware Cache Selection for Efficient LLM ServingabstractServing large language models (LLMs) efficiently remains challenging due to the high memory and latency overhead of key-value (KV) cache access during autoregressive decoding. We present TinyServe, a lightweight and extensible serving system for deploying tiny LLMs (e.g., TinyLLaMA, GPT2-345M) with support for structured KV sparsity, plugin-based token selection, and hardware-efficient attention kernels. Unlike prior simulation frameworks, TinyServe executes real-time decoding with configurable sparsity strategies and fine-grained instrumentation. To reduce decoding cost, we introduce a query-aware page selection mechanism that leverages bounding-box metadata to estimate attention relevance between the query and KV cache blocks. This enables selective KV loading with minimal overhead and no model modifications. Our fused CUDA kernel integrates page scoring, sparse memory access, and masked attention in a single pass. Experiments show that TinyServe achieves up to 3.4× speedup and over 2× memory savings with negligible accuracy drop. Additional analysis of cache reuse, page hit rate, and multi-GPU scaling confirms its practicality as an efficient system-level design for LLM training and inference research on resource-constrained hardware. Dong Liu 0049, Yanxuan Yu |
ACM Multimedia | 1 |