Dong Liu 0049

dblp:98/1737-49 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2026
0009-0009-6815-8297ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 59% Hardware accelerators and domain-specific architectures · 33% Reconfigurable computing and FPGAs · 8%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory disaggregation
CXL memory disaggregation
1.012026
CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving · FPGA 2026
Memory systems
memory disaggregation
1.012026
CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving · FPGA 2026
Machine learning › Efficient and distributed learning
inference serving
0.912025
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving · ACM Multimedia 2025
Machine learning › Efficient and distributed learning › inference serving
large language model serving
0.912025
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving · ACM Multimedia 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving · ACM Multimedia 2025
Reconfigurable computing and FPGAs
FPGA accelerator
0.312026
CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving · FPGA 2026
Memory systems
cache management
0.312025
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

query-aware page selection · 1.7fused CUDA kernel · 1.7bounding-box metadata · 1.7speculative prefetching · 1.0KV-cache compression · 1.0
YearPublicationVenuePosition
2026 MKA: Memory-Keyed Attention for Efficient Long-Context Reasoning
abstract
As long-context language modeling becomes increasingly important, the cost of maintaining and attending to large Key/Value (KV) caches grows rapidly, becoming a major bottleneck in both training and inference. While prior works such as Multi-Query Attention (MQA) and Multi-Latent Attention (MLA) reduce memory by sharing or compressing KV features, they often trade off representation quality or incur runtime overhead. We propose Memory-Keyed Attention (MKA), a hierarchical attention mechanism that integrates multi-level KV caches—local, session, and long-term—and learns to route attention across them dynamically. We further introduce Route-Fused MKA (FastMKA), a broadcast-routed variant that fuses memory sources before attention computation for enhanced efficiency. Experiments on different sequence lengths show that FastMKA achieves a favorable accuracy-efficiency trade-off: comparable perplexity to MLA while achieving up to 5 × faster training throughput and 1.8 × lower evaluation latency. These results highlight MKA as a practical and extensible framework for efficient long-context attention.
Dong Liu 0049, Yanxuan Yu, Ben Lengerich, Ying Nian Wu
CF1
2026 CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving
abstract
Large Language Models (LLMs) have revolutionized natural language processing tasks, but their deployment in datacenter environments faces significant challenges due to the massive memory requirements of key-value (KV) caches. During the autoregressive decoding process, KV caches consume substantial GPU memory, limiting batch sizes and overall system throughput. To address these challenges, we propose CXL-SpecKV, a novel disaggregated KV-cache architecture that leverages Compute Express Link (CXL) interconnects and FPGA accelerators to enable efficient speculative execution and memory disaggregation. Our approach introduces three key innovations: (i) a CXL-based memory disaggregation framework that offloads KV-caches to remote FPGA memory with low latency, (ii) a speculative KV-cache prefetching mechanism that predicts and preloads future tokens' cache entries, and (iii) an FPGA-accelerated KV-cache compression and decompression engine that reduces memory bandwidth requirements by up to 4×. When evaluated on state-of-the-art LLM models, CXL-SpecKV achieves up to 3.2× higher throughput compared to GPU-only baselines, while reducing memory costs by 2.8× and maintaining accuracy. Our system demonstrates that intelligent memory disaggregation combined with speculative execution can effectively address the memory wall challenge in large-scale LLM serving.
Dong Liu 0049, Yanxuan Yu
FPGA1
2025 TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
abstract
Serving large language models (LLMs) efficiently remains challenging due to the high memory and latency overhead of key-value (KV) cache access during autoregressive decoding. We present TinyServe, a lightweight and extensible serving system for deploying tiny LLMs (e.g., TinyLLaMA, GPT2-345M) with support for structured KV sparsity, plugin-based token selection, and hardware-efficient attention kernels. Unlike prior simulation frameworks, TinyServe executes real-time decoding with configurable sparsity strategies and fine-grained instrumentation. To reduce decoding cost, we introduce a query-aware page selection mechanism that leverages bounding-box metadata to estimate attention relevance between the query and KV cache blocks. This enables selective KV loading with minimal overhead and no model modifications. Our fused CUDA kernel integrates page scoring, sparse memory access, and masked attention in a single pass. Experiments show that TinyServe achieves up to 3.4× speedup and over 2× memory savings with negligible accuracy drop. Additional analysis of cache reuse, page hit rate, and multi-GPU scaling confirms its practicality as an efficient system-level design for LLM training and inference research on resource-constrained hardware.
Dong Liu 0049, Yanxuan Yu
ACM Multimedia1