VLDB 2026 Research / reviewers in the wild / expert
Weishu Deng
dblp:367/4239
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0000-6550-7484ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 51% GPUs and heterogeneous computing · 24% Hardware accelerators and domain-specific architectures · 24% | |
| Software engineering, system software, and programming languages
1 paper |
Operating systems · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU computing |
1.0 | 1 | 2026 | Scaling Attention Beyond GPUs for LLM Inference · HPDC 2026 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference
LLM inference |
1.0 | 1 | 2026 | Scaling Attention Beyond GPUs for LLM Inference · HPDC 2026 |
Memory systems › virtual memory management
page migration |
0.8 | 1 | 2024 | Nomad: Non-Exclusive Memory Tiering via Transactional Page Migration · OSDI 2024 |
Memory systems
tiered memory |
0.8 | 1 | 2024 | Nomad: Non-Exclusive Memory Tiering via Transactional Page Migration · OSDI 2024 |
Memory systems › memory offloading
KV cache offloading |
0.3 | 1 | 2026 | Scaling Attention Beyond GPUs for LLM Inference · HPDC 2026 |
Memory systems
memory hierarchy |
0.3 | 1 | 2026 | Scaling Attention Beyond GPUs for LLM Inference · HPDC 2026 |
Operating systems › resource management › memory management
virtual memory |
0.2 | 1 | 2024 | Nomad: Non-Exclusive Memory Tiering via Transactional Page Migration · OSDI 2024 |
Methods — techniques the papers use, named apart from their topics
sparse attention · 1.0log-sum-exp fusion · 1.0CPU-GPU hybrid attention · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling Attention Beyond GPUs for LLM InferenceabstractScaling inference for large language models is increasingly constrained by limited GPU memory, primarily due to the expanding intermediate states (KV caches) required for long-context generation and multi-user workloads. Once the KV cache exceeds the capacity of high-bandwidth memory, it must be offloaded to host memory and reloaded on demand, a workflow severely bottlenecked by the CPU–GPU interconnect, typically PCIe. Existing approaches exploiting offload KV caches to CPU memory and selectively reload partial segments for attention computation often underutilize CPU compute resources and suffer from accuracy degradation. We present Beyond, a drop-in runtime that integrates a smart offloading scheme to selectively identify and retain salient KV entries across continuous decoding sessions, together with a hybrid CPU–GPU attention mechanism for scalable inference. Beyond executes dense attention over recent KV entries stored in GPU memory while performing parallel, per-head sparse attention on salient contextual KV entries residing in CPU memory. The outputs are fused efficiently through a log-sum-exp scheme. During the bandwidth-constrained decoding phase, oversized KV caches are processed cooperatively by the aggregated CPU and GPU memory bandwidth, with only minimal PCIe data movement. Experiments across diverse models and workloads demonstrate that Beyond improves scalability, supports longer sequences and larger batch sizes, and outperforms existing sparse attention baselines in both efficiency and accuracy—all on commodity GPU hardware. Weishu Deng, Peiran Du, Lingfeng Xiang, Chen Zhong 0002, Faraz Ahmed, Lianjie Cao, Puneet Sharma 0001, Song Jiang 0001, Hui Lu 0001, Jia Rao |
HPDC | 1 |
| 2024 | Mega: More Efficient Graph Attention for GNNsabstractGraph neural networks (GNNs) have demonstrated effectiveness across diverse application domains by leveraging graph information to uncover intrinsic correlations alongside feature representation. This enables GNNs to explore richer information compared to conventional neural networks, resulting in enhanced predictive performance. However, the integration of graphstructured data into the learning process poses two challenges. First, the sparsity and irregularity of the graph representation result in inefficient and expensive memory accesses on throughput-oriented accelerators, such as GPUs. Second, as GNN training involves interleaved graph operations to extract topological information and neural operations to update node or edge embeddings, the joint optimization of these two operations on accelerators is challenging due to their distinct resource requirements. Among GNN graph operations, graph attention which helps focus GNN training on highly correlated nodes is critical to training performance and model accuracy. However, our profiling of representative GNNs reveals that irregular memory access during graph attention accounts for the dominating overhead in GNN training. To address this issue, this paper proposes a more efficient graph attention method Megato accelerate GNN training. MegA converts the original graph representation into one that regularizes memory access patterns for graph attention. Specifically, during preprocessing, Megatraverses a graph to derive a schedule for graph attention and uses the schedule to reorganize the graph representation for optimized memory access. MegA explores several techniques to balance memory access efficiency and preserve the original graph properties to avoid the loss of model accuracy. Experimental results with representative GNNs and graph data sets show that Megaconsistently outperforms conventional graph attention methods with up to 3x speedup. Weishu Deng, Jia Rao |
ICDCS | 1 |
| 2024 | Nomad: Non-Exclusive Memory Tiering via Transactional Page Migration
Lingfeng Xiang, Weishu Deng, Hui Lu 0001, Jia Rao, Ren Wang 0001 |
OSDI | 3 |