VLDB 2026 Research / reviewers in the wild / expert
Mengfei Rong
dblp:409/9802
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0009-2202-2327ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 75% GPUs and heterogeneous computing · 25% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU kernel optimization |
1.0 | 1 | 2026 | Accelerating Sparse Transformer Inference on GPU · PPoPP 2026 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
1.0 | 1 | 2026 | Accelerating Sparse Transformer Inference on GPU · PPoPP 2026 |
Hardware accelerators and domain-specific architectures › dataflow optimization
operator fusion |
1.0 | 1 | 2026 | Accelerating Sparse Transformer Inference on GPU · PPoPP 2026 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
transformer inference accelerator |
1.0 | 1 | 2026 | Accelerating Sparse Transformer Inference on GPU · PPoPP 2026 |
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating Sparse Transformer Inference on GPUabstractLarge language models (LLMs) are popular around the world due to their powerful understanding capabilities. As the core component of LLMs, accelerating Transformer through parallelization has gradually become a hot research topic. Mask layers introduce sparsity into Transformer to reduce calculations. However, previous works rarely focus on the performance optimization of sparse Transformer. In addition, current static operator fusion schemes fail to adapt to diverse application scenarios. To address the above problems, we propose STOF, a framework that incorporates optimizations for Sparse Transformer that enables flexible masking and Operator Fusion on GPU. For multi-head attention (MHA) structure, STOF maps the computation to row-wise or block-wise kernels with unique storage formats according to analytical modeling. For downstream operators, STOF maps the fusion scheme to compilation templates and determines the optimal running configuration through two-stage searching. The experimental results show that compared to the state-of-the-art work, STOF achieves maximum speedups of 1.6× in MHA computation and 1.4× in end-to-end inference. Wenhao Dai, Haodong Deng, Mengfei Rong, Fangxin Liu, Hailong Yang 0002, Qianwen Cao, Qingxiao Sun |
PPoPP | 3 |