EDBT 2026 Demo / reviewers in the wild / expert
Pengmiao Zhang
dblp:264/6380
· DBLP profile ↗
10ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-5411-3305ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Net2Tab: Tabularizing neural networks with applications to data prefetching
Pengmiao Zhang, Neelesh Gupta, Rajgopal Kannan, Viktor Prasanna 0001 |
J. Parallel Distributed Comput. | 1 |
| 2025 | GraFetch: Accelerating Graph Applications Through Domain Specific Hierarchical Hybrid PrefetchingabstractMemory performance bottlenecks the execution of graph applications, from traditional graph analytics (GA) to rapidly evolving graph neural networks (GNNs), due to the large size and complexity of graphs. While machine learning (ML) algorithms have shown potential in data prefetching to hide memory access latency, existing approaches face challenges with phase transitions and irregular memory access patterns in graph applications. To address these challenges, we introduce GraFetch, a specialized prefetching system for accelerating graph applications. GraFetch comprises of 1) a novel Hierarchical Hybrid Prefetching (HHP) framework that supports the cooperation of phase-specific ML predictors for high-complexity pattern prefetching and rule-based prefetchers for low-complexity pattern prefetching; and 2) Domain Specific Machine Learning (DSML) models integrated in the framework, which incorporate domain knowledge of graph applications to detect phases, recognize patterns, and predict memory accesses. We evaluate our approach using popular GA frameworks GPOP and X-Stream, and state-of-the-art GNN frameworks PyG and DGL. Our domain specific attention-based memory access predictors achieve 7.4% higher F1-score for delta (consecutive address jump) prediction and 15.35% higher accuracy@10 for page prediction compared with basic attention models. GraFetch achieves an average IPC improvement of 12.47% for GA and 4.18% for GNNs over a system with no prefetcher. This outperforms state-of-the-art rule-based prefetchers BO (7.12% for GA, 1.10% for GNNs), ISB (3.82% for GA, 1.60% for GNNs), and IMP (8.47% for GA, 2.20% for GNNs), as well as ML-based prefetchers Voyager (9.61% for GA, 3.14% for GNNs) and TransFetch (10.98% for GA, 2.48% for GNNs). Pengmiao Zhang, Rajgopal Kannan, Viktor Prasanna 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | TabConv: Low-Computation CNN Inference via Table LookupsabstractConvolutional Neural Networks (CNNs) have demonstrated remarkable ability throughout the field of computer vision. However, CNN inference requires a large number of arithmetic operations making them expensive to deploy in hardware. Current approaches alleviate this issue by developing hardware-supported, algorithmic processes to simplify spatial convolution functions. However, these methods still heavily rely on matrix multiplication, leading to significant computational overhead. To bridge the gap between hardware, algorithmic acceleration, and approximate matrix multiplication, we propose TabConv, a novel, table-based approximation for convolution to significantly reduce arithmetic operations during inference. Additionally, we introduce a priority masking technique based on cosine similarity to select layers for table-based approximation, thereby maintaining the model performance. We evaluate our approach on popular CNNs: ResNet-18, ResNet-34, and NetworkIn-Network (NIN). TabConv preserves over 93% of the original model's performance while reducing arithmetic operations by 36.5%, 25.8%, and 99.4% for ResNet-18 on CIFAR-10, CIFAR-100, and MNIST, respectively, 35.6% and 99.3% for ResNet-34 on CIFAR-10 and MNIST, and 98.9% for NIN on MNIST, achieving low-computation inference. Neelesh Gupta, Narayanan Kannan, Pengmiao Zhang, Viktor Prasanna 0001 |
CF | 3 |
| 2024 | Attention, Distillation, and Tabularization: Towards Practical Neural Network-Based PrefetchingabstractAttention-based Neural Networks (NN) have demonstrated their effectiveness in accurate memory access prediction, an essential step in data prefetching. However, the substantial computational overheads associated with these models result in high inference latency, limiting their feasibility as practical prefetchers. To close the gap, we propose a new approach based on tabularization that significantly reduces model complexity and inference latency without sacrificing prediction accuracy. Our novel tabularization methodology takes input as a distilled, yet highly accurate attention-based model for memory access prediction and efficiently converts its expensive matrix multiplications into a hierarchy of fast table lookups. As an exemplar of the above approach, we develop DART, a prefetcher comprised of a simple hierarchy of tables. With a modest 0.09 drop in F1-score, DART reduces 99.99% of arithmetic operations from the original attention-based model and 91.83% from the distilled model. DART accelerates the large model inference by 170× and the distilled model by 9.4×. DART has comparable latency and storage costs as state-of-the-art rule-based prefetcher BO but surpasses it by 6.1% in IPC improvement. DART outperforms state-of-the-art NN-based prefetchers TransFetch by 33.1% and Voyager by 37.2% in terms of IPC improvement, primarily due to its low prefetching latency. Pengmiao Zhang, Neelesh Gupta, Rajgopal Kannan, Viktor Prasanna 0001 |
IPDPS | 1 |
| 2023 | ME- ViT: A Single-Load Memory-Efficient FPGA Accelerator for Vision TransformersabstractVision Transformers (ViTs) have emerged as a state-of-the-art solution for object classification tasks. However, their computational demands and high parameter count make them unsuitable for real-time inference, prompting the need for efficient hardware implementations. Existing hardware accelerators for ViTs suffer from frequent off-chip memory access, restricting the achievable throughput by memory bandwidth. In devices with a high compute-to-communication ratio (e.g., edge FPGAs with limited bandwidth), off-chip memory access imposes a severe bottleneck on overall throughput. This work proposes ME-ViT, a novel Memory Efficient FPGA accelerator for ViT inference that minimizes memory traffic. We propose a single-load policy in designing ME-ViT: model parameters are only loaded once, intermediate results are stored on-chip, and all operations are implemented in a single processing element. To achieve this goal, we design a memory-efficient processing element (ME-PE), which processes multiple key operations of ViT inference on the same architecture through the reuse of multi-purpose buffers. We also integrate the Softmax and LayerNorm functions into the ME-PE, minimizing stalls between matrix multiplications. We evaluate ME-ViT on systolic array sizes of 32 and 16, achieving up to a 9.22x and 17.89x overall improvement in memory bandwidth, and a 2.16 x improvement in throughput per DSP for both designs over state-of-the-art ViT accelerators on FPGA. ME-ViT achieves a power efficiency improvement of up to 4.00 x (1.03x) over a GPU (FPGA) baseline. ME-ViT enables up to 5 ME-PE instantiations on a Xilinx Alveo U200, achieving a 5.10 x improvement in throughput over the state-of-the art FPGA baseline, and a 5.85 x (1.51 x) improvement in power efficiency over the GPU (FPGA) baseline. Kyle Marino, Pengmiao Zhang, Viktor Prasanna 0001 |
HiPC | 2 |
| 2023 | Phases, Modalities, Spatial and Temporal Locality: Domain Specific ML Prefetcher for Accelerating Graph AnalyticsabstractMemory performance is a key bottleneck in accelerating graph analytics. Existing Machine Learning (ML) prefetchers encounter challenges with phase transitions and irregular memory accesses in graph processing. We propose MPGraph, an ML-based Prefetcher for Graph analytics using domain specific models. MPGraph introduces three novel optimizations: soft detection of phase transitions, phase-specific multi-modality models for access delta and page predictions, and chain spatio-temporal prefetching (CSTP) for prefetch control. Pengmiao Zhang, Rajgopal Kannan, Viktor Prasanna 0001 |
SC | 1 |
| 2022 | Fine-grained address segmentation for attention-based variable-degree prefetchingabstractMachine learning algorithms have shown potential to improve prefetching performance by accurately predicting future memory accesses. Existing approaches are based on the modeling of text prediction, considering prefetching as a classification problem for sequence prediction. However, the vast and sparse memory address space leads to large vocabulary, which makes this modeling impractical. The number and order of outputs for multiple cache line prefetching are also fundamentally different from text prediction. Pengmiao Zhang, Ajitesh Srivastava, Anant Nori, Rajgopal Kannan, Viktor Prasanna 0001 |
CF | 1 |
| 2022 | A2P: Attention-based Memory Access Prediction for Graph Analytics
Pengmiao Zhang, Rajgopal Kannan, Anant Nori, Viktor Prasanna 0001 |
DATA | 1 |
| 2022 | ReSemble: Reinforced Ensemble Framework for Data PrefetchingabstractData prefetching hides memory latency by predicting and loading necessary data into cache beforehand. Most prefetchers in the literature are efficient for specific memory address patterns thereby restricting their utility to specialized applications-they do not perform well on hybrid applications with multifarious access patterns. Therefore we propose ReSem-ble: a Reinforcement Learning (RL) based adaptive enSemble framework that enables multiple prefetchers to complement each other on hybrid applications. Our RL trained ensemble controller takes prefetch suggestions from all prefetchers as input, selects the best suggestion dynamically, and learns online toward getting higher cumulative rewards, which are collected from prefetch hits/misses. Our ensemble framework using a simple multilayer perceptron as the controller achieves on the average 85.27 % (accuracy) and 44.22 % (coverage), leading to 31.02 % IPC improvement, which outperforms state-of-the-art individual prefetchers by 8.35%-26.11 %, while also outperforming SBP, a state-of-the-art (non-RL) ensemble prefetcher by 5.69%. Pengmiao Zhang, Rajgopal Kannan, Ajitesh Srivastava, Anant Nori, Viktor Prasanna 0001 |
SC | 1 |
| 2020 | MemMAP: Compact and Generalizable Meta-LSTM Models for Memory Access Prediction
Ajitesh Srivastava, Ta-Yang Wang, Pengmiao Zhang, César A. F. De Rose, Rajgopal Kannan, Viktor Prasanna 0001 |
PAKDD (2) | 3 |