VLDB 2026 Research / reviewers in the wild / expert
Morteza Baradaran
dblp:350/5417
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0004-0705-2820ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 67% Performance modeling and evaluation · 33% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
benchmarking |
1.0 | 1 | 2026 | Characterizing Digital DRAM PIM through Modeling and Benchmarking · ACM Trans. Archit. Code Optim. 2026 |
Memory systems
processing-in-memory |
1.0 | 1 | 2026 | Characterizing Digital DRAM PIM through Modeling and Benchmarking · ACM Trans. Archit. Code Optim. 2026 |
Memory systems › processing-in-memory
processing-using-DRAM |
1.0 | 1 | 2026 | Characterizing Digital DRAM PIM through Modeling and Benchmarking · ACM Trans. Archit. Code Optim. 2026 |
Methods — techniques the papers use, named apart from their topics
roofline analysis · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HARMONI: Hierarchical ARchitecture MOdeling for LLMs with Near/In Memory ComputingabstractLLM inference has emerged as a strongly memory-bound workload suitable for Processing-in-Memory (PIM) and Processing-near-Memory (PNM) architectures. Yet, most existing PIM-LLM evaluation frameworks are in-house, closed-source, or difficult to extend, and often fail to model end-to-end inference behavior or tensor-allocation and communication effects that critically shape performance. We introduce HARMONI, a fast, modular, memory-centric performance modeling framework designed specifically for hierarchical PIM/PNM architectures running LLM workloads. HARMONI captures the complete inference execution through a task-graph representation and supports user-defined tensor allocation with address-interleaving schemes. It models computation, communication, and queuing delays across logic nodes, enabling realistic mapping of inference kernels onto heterogeneous logic nodes. HARMONI outputs detailed, kernel-wise and resource-wise time and energy breakdowns, providing actionable visibility into architectural bottlenecks. Together, these capabilities make HARMONI a practical tool for rapidly evaluating PIM/PNM designs and DRAM standards (e.g., DDR4, DDR5, GDDR6), enabling systematic co-evaluation of inference workloads and memory-centric architecture designs. Khyati Kiyawat, Yasas Seneviratne, Zhenxing Fan, Morteza Baradaran, Kevin Skadron |
ISPASS | 4 |
| 2026 | Characterizing Digital DRAM PIM through Modeling and BenchmarkingabstractThe disparity between processor speed and memory bandwidth has become a growing performance bottleneck, particularly for memory-intensive workloads. Processing-in-Memory (PIM) mitigates this bottleneck by integrating computation directly within DRAM. However, the effectiveness of PIM varies significantly across workloads, architectures, and DRAM technologies, yet it is often assessed using tightly coupled simulators and benchmarks that lack portability and generality. This article extends PIMbench and PIMeval —a generalizable benchmark suite and an extensible PIM simulator to support a broader range of workloads, PIM architectures, and DRAM technologies. This evaluation incorporates roofline analysis and a breakdown of intra-memory execution stages to identify PIM-specific bottlenecks and performance scaling limits. The evaluation spans three classes of digital PIM architectures: subarray-level bit-serial, subarray-level bit-parallel, and bank-level bit-parallel. It further demonstrates how internal DRAM parameters such as subarray count and GDL width impact PIM performance. The code is publicly available at: https://github.com/UVA-LavaLab/PIMeval-PIMbench . Farzana Siddique, Deyuan Guo, Hugo Abbot, Kyle Durrer, MohammadHosein Gholamrezaei, Morteza Baradaran, Ethan Ermovick, Alif Ahmed, Zhenxing Fan, Beenish Gul, Ashish Venkat, Kevin Skadron |
ACM Trans. Archit. Code Optim. | 6 |