Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Morteza Baradaran

dblp:350/5417 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0004-0705-2820ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 67% Performance modeling and evaluation · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
benchmarking
1.012026
Characterizing Digital DRAM PIM through Modeling and Benchmarking · ACM Trans. Archit. Code Optim. 2026
Memory systems
processing-in-memory
1.012026
Characterizing Digital DRAM PIM through Modeling and Benchmarking · ACM Trans. Archit. Code Optim. 2026
Memory systems › processing-in-memory
processing-using-DRAM
1.012026
Characterizing Digital DRAM PIM through Modeling and Benchmarking · ACM Trans. Archit. Code Optim. 2026

Methods — techniques the papers use, named apart from their topics

roofline analysis · 1.0
YearPublicationVenuePosition
2026 HARMONI: Hierarchical ARchitecture MOdeling for LLMs with Near/In Memory Computing
abstract
LLM inference has emerged as a strongly memory-bound workload suitable for Processing-in-Memory (PIM) and Processing-near-Memory (PNM) architectures. Yet, most existing PIM-LLM evaluation frameworks are in-house, closed-source, or difficult to extend, and often fail to model end-to-end inference behavior or tensor-allocation and communication effects that critically shape performance. We introduce HARMONI, a fast, modular, memory-centric performance modeling framework designed specifically for hierarchical PIM/PNM architectures running LLM workloads. HARMONI captures the complete inference execution through a task-graph representation and supports user-defined tensor allocation with address-interleaving schemes. It models computation, communication, and queuing delays across logic nodes, enabling realistic mapping of inference kernels onto heterogeneous logic nodes. HARMONI outputs detailed, kernel-wise and resource-wise time and energy breakdowns, providing actionable visibility into architectural bottlenecks. Together, these capabilities make HARMONI a practical tool for rapidly evaluating PIM/PNM designs and DRAM standards (e.g., DDR4, DDR5, GDDR6), enabling systematic co-evaluation of inference workloads and memory-centric architecture designs.
Khyati Kiyawat, Yasas Seneviratne, Zhenxing Fan, Morteza Baradaran, Kevin Skadron
ISPASS4
2026 Characterizing Digital DRAM PIM through Modeling and Benchmarking
abstract
The disparity between processor speed and memory bandwidth has become a growing performance bottleneck, particularly for memory-intensive workloads. Processing-in-Memory (PIM) mitigates this bottleneck by integrating computation directly within DRAM. However, the effectiveness of PIM varies significantly across workloads, architectures, and DRAM technologies, yet it is often assessed using tightly coupled simulators and benchmarks that lack portability and generality. This article extends PIMbench and PIMeval —a generalizable benchmark suite and an extensible PIM simulator to support a broader range of workloads, PIM architectures, and DRAM technologies. This evaluation incorporates roofline analysis and a breakdown of intra-memory execution stages to identify PIM-specific bottlenecks and performance scaling limits. The evaluation spans three classes of digital PIM architectures: subarray-level bit-serial, subarray-level bit-parallel, and bank-level bit-parallel. It further demonstrates how internal DRAM parameters such as subarray count and GDL width impact PIM performance. The code is publicly available at: https://github.com/UVA-LavaLab/PIMeval-PIMbench .
Farzana Siddique, Deyuan Guo, Hugo Abbot, Kyle Durrer, MohammadHosein Gholamrezaei, Morteza Baradaran, Ethan Ermovick, Alif Ahmed, Zhenxing Fan, Beenish Gul, Ashish Venkat, Kevin Skadron
ACM Trans. Archit. Code Optim.6