EDBT 2026 Demo / reviewers in the wild / expert
Maryam Babaie
dblp:62/11054
· DBLP profile ↗
3ranked-venue papers
3as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 91% High-performance computing · 9% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache design |
0.9 | 1 | 2025 | Efficient Caching with A Tag-enhanced DRAM · HPCA 2025 |
Memory systems › cache
DRAM cache |
0.9 | 1 | 2025 | Efficient Caching with A Tag-enhanced DRAM · HPCA 2025 |
Memory systems › DRAM
DRAM microarchitecture |
0.9 | 1 | 2025 | Efficient Caching with A Tag-enhanced DRAM · HPCA 2025 |
Methods — techniques the papers use, named apart from their topics
full-system simulation · 0.9flush buffer · 0.9early tag probing · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Characterization of Coherent Data Movement in Scaled-Out Shared Memory SystemsabstractAs core counts expand through chiplet integration, advanced packaging, and memory expansion technologies, modern systems increasingly rely on large-scale and coherent shared-memory architectures. While coherence protocols are well studied, the empirical behavior of coherent data movement at system scale remains poorly understood. This work presents a detailed characterization of coherent data movement using cyclelevel full-system simulation of a modeled large-scale hierarchical shared-memory system. Across representative HPC and graph processing workloads, we find that coherence overheads are dominated by a small subset of high-impact pages, despite most pages exhibiting limited sharing. These pages combine wide socket span, high access frequency, and millisecond-scale temporal persistence, generating sustained long-distance coherence traffic that frequently propagates beyond the local chassis into higher levels of the interconnect. Together, these observations indicate that coherence cost and challenges are driven by a sparse but dominant tail of shared data, motivating architectural mechanisms that adapt data placement and coherence anchoring to page-level access behavior to make coherence mechanisms more efficient. Maryam Babaie, Michael R. Miller, Wendy Elsasser, Steven C. Woo, Jason Lowe-Power |
ISPASS | 1 |
| 2025 | Efficient Caching with A Tag-enhanced DRAMabstractAs SRAM-based caches are hitting a scaling wall, manufacturers are integrating DRAM-based caches into system designs to continue increasing cache sizes. While DRAM caches can improve the performance of memory systems, existing DRAM cache designs suffer from high miss penalties, wasted data movement, and interference between misses and demands. In this paper, we propose TDRAM, a novel DRAM microarchitecture tailored for caching. TDRAM enhances existing DRAM, such as HBM3, by adding small, low-latency mats to store tags and metadata on the same die as the data mats. These mats enable tag and data access in lockstep, in-DRAM tag comparison, and conditional data response based on the comparison result (reducing wasted data transfers), akin to SRAM cache mechanisms. TDRAM further optimizes hit and miss latencies through opportunistic early tag probing. Moreover, TDRAM introduces a flush buffer to store conflicting dirty data on write misses, eliminating data bus turnaround delays on write demands. We evaluate TDRAM in a full-system simulation using a set of HPC workloads with large memory footprints, showing that TDRAM, on average, provides $2.65 \times$ faster tag checks, $1.23 \times$ speedup, and 21% less energy consumption compared to state-of-the-art commercial and research designs. Maryam Babaie, Ayaz Akram, Wendy Elsasser, Brent Haukness, Michael R. Miller, Taeksang Song, Thomas Vogelsang, Steven C. Woo, Jason Lowe-Power |
HPCA | 1 |
| 2023 | Enabling Design Space Exploration of DRAM Caches for Emerging Memory SystemsabstractThe increasing growth of applications’ memory capacity and performance demands has led the CPU vendors to deploy heterogeneous memory systems either within a single system or via disaggregation. DRAM caches are one way to enable heterogeneity and disaggregation in such systems. While there is significant research investigating the designs of DRAM caches, there has been little research investigating DRAM caches from a full system point of view, because there is not a suitable model available to the community to accurately study large-scale systems with DRAM caches at a cycle-level. In this work we describe a new cycle-level DRAM cache model in the gem5 simulator which can be used for emerging heterogeneous and disaggregated memory systems. Maryam Babaie, Ayaz Akram, Jason Lowe-Power |
ISPASS | 1 |