EDBT 2026 Demo / reviewers in the wild / expert
Michael R. Miller
dblp:64/3665
· DBLP profile ↗
3ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 91% High-performance computing · 9% | |
| Artificial intelligence
1 paper |
Reinforcement learning · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache design |
0.9 | 1 | 2025 | Efficient Caching with A Tag-enhanced DRAM · HPCA 2025 |
Memory systems › cache
DRAM cache |
0.9 | 1 | 2025 | Efficient Caching with A Tag-enhanced DRAM · HPCA 2025 |
Memory systems › DRAM
DRAM microarchitecture |
0.9 | 1 | 2025 | Efficient Caching with A Tag-enhanced DRAM · HPCA 2025 |
Machine learning › Reinforcement learning
preference learning |
0.0 | 1 | 2000 | Learning Priorities From Noisy Examples · ICML 2000 |
Methods — techniques the papers use, named apart from their topics
full-system simulation · 0.9flush buffer · 0.9early tag probing · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Characterization of Coherent Data Movement in Scaled-Out Shared Memory SystemsabstractAs core counts expand through chiplet integration, advanced packaging, and memory expansion technologies, modern systems increasingly rely on large-scale and coherent shared-memory architectures. While coherence protocols are well studied, the empirical behavior of coherent data movement at system scale remains poorly understood. This work presents a detailed characterization of coherent data movement using cyclelevel full-system simulation of a modeled large-scale hierarchical shared-memory system. Across representative HPC and graph processing workloads, we find that coherence overheads are dominated by a small subset of high-impact pages, despite most pages exhibiting limited sharing. These pages combine wide socket span, high access frequency, and millisecond-scale temporal persistence, generating sustained long-distance coherence traffic that frequently propagates beyond the local chassis into higher levels of the interconnect. Together, these observations indicate that coherence cost and challenges are driven by a sparse but dominant tail of shared data, motivating architectural mechanisms that adapt data placement and coherence anchoring to page-level access behavior to make coherence mechanisms more efficient. Maryam Babaie, Michael R. Miller, Wendy Elsasser, Steven C. Woo, Jason Lowe-Power |
ISPASS | 2 |
| 2025 | Efficient Caching with A Tag-enhanced DRAMabstractAs SRAM-based caches are hitting a scaling wall, manufacturers are integrating DRAM-based caches into system designs to continue increasing cache sizes. While DRAM caches can improve the performance of memory systems, existing DRAM cache designs suffer from high miss penalties, wasted data movement, and interference between misses and demands. In this paper, we propose TDRAM, a novel DRAM microarchitecture tailored for caching. TDRAM enhances existing DRAM, such as HBM3, by adding small, low-latency mats to store tags and metadata on the same die as the data mats. These mats enable tag and data access in lockstep, in-DRAM tag comparison, and conditional data response based on the comparison result (reducing wasted data transfers), akin to SRAM cache mechanisms. TDRAM further optimizes hit and miss latencies through opportunistic early tag probing. Moreover, TDRAM introduces a flush buffer to store conflicting dirty data on write misses, eliminating data bus turnaround delays on write demands. We evaluate TDRAM in a full-system simulation using a set of HPC workloads with large memory footprints, showing that TDRAM, on average, provides $2.65 \times$ faster tag checks, $1.23 \times$ speedup, and 21% less energy consumption compared to state-of-the-art commercial and research designs. Maryam Babaie, Ayaz Akram, Wendy Elsasser, Brent Haukness, Michael R. Miller, Taeksang Song, Thomas Vogelsang, Steven C. Woo, Jason Lowe-Power |
HPCA | 5 |
| 2000 | Learning Priorities From Noisy Examples
Geoffrey G. Towell, Thomas Petsche, Michael R. Miller |
ICML | 3 |