Fatemeh Golshan

dblp:274/1574 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2024
0000-0003-4250-1623ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 65% Parallel and multicore computing · 19% Processor architecture and microarchitecture · 16%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › DRAM › DRAM architecture
3D-stacked DRAM
0.812024
Blenda: Dynamically-Reconfigurable Stacked DRAM · MICRO 2024
Parallel and multicore computing › task partitioning
dynamic partitioning
0.812024
Blenda: Dynamically-Reconfigurable Stacked DRAM · MICRO 2024
Memory systems
memory management
0.812024
Blenda: Dynamically-Reconfigurable Stacked DRAM · MICRO 2024
Memory systems
cache
0.712023
MANA: Microarchitecting a Temporal Instruction Prefetcher · IEEE Trans. Computers 2023
Processor architecture and microarchitecture › instruction fetch
instruction prefetching
0.712023
MANA: Microarchitecting a Temporal Instruction Prefetcher · IEEE Trans. Computers 2023
Memory systems
DRAM
0.212024
Blenda: Dynamically-Reconfigurable Stacked DRAM · MICRO 2024
Memory systems › cache
cache miss reduction
0.212023
MANA: Microarchitecting a Temporal Instruction Prefetcher · IEEE Trans. Computers 2023

Methods — techniques the papers use, named apart from their topics

dynamic partitioning · 0.8cache filtering · 0.8storage cost reduction · 0.7metadata record design · 0.7
YearPublicationVenuePosition
2024 Blenda: Dynamically-Reconfigurable Stacked DRAM
abstract
This paper proposes Blenda, a dynamically-partitioned memory-cache blend architecture for giga-scale die-stacked DRAMs. Blenda architects the stacked DRAM partly as memory and partly as cache, and dynamically adjusts each part's size to workloads' demands. The memory part hosts hot data objects and serves requests to them efficiently (i.e., without metadata overheads). The cache part captures transient data and filters requests to bandwidth-limited off-chip DRAM. Blenda provides three key contributions: (i) Blenda partitions stacked DRAM's capacity in a workload-aware manner: different workloads enjoy different memory-cache configurations. (ii) Blenda is reactive: the configuration is adjusted to workloads' phases dynamically and application-transparently: no reboot or user involvement are needed. (iii) Blenda gracefully transitions among configurations: no data invalidation is required upon most reconfigurations. We simulate 15 diverse big-data workloads running on a state-of-the-art processor and show that Blenda outperforms the best-performing prior architecture by 34%. Blenda's total storage overhead is less than 100 bytes per core.
Mohammad Bakhshalipour, Hamidreza Zare, Farid Samandi, Fatemeh Golshan, Pejman Lotfi-Kamran, Hamid Sarbazi-Azad
MICRO4
2023 MANA: Microarchitecting a Temporal Instruction Prefetcher
abstract
L1 instruction(L1-l) cache misses are a source of performance bottleneck. While many instruction prefetchers have been proposed, most of them leave a considerable potential uncovered. In 2011, Proactive Instruction Fetch (PIF) showed that a hardware prefetcher could effectively eliminate all instruction-cache misses. However, its enormous storage cost makes it impractical. Consequently, reducing the storage cost was the main research focus in instruction prefetching in the past decade. Several instruction prefetchers, including RDIP and Shotgun, were proposed to offer PIF-level performance with significantly lower storage overhead. However, our findings show that there is a considerable performance gap between these proposals and PIF. While these proposals use different mechanisms for prefetching, the performance gap is mainly not because of the mechanism, and instead, is due to not having sufficient storage. We make the case that the key to designing a powerful and cost-effective instruction prefetcher is choosing a metadata record and microarchitecting the prefetcher to minimize the storage. Our proposal, MANA, offers PIF-level performance with 15.7x lower storage cost. MANA outperforms RDIP and Shotgun by 12.5 and 29%, respectively. We also evaluate a version of MANA with no storage overhead and show that it offers 98% of the peak performance benefits.
Ali Ansari 0001, Fatemeh Golshan, Rahil Barati, Pejman Lotfi-Kamran, Hamid Sarbazi-Azad
IEEE Trans. Computers2