Negar Akbarzadeh

dblp:240/3581 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-8040-4317ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 6 since 2021
YearPublicationVenuePosition
2026 A comprehensive survey on multi-GPU systems
Atiyeh Gheibi-Fetrat, Arad Maleki, Sahand Zoufan, Masoud Mohammadi-Lak, Amirsaeed Ahmadi-Tonekaboni, Mahdi Alinejad, Komeil Yahyazadeh, Mohammad Alizadeh, Negar Akbarzadeh, Sina Darabi-moghaddam, Shaahin Hessabi, Hamid Sarbazi-Azad
Parallel Comput.9
2025 A comprehensive review and classification of micro-scale and macro-scale interconnection network simulators for research and education in network-based computing systems
Atiyeh Gheibi-Fetrat, Fatemeh Serajeh-hassani, Negar Akbarzadeh, Amir Mirzaei, Mahmoud Reza Kheyrati-Fard, Ahmad Javadi Nezhad, Jeong-A Lee, Hamid Sarbazi-Azad
J. Supercomput.3
2025 A survey of SSD simulators and emulators
Atiyeh Gheibi-Fetrat, Fatemeh Serajeh-hassani, Masoud Mohammadi-Lak, Amir Mirzaei, Negar Akbarzadeh, Mahmoud Reza Kheyrati-Fard, Mohammad Hosseini 0001, Ahmad Javadi Nezhad, Arash Tavakkol, Jeong-A Lee, Hamid Sarbazi-Azad
J. Supercomput.5
2025 MQSimNet: an open-source simulator for next-generation network-based SSDs
Amir Mirzaei, Fatemeh Serajeh-hassani, Atiyeh Gheibi-Fetrat, Mina Zabihi, Sina Ghorbani-Jabbedar, Mahmoud Reza Kheyrati-Fard, Ahmad Javadi Nezhad, Mohammad Hosseini 0001, Negar Akbarzadeh, Jeong-A Lee, Hamid Sarbazi-Azad
J. Supercomput.9
2022 Morpheus: Extending the Last Level Cache Capacity in GPU Systems Using Idle GPU Core Resources
abstract
Graphics Processing Units (GPUs) are widely-used accelerators for data-parallel applications. In many GPU applications, GPU memory bandwidth bottlenecks performance, causing underutilization of GPU cores. Hence, disabling many cores does not affect the performance of memory-bound workloads. While simply power-gating unused GPU cores would save energy, prior works attempt to better utilize GPU cores for other applications (ideally compute-bound), which increases the GPU’s total throughput. In this paper, we introduce Morpheus, a new hardware/software co-designed technique to boost the performance of memory-bound applications. The key idea of Morpheus is to exploit unused core resources to extend the GPU last level cache (LLC) capacity. In Morpheus, each GPU core has two execution modes: compute mode and cache mode. Cores in compute mode operate conventionally and run application threads. However, for the cores in cache mode, Morpheus invokes a software helper kernel that uses the cores’ on-chip memories (i.e., register file, shared memory, and L1) in a way that extends the LLC capacity for a running memory-bound workload. Morpheus adds a controller to the GPU hardware to forward LLC requests to either the conventional LLC (managed by hardware) or the extended LLC (managed by the helper kernel). Our experimental results show that Morpheus improves the performance and energy efficiency of a baseline GPU architecture by an average of 39% and 58%, respectively, across several memory-bound workloads. Morpheus’ performance is within 3% of a GPU design that has a quadruple-sized conventional LLC. Morpheus can thus contribute to reducing the hardware dedicated to a conventional LLC by exploiting idle cores’ on-chip memory resources as additional cache capacity.
Sina Darabi, Mohammad Sadrosadati, Negar Akbarzadeh, Joël Lindegger, Mohammad Hosseini 0001, Jisung Park 0001, Juan Gómez-Luna, Onur Mutlu, Hamid Sarbazi-Azad
MICRO3
2022 OSM: Off-Chip Shared Memory for GPUs
abstract
Graphics Processing Units (GPUs) employ a shared memory, a software-managed cache for programmers, in each streaming multiprocessor to accelerate data sharing among the threads in a thread block. Although 60% of the shared memory space is underutilized, on average, there are some workloads that demand higher shared memory capacities. Therefore, improving shared memory utilization while satisfying the needs of shared memory intensive workloads is challenging. We make a key observation that the lifetime of each shared memory address is significantly shorter than the execution time of a thread block. In this paper, we first propose Off-Chip Shared Memory (OSM) that allocates shared memory space in the off-chip memory, and accelerates accesses to it via a small on-chip cache. Using an 8 KB cache for shared memory addresses, OSM provides almost the same performance as the baseline GPU that uses 96 KB on-chip shared memory. OSM improves GPU performance in two ways. First, it allocates higher shared memory capacities in the off-chip memory, and improves thread-level parallelism (TLP). Second, it designs a unified cache for shared memory and global address spaces, providing more caching space for global memory address space even for the workloads with high shared memory utilization. Our experimental results show an average 21% and 18% IPC improvement compared to the baseline and the state-of-the-art architectures.
Sina Darabi, Ehsan Yousefzadeh-Asl-Miandoab, Negar Akbarzadeh, Hajar Falahati, Pejman Lotfi-Kamran, Mohammad Sadrosadati, Hamid Sarbazi-Azad
IEEE Trans. Parallel Distributed Syst.3