EDBT 2026 Demo / reviewers in the wild / expert
Ahmad Javadi Nezhad
dblp:401/2450
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0004-4444-150XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Low-latency On-chip Cache Hierarchy for Load-to-use Stall Reduction in GPUsabstractMemory hierarchy in Graphics Processing Units (GPUs) is conventionally designed to provide high bandwidth rather than low latency. In particular, because of the high tolerance to load-to-use latency (i.e., the time that warps wait for data fetched by memory loads), GPU L1D caches are optimized for density, capacity, and low power with latencies that are often orders of magnitude longer than conventional CPU caches. However, there are many important classes of data-parallel applications (e.g., graph, tree, priority queue processing, and sparse deep learning applications) that benefit from lower load-to-use latency than that offered by modern GPUs due to their inherent divergence and low effective Thread-Level Parallelism (TLP). This article introduces an innovative on-chip cache hierarchy that incorporates a decoupled L1D cache with reduced latency (LoTUS) and its management scheme. LoTUS is a minimally sized fully associative cache placed in each GPU subcore that captures the primary working set of data-parallel applications. It exploits conventional high-performance low-density SRAM cells and dramatically reduces load-to-use latency. We also propose an intelligent extension of LoTUS, called LoTUSage, which employs a lightweight learning-based model to predict the utility of caching requests in LoTUS. Evaluation results show that LoTUS and LoTUSage improve the average performance by 23.9% and 35.4% and reduce the average energy consumption by 27.8% and 38.5%, respectively, for the applications suffering from high load-to-use stalls with negligible area and power overheads. Negin Mahani, Hajar Falahati, Sina Darabi, Ahmad Javadi Nezhad, Yunho Oh, Mohammad Sadrosadati, Hamid Sarbazi-Azad, Babak Falsafi |
ACM Trans. Archit. Code Optim. | 4 |
| 2025 | A comprehensive review and classification of micro-scale and macro-scale interconnection network simulators for research and education in network-based computing systems
Atiyeh Gheibi-Fetrat, Fatemeh Serajeh-hassani, Negar Akbarzadeh, Amir Mirzaei, Mahmoud Reza Kheyrati-Fard, Ahmad Javadi Nezhad, Jeong-A Lee, Hamid Sarbazi-Azad |
J. Supercomput. | 6 |
| 2025 | A survey of SSD simulators and emulators
Atiyeh Gheibi-Fetrat, Fatemeh Serajeh-hassani, Masoud Mohammadi-Lak, Amir Mirzaei, Negar Akbarzadeh, Mahmoud Reza Kheyrati-Fard, Mohammad Hosseini 0001, Ahmad Javadi Nezhad, Arash Tavakkol, Jeong-A Lee, Hamid Sarbazi-Azad |
J. Supercomput. | 8 |
| 2025 | MQSimNet: an open-source simulator for next-generation network-based SSDs
Amir Mirzaei, Fatemeh Serajeh-hassani, Atiyeh Gheibi-Fetrat, Mina Zabihi, Sina Ghorbani-Jabbedar, Mahmoud Reza Kheyrati-Fard, Ahmad Javadi Nezhad, Mohammad Hosseini 0001, Negar Akbarzadeh, Jeong-A Lee, Hamid Sarbazi-Azad |
J. Supercomput. | 7 |