EDBT 2026 Demo / reviewers in the wild / expert
Alireza Monemi
dblp:132/2980
· DBLP profile ↗
7ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-3438-3877ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FetchFlare: An Open-Source Strided Data Prefetcher for High-Performance Cache HierarchiesabstractIn recent years, the rise of open-source hardware has transformed the landscape of technology development. In particular, RISC-V has offered hardware designers the possibility of designing processors in a much cheaper way by leveraging a rich ecosystem of open-source designs that can be easily reused, extended, and customized. Although the RISC-V ecosystem is rapidly growing and open-source processors are becoming increasingly sophisticated, some advanced architectural techniques typically employed in commercial high-performance processors are still not prevalent in RISC-V open-source architectures. Among them, hardware prefetchers have been ubiquitous in highend processors for many years, but they are not as commonly found in open-source RISC-V processors. To bridge this gap, this work presents FetchFlare, a stride prefetcher for highperformance cache hierarchies. FetchFlare is able to capture the memory access patterns of applications, predict future memory accesses, and issue prefetch requests for them. We provide an open-source RTL implementation of FetchFlare and integrate it into a complete open-source setup formed by the OpenPiton framework, the Sargantana core, and the High-Performance Data Cache (HPDCache). Compared to a baseline system without prefetching, FetchFlare achieves an average speedup of $63 \%$, avoids cache misses in the L1D and the L2 caches, and presents an average accuracy, coverage, and timeliness of $86 \%, 39 \%$, and 99%, respectively. Golnaz Korkian, Neiel Leyva, Arnau Bigas, Noelia Oliete-Escuín, Abbas Haghi, Alireza Monemi, César Fuguet Tortolero, Lluc Alvarez |
DSD | 6 |
| 2024 | A Mess of Memory System Benchmarking, Simulation and Application ProfilingabstractThe Memory stress (Mess) framework provides a unified view of the memory system benchmarking, simulation and application profiling. The Mess benchmark provides a holistic and detailed memory system characterization. It is based on hundreds of measurements that are represented as a family of bandwidth-latency curves. The benchmark increases the coverage of all the previous tools and leads to new findings in the behavior of the actual and simulated memory systems. We deploy the Mess benchmark to characterize Intel, AMD, IBM, Fujitsu, Amazon and NVIDIA servers with DDR4, DDR5, HBM2 and HBM2E memory. The Mess memory simulator uses bandwidth-latency concept for the memory performance simulation. We integrate Mess with widely-used CPUs simulators enabling modeling of all high-end memory technologies. The Mess simulator is fast, easy to integrate and it closely matches the actual system performance. By design, it enables a quick adoption of new memory technologies in hardware simulators. Finally, the Mess application profiling positions the application in the bandwidth-latency space of the target memory system. This information can be correlated with other application runtime activities and the source code, leading to a better overall understanding of the application's behavior. The current Mess benchmark release covers all major CPU and GPU ISAs, x86, ARM, Power, RISC-V, and NVIDIA's PTX. We also release as open source the ZSim, gem5 and OpenPiton Metro-MPI integrated with the Mess simulator for DDR4, DDR5, Optane, HBM2, HBM2E and CXL memory expanders. The Mess application profiling is already integrated into a suite of production HPC performance analysis tools. Pouya Esmaili-Dokht, Francesco Sgherzi, Valéria Soldera Girelli, Isaac Boixaderas, Mariana Carmin, Alireza Monemi, Adrià Armejach, Estanislao Mercadal, Germán Llort, Petar Radojkovic, Miquel Moretó, Judit Giménez, Xavier Martorell, Eduard Ayguadé, Jesús Labarta, Emanuele Confalonieri, Rishabh Dubey, Jason Adlard |
MICRO | 6 |
| 2022 | Parallel IFFT/FFT for MIMO-OFDM LTE on NoC-Based FPGA
Kais Jallouli, Azer Hasnaoui, Jean-Philippe Diguet, Alireza Monemi, Salem Hasnaoui |
AINA (1) | 4 |
| 2022 | MIMO-OFDM LTE System based on a parallel IFFT/FFT on a multiprocessor platformabstractThis paper proposes a software workflow used to develop and evaluate a real time MIMO-OFDM LTE communication system. The central focus in this work is on OFDM modulation/demodulation functions which induce most of the processing time. To guarantee low latency, low energy dissipation, and high bandwidth requirements, we propose a multicore framework based on a parallel IFFT algorithm. We perform an accurate design space exploration by varying the number of tiles and by changing some NoC parameters. Based on our DSE study, the latency, bandwidth, and total energy dissipation are analyzed to select the best architecture design for the OFDM LTE system. In the worst case of a 20 MHz channel bandwidth, the best architecture selected for the OFDM LTE system is implemented with a hybrid network-on-chip using 16 tiles computing IFFT tasks. It leads to 84% reduction in latency, 73% increase in bandwidth in comparison with traditional OFDM LTE system using a single processing tile for the IFFT task. Kais Jallouli, Azer Hasnaoui, Jean-Philippe Diguet, Alireza Monemi, Salem Hasnaoui |
IWCMC | 4 |
| 2021 | PIugSMART: a pluggable open-source module to implement multihop bypass in networks-on-chipabstractThe integration of many processing elements per die makes it more difficult to provide low latency in the Network-on-Chip (NoC). Multihop bypass proposals, such as SMART, attack this problem by allowing flits to skip multiple routers in the path in a single cycle, drastically reducing latency while preserving a regular tiled layout. However, multihop bypass routers are more complex and relatively different from traditional NoC routers, since they rely on global broadcast signals and global allocation mechanisms. Additionally, the maximum number of nodes that can be bypassed within a single cycle is limited by the Critical Path Delay (CPD) of the NoC. Hence, a practical multihop bypass mechanism must also minimize this delay. Alireza Monemi, Ivan Perez 0004, Neiel Leyva, Enrique Vallejo 0001, Ramón Beivide, Miquel Moretó |
NOCS | 1 |
| 2015 | Virtual Channel and Switch Allocation for Low Latency Network-on-Chip RoutersabstractNetwork-on-chip (NoC) is an emerging interconnect infrastructure to address the scalability limitation of conventional shared bus. To reduce the latency of NoC router, virtual channel (VC) and switch allocation stages are performed concurrently by relaxing the dependency between these two stages. Alireza Monemi, Chia Yee Ooi, Muhammad N. Marsono |
FCCM | 1 |
| 2013 | Online NetFPGA decision tree statistical traffic classifier
Alireza Monemi, Roozbeh Zarei, Muhammad N. Marsono |
Comput. Commun. | 1 |